import pingouin as pg
import pandas as pd
# One-sample t-test
ttest = pg.ttest([5.5, 2.4, 6.8, 9.6, 4.2], 4.0)
print("One-sample t-test:")
print(ttest.to_string())
# Pearson correlation
df = pd.DataFrame({
'x': [1, 2, 3, 4, 5, 6, 7],
'y': [2, 4, 5, 4, 5, 7, 9]
})
corr = pg.corr(df['x'], df['y'])
print("\nPearson correlation:")
print(corr.to_string())
# One-way ANOVA
data = pd.DataFrame({
'score': [23, 25, 28, 30, 22, 26, 33, 35, 38, 40, 29, 32],
'group': ['A']*4 + ['B']*4 + ['C']*4
})
aov = pg.anova(data=data, dv='score', between='group')
print("\nOne-way ANOVA:")
print(aov.to_string())Pingouin is a statistics package built on pandas and SciPy. One function call runs a common test, such as a t-test, an ANOVA or a correlation, and returns a pandas DataFrame with the statistic, the p-value, a confidence interval and an effect size, often with the power and a Bayes factor as well. Students, psychologists and other researchers use it for the tests in a typical results section. This page is an online Pingouin compiler: the code runs in your browser, so you can try it without installing anything. Run the example first, then paste any snippet below into a new cell to try it.
pg.ttest() with a list and a single number runs a one-sample t-test,
which checks whether the mean of the five values, 5.7, differs from
4.0. Each function returns a DataFrame, and to_string() prints every
column. The row holds T, the degrees of freedom, the p-value (0.235),
the 95% confidence interval of the mean, Cohen's d, the power and the
Bayes factor BF10. pg.corr() finds a Pearson correlation of 0.92 between
x and y, with p = 0.003. pg.anova() takes a long-format DataFrame:
dv names the column of measurements and between the column of group
labels. With F = 2.96 and p_unc = 0.103, the three group means do not
differ significantly at the 5% level. np2 is the effect size, partial
eta squared.
Two samples give an independent t-test. The groups have different
sizes, so Pingouin applies Welch's correction, which shows as a
fractional dof. pg.mwu() runs the Mann-Whitney U test, which does
not assume normal data:
import numpy as np
import pingouin as pg
rng = np.random.default_rng(seed=1)
control = rng.normal(50, 8, 30)
treated = rng.normal(55, 8, 25)
print(pg.ttest(treated, control).round(3).to_string())
print(pg.mwu(treated, control).round(3).to_string())
With equal group sizes, pass correction=True to get Welch's test.
A significant ANOVA says that the means are not all equal, not which
ones differ. pg.pairwise_tukey() compares every pair and adjusts the
p-values for the number of comparisons:
import pandas as pd
import pingouin as pg
df = pd.DataFrame({
"score": [23, 25, 28, 30, 22, 26, 33, 35, 38, 40, 29, 32,
41, 44, 39, 45],
"group": ["A"] * 4 + ["B"] * 4 + ["C"] * 4 + ["D"] * 4,
})
print(pg.anova(data=df, dv="score", between="group").round(4).to_string())
print(pg.pairwise_tukey(data=df, dv="score", between="group").round(3).to_string())
The ANOVA gives p = 0.0015, but only A–D and B–D have a p_tukey
below 0.05.
When each subject is measured twice, pass paired=True. pg.wilcoxon()
is the rank-based alternative:
import pingouin as pg
before = [72, 80, 65, 90, 77, 84, 69, 75]
after = [70, 76, 66, 85, 72, 80, 64, 73]
print(pg.ttest(before, after, paired=True).round(3).to_string())
print(pg.wilcoxon(before, after).round(3).to_string())
The scores drop by 3.25 points on average, and CI95 is the 95%
confidence interval for that drop.
pg.normality() runs a Shapiro-Wilk test for each group, and
pg.homoscedasticity() runs Levene's test for equal variances. When the
variances differ, use pg.welch_anova() instead of pg.anova():
import numpy as np
import pandas as pd
import pingouin as pg
rng = np.random.default_rng(seed=4)
df = pd.DataFrame({
"score": np.concatenate([rng.normal(60, 3, 20), rng.normal(64, 3, 20),
rng.normal(66, 9, 20)]),
"group": ["A"] * 20 + ["B"] * 20 + ["C"] * 20,
})
print(pg.normality(df, dv="score", group="group").round(3))
print(pg.homoscedasticity(df, dv="score", group="group").round(3))
print(pg.welch_anova(data=df, dv="score", between="group").round(4).to_string())
Group C is three times as spread out as A and B, so Levene's test rejects equal variances with p = 0.007.
p-val, CI95% and cohen-d to
p_val, CI95 and cohen_d. Code from older tutorials that reads
res["p-val"] raises a KeyError.normality() and homoscedasticity() still call their p-value
column pval, without the underscore.