Online Pingouin Compiler

Run Pingouin code in your browser. Run t-tests, ANOVA with post-hoc tests, correlations and assumption checks, with effect sizes.

Python
import pingouin as pg
import pandas as pd

# One-sample t-test
ttest = pg.ttest([5.5, 2.4, 6.8, 9.6, 4.2], 4.0)
print("One-sample t-test:")
print(ttest.to_string())

# Pearson correlation
df = pd.DataFrame({
    'x': [1, 2, 3, 4, 5, 6, 7],
    'y': [2, 4, 5, 4, 5, 7, 9]
})
corr = pg.corr(df['x'], df['y'])
print("\nPearson correlation:")
print(corr.to_string())

# One-way ANOVA
data = pd.DataFrame({
    'score': [23, 25, 28, 30, 22, 26, 33, 35, 38, 40, 29, 32],
    'group': ['A']*4 + ['B']*4 + ['C']*4
})
aov = pg.anova(data=data, dv='score', between='group')
print("\nOne-way ANOVA:")
print(aov.to_string())
One-sample t-test:
               T  dof alternative     p_val          CI95   cohen_d     power   BF10
T_test  1.397391    4   two-sided  0.234824  [2.32, 9.08]  0.624932  0.191796  0.766

Pearson correlation:
         n         r          CI95     p_val    BF10     power
pearson  7  0.918559  [0.54, 0.99]  0.003478  14.059  0.910838

One-way ANOVA:
  Source  ddof1  ddof2         F     p_unc       np2
0  group      2      9  2.958668  0.102915  0.396675

Pingouin is a statistics package built on pandas and SciPy. One function call runs a common test, such as a t-test, an ANOVA or a correlation, and returns a pandas DataFrame with the statistic, the p-value, a confidence interval and an effect size, often with the power and a Bayes factor as well. Students, psychologists and other researchers use it for the tests in a typical results section. This page is an online Pingouin compiler: the code runs in your browser, so you can try it without installing anything. Run the example first, then paste any snippet below into a new cell to try it.

What the example does

pg.ttest() with a list and a single number runs a one-sample t-test, which checks whether the mean of the five values, 5.7, differs from 4.0. Each function returns a DataFrame, and to_string() prints every column. The row holds T, the degrees of freedom, the p-value (0.235), the 95% confidence interval of the mean, Cohen's d, the power and the Bayes factor BF10. pg.corr() finds a Pearson correlation of 0.92 between x and y, with p = 0.003. pg.anova() takes a long-format DataFrame: dv names the column of measurements and between the column of group labels. With F = 2.96 and p_unc = 0.103, the three group means do not differ significantly at the 5% level. np2 is the effect size, partial eta squared.

Compare two groups

Two samples give an independent t-test. The groups have different sizes, so Pingouin applies Welch's correction, which shows as a fractional dof. pg.mwu() runs the Mann-Whitney U test, which does not assume normal data:

import numpy as np
import pingouin as pg

rng = np.random.default_rng(seed=1)
control = rng.normal(50, 8, 30)
treated = rng.normal(55, 8, 25)

print(pg.ttest(treated, control).round(3).to_string())
print(pg.mwu(treated, control).round(3).to_string())

With equal group sizes, pass correction=True to get Welch's test.

Find which groups differ after an ANOVA

A significant ANOVA says that the means are not all equal, not which ones differ. pg.pairwise_tukey() compares every pair and adjusts the p-values for the number of comparisons:

import pandas as pd
import pingouin as pg

df = pd.DataFrame({
    "score": [23, 25, 28, 30, 22, 26, 33, 35, 38, 40, 29, 32,
              41, 44, 39, 45],
    "group": ["A"] * 4 + ["B"] * 4 + ["C"] * 4 + ["D"] * 4,
})
print(pg.anova(data=df, dv="score", between="group").round(4).to_string())
print(pg.pairwise_tukey(data=df, dv="score", between="group").round(3).to_string())

The ANOVA gives p = 0.0015, but only A–D and B–D have a p_tukey below 0.05.

Compare paired measurements

When each subject is measured twice, pass paired=True. pg.wilcoxon() is the rank-based alternative:

import pingouin as pg

before = [72, 80, 65, 90, 77, 84, 69, 75]
after = [70, 76, 66, 85, 72, 80, 64, 73]

print(pg.ttest(before, after, paired=True).round(3).to_string())
print(pg.wilcoxon(before, after).round(3).to_string())

The scores drop by 3.25 points on average, and CI95 is the 95% confidence interval for that drop.

Check the assumptions

pg.normality() runs a Shapiro-Wilk test for each group, and pg.homoscedasticity() runs Levene's test for equal variances. When the variances differ, use pg.welch_anova() instead of pg.anova():

import numpy as np
import pandas as pd
import pingouin as pg

rng = np.random.default_rng(seed=4)
df = pd.DataFrame({
    "score": np.concatenate([rng.normal(60, 3, 20), rng.normal(64, 3, 20),
                             rng.normal(66, 9, 20)]),
    "group": ["A"] * 20 + ["B"] * 20 + ["C"] * 20,
})
print(pg.normality(df, dv="score", group="group").round(3))
print(pg.homoscedasticity(df, dv="score", group="group").round(3))
print(pg.welch_anova(data=df, dv="score", between="group").round(4).to_string())

Group C is three times as spread out as A and B, so Levene's test rejects equal variances with p = 0.007.

Good to know

  • Pingouin 0.6 renamed the columns p-val, CI95% and cohen-d to p_val, CI95 and cohen_d. Code from older tutorials that reads res["p-val"] raises a KeyError.
  • normality() and homoscedasticity() still call their p-value column pval, without the underscore.