import seaborn as sns sns.lineplot(x=[1, 2, 3], y=[2, 5, 12])
Seaborn is a Python library for statistical charts, built on Matplotlib. You pass it a pandas DataFrame and name the columns to plot, and seaborn handles grouping, colors, legends and summaries such as means and confidence intervals. Data scientists use it to explore data: to see how values are spread, compare groups and spot relationships between columns. This page is an online seaborn compiler: the code runs in your browser, so you can try it without installing anything. Run the example first, then paste any snippet below into a new cell to try it.
sns.lineplot() draws a line through the points (1, 2), (2, 5) and
(3, 12). With plain lists the axes have no labels. Pass a DataFrame as
data and column names as x and y, and seaborn labels each axis
with its column name. The call returns the Matplotlib Axes it drew
on, which the cell displays because it is the last line. When several
rows share an x value, lineplot() draws their mean and shades a 95%
confidence interval around it.
hue colors each point by the value in another column and adds a
legend. style and size do the same for the marker shape and size:
import numpy as np
import pandas as pd
import seaborn as sns
rng = np.random.default_rng(seed=0)
hours = rng.uniform(0, 10, 60)
df = pd.DataFrame({
"hours": hours,
"score": 50 + 4 * hours + rng.normal(0, 6, 60),
"course": rng.choice(["Math", "History"], 60),
})
sns.scatterplot(data=df, x="hours", y="score", hue="course")
histplot() counts the values that fall in each bin. With hue, it
draws one histogram per group, overlapping in the same axes, and
kde=True adds a smoothed density curve for each. Pass
multiple="stack" to stack them instead:
import numpy as np
import pandas as pd
import seaborn as sns
rng = np.random.default_rng(seed=1)
df = pd.DataFrame({
"minutes": np.concatenate([rng.normal(30, 5, 200), rng.normal(45, 8, 200)]),
"mode": ["Train"] * 200 + ["Bus"] * 200,
})
sns.histplot(data=df, x="minutes", hue="mode", kde=True, bins=20)
Each box spans the middle half of a group's values, with a line at the median. The whiskers reach the furthest values within 1.5 box lengths of the box, and anything beyond them is drawn as a separate point, like Wednesday's 180 here:
import pandas as pd
import seaborn as sns
df = pd.DataFrame({
"day": ["Mon"] * 5 + ["Tue"] * 5 + ["Wed"] * 5,
"sales": [120, 135, 128, 150, 110, 140, 155, 160, 138, 149,
90, 105, 98, 180, 101],
})
sns.boxplot(data=df, x="day", y="sales")
sns.barplot() takes the same arguments and draws each group's mean,
with a 95% confidence interval as the error bar.
df.corr() gives the correlation between each pair of columns, from -1
to 1. annot=True writes each value in its cell, and vmin=-1 and
vmax=1 fix the color scale: with cmap="coolwarm", negative values
are blue, positive ones red and values near 0 gray:
import numpy as np
import pandas as pd
import seaborn as sns
rng = np.random.default_rng(seed=2)
temp = rng.normal(20, 5, 100)
df = pd.DataFrame({
"temperature": temp,
"ice_cream": 3 * temp + rng.normal(0, 5, 100),
"umbrellas": -2 * temp + rng.normal(0, 10, 100),
"noise": rng.normal(0, 1, 100),
})
sns.heatmap(df.corr(), annot=True, fmt=".2f", cmap="coolwarm", vmin=-1, vmax=1)
relplot(), displot(), catplot(), pairplot() and jointplot()
return a grid object, not an Axes, and on this page the cell prints
the object instead of the chart. Add import matplotlib.pyplot as plt
and end the cell with plt.show() to display it.scatterplot() and histplot() in one cell draw on
the same Axes. For side-by-side plots, create
fig, axes = plt.subplots(1, 2) and pass ax=axes[0] to one call and
ax=axes[1] to the other.sns.load_dataset(), which the seaborn docs use for their examples,
downloads the data from GitHub, so it needs an internet connection.