Browse by section

Python 日本語

Data Analysis and Visualisation with Pandas and Seaborn

The reason to use Seaborn is that one line gives you aggregation, plotting, a legend, and error bars. What takes ten lines in Matplotlib becomes sns.barplot(data=df, x="day", y="total_bill").

The catch is that Seaborn assumes tidy data, so if your data is not shaped correctly nothing works the way you expect. Labels in non-Latin scripts also need font configuration, exactly as in Matplotlib.

This article covers reshaping in Pandas, the main Seaborn chart types, and the seaborn.objects interface the project now recommends for new code.

Sponsored

Setup and font configuration

pip install pandas seaborn matplotlib matplotlib-fontja
import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt
import matplotlib_fontja  # prevents CJK text rendering as boxes

sns.set_theme(style="whitegrid")
plt.rcParams['axes.unicode_minus'] = False

One ordering detail matters: sns.set_theme() overwrites font settings. Calling set_theme() before importing matplotlib_fontja can bring the boxes back. Import first, then call set_theme(), as above.

If characters still break, pass the font to the theme:

sns.set_theme(style="whitegrid", font="Hiragino Sans")  # macOS
# sns.set_theme(style="whitegrid", font="Yu Gothic")    # Windows

The data shape Seaborn expects

This is the first hurdle. Seaborn assumes long format: one row per observation, one column per variable. The “months across the top” table people build in Excel—wide format—has no column to pass to x= or hue=.

# wide format (the usual spreadsheet shape)
wide = pd.DataFrame({
    "product": ["A", "B"],
    "Jan": [100, 80],
    "Feb": [120, 90],
    "Mar": [140, 70],
})

# convert to long format
long = wide.melt(id_vars="product", var_name="month", value_name="sales")
print(long)
#   product month  sales
# 0       A   Jan    100
# 1       B   Jan     80
# 2       A   Feb    120
# ...

Once it is long, you can pass it straight in:

sns.lineplot(data=long, x="month", y="sales", hue="product", marker="o")
plt.show()

Passing a column name to hue= splits the series and builds the legend for you—that is Seaborn’s core value. Conversely, without the right shape you get almost none of it. When a Seaborn chart will not come out right, the data shape is the cause about nine times out of ten.

Sponsored

Preprocessing in Pandas

df = pd.read_csv("data.csv", encoding="utf-8-sig", parse_dates=["date"])

print(df.isna().sum())

df = df.dropna(subset=["value"])
df["value"] = df["value"].fillna(df["value"].mean())

adults = df[df["age"] >= 20]

summary = df.groupby("category", as_index=False)["value"].agg(["mean", "sum", "count"])

encoding="utf-8-sig" handles the BOM in CSVs exported from Excel; reading as utf-8 corrupts the first column name.

Avoid inplace=True and assign the result back, as above. It is being deprecated in pandas 2.x and makes copy semantics hard to reason about.

The four charts you will use most

barplot: comparing means

tips = sns.load_dataset("tips")

sns.barplot(data=tips, x="day", y="total_bill", hue="sex")
plt.title("Average bill by day and sex")
plt.show()

barplot plots the mean, not the sum, and the vertical line is a 95% confidence interval. For totals, pass estimator="sum" or aggregate with groupby().sum() first. This is the most common misreading.

boxplot: spread

sns.boxplot(data=tips, x="day", y="total_bill")
plt.show()

With small samples—a few dozen points—stripplot or swarmplot showing the actual points is more honest. Quartiles of three data points mean nothing.

sns.boxplot(data=tips, x="day", y="total_bill", showfliers=False)
sns.stripplot(data=tips, x="day", y="total_bill", color=".25", size=3, alpha=.5)
plt.show()

histplot: distribution

sns.histplot(data=tips, x="total_bill", bins="auto", kde=True)
plt.show()

kde=True overlays a density estimate. Leave bins unset the first time and look at the shape before tuning.

scatterplot: relationships

sns.scatterplot(data=tips, x="total_bill", y="tip", hue="time", size="size", alpha=.7)
plt.show()

You can map separate columns to hue (colour), size, and style (marker shape), encoding up to four variables in one chart. In practice, stop at three—beyond that it stops being readable.

Sponsored

Correlation heatmaps: a pandas 2.x change

Calling df.corr() with string columns present now raises an error. pandas 1.x silently used only numeric columns; 2.x requires you to say so.

# raises ValueError when string columns are present
# corr = tips.corr()

corr = tips.corr(numeric_only=True)

fig, ax = plt.subplots(figsize=(6, 5))
sns.heatmap(corr, annot=True, fmt=".2f", cmap="coolwarm",
            vmin=-1, vmax=1, square=True, ax=ax)
ax.set_title("Correlation heatmap")
plt.show()

vmin=-1, vmax=1 is essential. Without it, the colour scale fits the data range and a correlation of 0.3 can look bright red. Pin the scale for correlation matrices.

pairplot: everything against everything

iris = sns.load_dataset("iris")

sns.pairplot(iris, hue="species", diag_kind="hist", corner=True)
plt.show()

corner=True drops the upper triangle, removing duplicated information and making it far easier to read.

Be careful: pairplot scales with the square of the column count. Calling it on a frame with more than twenty columns will freeze for minutes. Narrow it with vars=["colA", "colB", "colC"].

Regression plots need numeric columns

regplot overlays a regression line on a scatter plot, and both x and y must be numeric. Passing a categorical column—city names, product names—raises an error.

# a string category cannot go on y
# sns.regplot(data=df, x="age", y="city")

sns.regplot(data=tips, x="total_bill", y="tip", scatter_kws={"alpha": .4})
plt.show()

To fit separate lines per category, use lmplot:

sns.lmplot(data=tips, x="total_bill", y="tip", hue="smoker", height=5)
plt.show()

Figure-level versus axes-level functions

Seaborn has two kinds of function, and mixing them up breaks your layout.

Kind Examples Behaviour
axes-level scatterplot, barplot, boxplot, histplot Accepts ax=; draws into an existing figure
figure-level relplot, catplot, displot, lmplot, pairplot Creates its own figure; ax= is not accepted
# axes-level: compose several
fig, axes = plt.subplots(1, 2, figsize=(12, 4))
sns.histplot(data=tips, x="total_bill", ax=axes[0])
sns.boxplot(data=tips, x="day", y="tip", ax=axes[1])
plt.tight_layout()
plt.show()

# figure-level: col= does the splitting for you
g = sns.relplot(data=tips, x="total_bill", y="tip",
                col="time", hue="smoker", height=4)
g.figure.suptitle("Bill and tip by time of day", y=1.02)
plt.show()

Figure-level functions return a FacetGrid, not an Axes. Use g.figure.suptitle() for the title and g.savefig() to save. Calling plt.title() only affects the last subplot.

For new code: the seaborn.objects interface

Seaborn 0.12 introduced a declarative interface that composes a chart from parts, closer to ggplot2 in R.

import seaborn.objects as so

(
    so.Plot(tips, x="total_bill", y="tip", color="time")
    .add(so.Dot(alpha=.5))
    .add(so.Line(), so.PolyFit())      # overlay a regression line
    .facet(col="smoker")               # split into columns
    .label(x="Bill", y="Tip", title="Bill versus tip")
    .show()
)

The classic functions still work, but separating “what to plot” from “how to show it” makes complex charts much easier to read. It is worth considering for anything new.

Saving

# axes-level
fig.savefig("chart.png", dpi=150, bbox_inches="tight")

# figure-level
g.savefig("chart.png", dpi=150, bbox_inches="tight")
  • bbox_inches="tight": stops labels and legends being clipped
  • Call it before plt.show(): after display the figure may be discarded, giving a blank file

Summary

  • Import matplotlib-fontja for CJK text, and call sns.set_theme() afterwards
  • Seaborn expects long format; convert wide data with melt()
  • barplot shows the mean, not the sum. Use estimator="sum" for totals
  • df.corr() needs numeric_only=True; heatmaps need vmin=-1, vmax=1
  • regplot requires numeric x and y. Use lmplot for per-category fits
  • pairplot scales quadratically—narrow it with vars=
  • Figure-level functions do not take ax= and return a FacetGrid
  • Consider seaborn.objects for new work

When you need fine control, dropping to Matplotlib directly is often faster. That is covered in data visualisation with Pandas and Matplotlib.