Browse by section

Python 日本語

Data Visualisation with Pandas and Matplotlib

Plotting with Pandas and Matplotlib, the first thing that goes wrong is non-Latin text rendering as boxes. If you need Japanese labels, the fix is pip install matplotlib-fontja and importing it. That is the whole solution.

The commonly repeated advice to “set plt.rcParams['font.family'] = 'Arial' for Japanese” is simply wrong: Arial contains no Japanese glyphs, so it fixes nothing.

This article starts from environment setup, then covers preprocessing in Pandas, line, bar, histogram, scatter and box plots, and finally the API changes in pandas 2.x and Matplotlib 3.10 that break older code.

Sponsored

Getting non-Latin labels to render

The shortest route is matplotlib-fontja, which detects and configures a Japanese font on import.

pip install pandas matplotlib matplotlib-fontja
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib_fontja  # importing is enough

plt.plot([1, 2, 3], [10, 20, 15])
plt.title("売上の推移")
plt.xlabel("月")
plt.ylabel("金額(万円)")
plt.show()

japanize-matplotlib used to be the standard choice, but it broke on import for a period after Python 3.12 removed distutils (since fixed). For new projects, matplotlib-fontja is the safer pick.

Without adding a library

If you cannot install anything extra, name fonts already present on the system.

import matplotlib.pyplot as plt

plt.rcParams['font.family'] = 'sans-serif'
plt.rcParams['font.sans-serif'] = [
    'Hiragino Sans',       # macOS
    'Yu Gothic',           # Windows
    'Meiryo',              # Windows
    'Noto Sans CJK JP',    # Linux
    'IPAexGothic',
]
plt.rcParams['axes.unicode_minus'] = False  # stop minus signs breaking

Include axes.unicode_minus = False. Once you switch to a CJK font, negative axis labels alone often turn into boxes, because most Japanese fonts lack the specific minus sign (U+2212) Matplotlib uses.

Loading and preparing data with Pandas

Fixing types at load time saves work later.

import pandas as pd

df = pd.read_csv(
    "sample_data.csv",
    parse_dates=["date"],   # parse dates on read
    encoding="utf-8-sig",   # handles the BOM Excel writes
)

print(df.head())
print(df.dtypes)

encoding="utf-8-sig" matters more than it looks. CSVs exported from Excel carry a BOM, and reading them as utf-8 gives you a first column named date that you cannot reference.

Missing values

# count missing values per column
print(df.isna().sum())

# 1. drop rows
df = df.dropna(subset=["value"])

# 2. fill with the mean
df["value"] = df["value"].fillna(df["value"].mean())

Avoid inplace=True. It is being deprecated in pandas 2.x, and it obscures whether a copy occurs. Assign the result back as above.

Sponsored

Line charts: time series

df = df.sort_values("date")   # always sort before plotting a time series

fig, ax = plt.subplots(figsize=(10, 4))
ax.plot(df["date"], df["value"], marker="o", linestyle="-")

ax.set_title("Value over time")
ax.set_xlabel("Date")
ax.set_ylabel("Value")
ax.grid(alpha=.3)

fig.autofmt_xdate()   # angle the date labels so they do not collide
plt.show()
  • Call sort_values() first—unsorted dates produce a line that zigzags backwards
  • Use the fig, ax = plt.subplots() form. Calling plt.plot() directly stops being predictable as soon as you have more than one chart

Bar charts: comparing categories

summary = (
    df.groupby("category", as_index=False)["value"]
      .sum()
      .sort_values("value", ascending=False)   # largest first
)

fig, ax = plt.subplots(figsize=(8, 4))
ax.bar(summary["category"], summary["value"])

ax.set_title("Total by category")
ax.set_xlabel("Category")
ax.set_ylabel("Total")

ax.bar_label(ax.containers[0], fmt="%.0f")
plt.show()

Always sort bars by value. Leaving them in alphabetical or insertion order defeats the purpose of a comparison chart. ax.bar_label() adds the numbers, which makes the chart usable in a report as-is.

Sponsored

Histograms: the shape of the distribution

fig, ax = plt.subplots(figsize=(8, 4))
ax.hist(df["value"].dropna(), bins=20, edgecolor="black")

ax.set_title("Distribution of values")
ax.set_xlabel("Value")
ax.set_ylabel("Frequency")
plt.show()

The bins count changes what you see. Too few flattens the peaks; too many produces noise. Start with bins="auto" and adjust from there.

Scatter plots

fig, ax = plt.subplots(figsize=(6, 6))
ax.scatter(df["feature1"], df["feature2"], alpha=0.4, s=20)

ax.set_title("Relationship between two variables")
ax.set_xlabel("Feature 1")
ax.set_ylabel("Feature 2")
plt.show()

Lowering alpha is standard practice. Past a few thousand points, overlap saturates and density differences disappear. If that is not enough, switch to ax.hexbin().

Box plots: an API that changed

This is where old code breaks. Matplotlib 3.10 (December 2024) deprecated the vert argument in favour of orientation.

groups = [g["value"].dropna().values for _, g in df.groupby("category")]
labels = [name for name, _ in df.groupby("category")]

fig, ax = plt.subplots(figsize=(8, 4))

# old: ax.boxplot(groups, vert=False)
# new: use orientation
ax.boxplot(groups, tick_labels=labels, orientation="horizontal")

ax.set_title("Spread by category")
ax.set_xlabel("Value")
plt.show()

The label argument was also renamed from labels to tick_labels. Running old code emits warnings for both, so update them together.

df.boxplot(column="value", by="category") also works, but it adds an automatic “Boxplot grouped by category” title that plt.title() cannot remove. For charts going into a document, calling ax.boxplot() directly is easier to control.

Correlation heatmaps: another pandas 2.x change

Calling df.corr() with string columns present now raises an error.

import seaborn as sns

# raises ValueError when string columns are present
# corr = df.corr()

# restrict to numeric columns
corr = df.corr(numeric_only=True)

fig, ax = plt.subplots(figsize=(8, 6))
sns.heatmap(corr, annot=True, fmt=".2f", cmap="coolwarm",
            vmin=-1, vmax=1, ax=ax)

ax.set_title("Correlation heatmap")
plt.show()

Do not omit vmin=-1, vmax=1. Without them the colour scale auto-fits the data range, and a correlation of 0.3 can appear bright red. Correlation matrices should always be pinned to −1 through 1.

Practice datasets

If you have no data to hand, seaborn’s bundled datasets are the quickest option and do not require installing scikit-learn.

import seaborn as sns

df = sns.load_dataset("iris")
# df = sns.load_dataset("titanic")
# df = sns.load_dataset("tips")

print(df.head())
print(df.describe())

Loading requires network access the first time (it fetches from GitHub). Offline, use load_iris() from scikit-learn.

df.hist(figsize=(10, 6), bins=20)
plt.tight_layout()   # resolve overlapping labels
plt.show()

Call plt.tight_layout() last whenever you have a grid of charts, or axis labels will run into the neighbouring plot.

Saving to a file

fig.savefig("chart.png", dpi=150, bbox_inches="tight")
  • bbox_inches="tight": prevents axis labels being cut off. Without it, rotated date labels lose their bottom edge
  • dpi=150: the default of 100 looks coarse when pasted into a document
  • Call it before plt.show(): after display the figure may be discarded, saving a blank image

Summary

  • Non-Latin labels: import matplotlib-fontja. Setting Arial does nothing
  • Broken minus signs: axes.unicode_minus = False
  • Excel CSVs: encoding="utf-8-sig"
  • Do not use inplace=True; assign the result back
  • Sort time series by date, and bar charts by value, before plotting
  • Box plots: vert is deprecated in Matplotlib 3.10—use orientation and tick_labels
  • df.corr() needs numeric_only=True; heatmaps need vmin=-1, vmax=1
  • Save with bbox_inches="tight", before plt.show()

If you want statistical charts with less code, Seaborn is shorter. I have also written about building a small image-compression GUI in Python.