Plotting with Pandas and Matplotlib, the first thing that goes wrong is non-Latin text rendering as boxes. If you need Japanese labels, the fix is pip install matplotlib-fontja and importing it. That is the whole solution.
The commonly repeated advice to “set plt.rcParams['font.family'] = 'Arial' for Japanese” is simply wrong: Arial contains no Japanese glyphs, so it fixes nothing.
This article starts from environment setup, then covers preprocessing in Pandas, line, bar, histogram, scatter and box plots, and finally the API changes in pandas 2.x and Matplotlib 3.10 that break older code.
Sponsored
Getting non-Latin labels to render
The shortest route is matplotlib-fontja, which detects and configures a Japanese font on import.
pip install pandas matplotlib matplotlib-fontja
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib_fontja # importing is enough
plt.plot([1, 2, 3], [10, 20, 15])
plt.title("売上の推移")
plt.xlabel("月")
plt.ylabel("金額(万円)")
plt.show()
japanize-matplotlib used to be the standard choice, but it broke on import for a period after Python 3.12 removed distutils (since fixed). For new projects, matplotlib-fontja is the safer pick.
Without adding a library
If you cannot install anything extra, name fonts already present on the system.
import matplotlib.pyplot as plt
plt.rcParams['font.family'] = 'sans-serif'
plt.rcParams['font.sans-serif'] = [
'Hiragino Sans', # macOS
'Yu Gothic', # Windows
'Meiryo', # Windows
'Noto Sans CJK JP', # Linux
'IPAexGothic',
]
plt.rcParams['axes.unicode_minus'] = False # stop minus signs breaking
Include axes.unicode_minus = False. Once you switch to a CJK font, negative axis labels alone often turn into boxes, because most Japanese fonts lack the specific minus sign (U+2212) Matplotlib uses.
Loading and preparing data with Pandas
Fixing types at load time saves work later.
import pandas as pd
df = pd.read_csv(
"sample_data.csv",
parse_dates=["date"], # parse dates on read
encoding="utf-8-sig", # handles the BOM Excel writes
)
print(df.head())
print(df.dtypes)
encoding="utf-8-sig" matters more than it looks. CSVs exported from Excel carry a BOM, and reading them as utf-8 gives you a first column named date that you cannot reference.
Missing values
# count missing values per column
print(df.isna().sum())
# 1. drop rows
df = df.dropna(subset=["value"])
# 2. fill with the mean
df["value"] = df["value"].fillna(df["value"].mean())
Avoid inplace=True. It is being deprecated in pandas 2.x, and it obscures whether a copy occurs. Assign the result back as above.
Sponsored
Line charts: time series
df = df.sort_values("date") # always sort before plotting a time series
fig, ax = plt.subplots(figsize=(10, 4))
ax.plot(df["date"], df["value"], marker="o", linestyle="-")
ax.set_title("Value over time")
ax.set_xlabel("Date")
ax.set_ylabel("Value")
ax.grid(alpha=.3)
fig.autofmt_xdate() # angle the date labels so they do not collide
plt.show()
- Call
sort_values()first—unsorted dates produce a line that zigzags backwards - Use the
fig, ax = plt.subplots()form. Callingplt.plot()directly stops being predictable as soon as you have more than one chart
Bar charts: comparing categories
summary = (
df.groupby("category", as_index=False)["value"]
.sum()
.sort_values("value", ascending=False) # largest first
)
fig, ax = plt.subplots(figsize=(8, 4))
ax.bar(summary["category"], summary["value"])
ax.set_title("Total by category")
ax.set_xlabel("Category")
ax.set_ylabel("Total")
ax.bar_label(ax.containers[0], fmt="%.0f")
plt.show()
Always sort bars by value. Leaving them in alphabetical or insertion order defeats the purpose of a comparison chart. ax.bar_label() adds the numbers, which makes the chart usable in a report as-is.
Sponsored
Histograms: the shape of the distribution
fig, ax = plt.subplots(figsize=(8, 4))
ax.hist(df["value"].dropna(), bins=20, edgecolor="black")
ax.set_title("Distribution of values")
ax.set_xlabel("Value")
ax.set_ylabel("Frequency")
plt.show()
The bins count changes what you see. Too few flattens the peaks; too many produces noise. Start with bins="auto" and adjust from there.
Scatter plots
fig, ax = plt.subplots(figsize=(6, 6))
ax.scatter(df["feature1"], df["feature2"], alpha=0.4, s=20)
ax.set_title("Relationship between two variables")
ax.set_xlabel("Feature 1")
ax.set_ylabel("Feature 2")
plt.show()
Lowering alpha is standard practice. Past a few thousand points, overlap saturates and density differences disappear. If that is not enough, switch to ax.hexbin().
Box plots: an API that changed
This is where old code breaks. Matplotlib 3.10 (December 2024) deprecated the vert argument in favour of orientation.
groups = [g["value"].dropna().values for _, g in df.groupby("category")]
labels = [name for name, _ in df.groupby("category")]
fig, ax = plt.subplots(figsize=(8, 4))
# old: ax.boxplot(groups, vert=False)
# new: use orientation
ax.boxplot(groups, tick_labels=labels, orientation="horizontal")
ax.set_title("Spread by category")
ax.set_xlabel("Value")
plt.show()
The label argument was also renamed from labels to tick_labels. Running old code emits warnings for both, so update them together.
df.boxplot(column="value", by="category") also works, but it adds an automatic “Boxplot grouped by category” title that plt.title() cannot remove. For charts going into a document, calling ax.boxplot() directly is easier to control.
Correlation heatmaps: another pandas 2.x change
Calling df.corr() with string columns present now raises an error.
import seaborn as sns
# raises ValueError when string columns are present
# corr = df.corr()
# restrict to numeric columns
corr = df.corr(numeric_only=True)
fig, ax = plt.subplots(figsize=(8, 6))
sns.heatmap(corr, annot=True, fmt=".2f", cmap="coolwarm",
vmin=-1, vmax=1, ax=ax)
ax.set_title("Correlation heatmap")
plt.show()
Do not omit vmin=-1, vmax=1. Without them the colour scale auto-fits the data range, and a correlation of 0.3 can appear bright red. Correlation matrices should always be pinned to −1 through 1.
Practice datasets
If you have no data to hand, seaborn’s bundled datasets are the quickest option and do not require installing scikit-learn.
import seaborn as sns
df = sns.load_dataset("iris")
# df = sns.load_dataset("titanic")
# df = sns.load_dataset("tips")
print(df.head())
print(df.describe())
Loading requires network access the first time (it fetches from GitHub). Offline, use load_iris() from scikit-learn.
df.hist(figsize=(10, 6), bins=20)
plt.tight_layout() # resolve overlapping labels
plt.show()
Call plt.tight_layout() last whenever you have a grid of charts, or axis labels will run into the neighbouring plot.
Saving to a file
fig.savefig("chart.png", dpi=150, bbox_inches="tight")
bbox_inches="tight": prevents axis labels being cut off. Without it, rotated date labels lose their bottom edgedpi=150: the default of 100 looks coarse when pasted into a document- Call it before
plt.show(): after display the figure may be discarded, saving a blank image
Summary
- Non-Latin labels: import
matplotlib-fontja. Setting Arial does nothing - Broken minus signs:
axes.unicode_minus = False - Excel CSVs:
encoding="utf-8-sig" - Do not use
inplace=True; assign the result back - Sort time series by date, and bar charts by value, before plotting
- Box plots:
vertis deprecated in Matplotlib 3.10—useorientationandtick_labels df.corr()needsnumeric_only=True; heatmaps needvmin=-1, vmax=1- Save with
bbox_inches="tight", beforeplt.show()
If you want statistical charts with less code, Seaborn is shorter. I have also written about building a small image-compression GUI in Python.