Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Multi-column `groupby` crashes on non-string observations in dotplot/matrixplot/heatmap

Open Beginner friendly
#4,384 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
82/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
numpy, pandas, python

Research direction

Start in plotting/legacy/_anndata.py at _prepare_dataframe, then trace how BasePlot.init invokes it. Reproduce the minimal dotplot example and check the corresponding matrixplot, stacked_violin, tracksplot, and heatmap paths; done means multi-column groupby accepts numeric and categorical observations without the reported errors.

Written by the indexing model from the issue text.

Description

Please make sure these conditions are met
  • I have checked that this issue has not already been reported.
  • I have confirmed this bug exists on the latest version of scanpy.
  • (optional) I have confirmed this bug exists on the main branch of scanpy.
What happened?

Passing a list of groupby keys fails with TypeError when any key is a non-string observation, whether an integer categorical or a plain numeric column. A single-key groupby handles the same column fine (is_numeric_dtype → pd.cut, or astype("category")), so this only affects the multi-key path.

Reproduces in sc.pl.dotplot, sc.pl.matrixplot, sc.pl.stacked_violin, sc.pl.tracksplot, and sc.pl.heatmap: all go through _prepare_dataframe in plotting/legacy/_anndata.py (BasePlot subclasses call it from BasePlot.__init__).

Two problems in the multi-key branch:

  1. obs_tidy[groupby].apply("_".join, axis=1) requires every joined value to be str. Integer categories or numeric columns raise TypeError: sequence item 1: expected str instance, int found.
  2. obs_tidy[g].cat.categories assumes each groupby column is categorical. A plain numeric column would raise AttributeError at this line if it got past the join.
Minimal code sample
import numpy as np
import pandas as pd
import scanpy as sc

adata = sc.datasets.pbmc68k_reduced()
adata.obs["num_grp"] = pd.Categorical(np.repeat([1, 2], adata.n_obs // 2))

# single key works
sc.pl.dotplot(adata, var_names=adata.var_names[:4], groupby="num_grp", show=False)

# multi key raises TypeError
sc.pl.dotplot(adata, var_names=adata.var_names[:4], groupby=["bulk_labels", "num_grp"], show=False)
Error output
TypeError: sequence item 1: expected str instance, int found

I have a fix ready: convert joined values with astype(str), take levels from .cat.categories only for categorical columns (np.unique otherwise), and fall back to end-ordering for joined labels missing from the order map (NaN-containing combinations). Happy to open a PR.

Dominant language
Python
Stars
2.6k
Forks
779
Avg merge
16h 57m
Merged PRs (30d)
21

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from scverse/scanpy

All issues in scverse/scanpy

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.