Shared sparse one-hot (indicator) helper for squidpy and scanpy
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 72/100
Research direction
Start with scanpy/get/_aggregated.py and squidpy/gr/_nhood.py, then review the proposed fast_array_utils.conv entry point and its sparse extra. Confirm the shared helper preserves missing labels, unused categories, mask handling, float64 output, and the stated matrix layouts; done means both callers use it without changing their current behavior.
Written by the indexing model from the issue text.
Description
prettified with ai, basically saying we have two different functions we can unify here. low prio but good to document.
scanpy (sparse_indicator in scanpy/get/_aggregated.py) and squidpy (_onehot in squidpy/gr/_nhood.py) each keep a private helper that turns category codes into a sparse float64 indicator matrix, with a missing label (-1) giving no entry. They differ only in interface: scanpy takes a pd.Categorical and returns (n_categories, n_obs) as a coo_array with an optional mask, while squidpy takes a pd.Series and returns a (n_obs, n_categories) csr_matrix.
Proposal for fast_array_utils.conv (needs the sparse extra; takes codes since fast-array-utils doesn't depend on pandas):
def sparse_indicator(
codes: NDArray[np.integer], n_categories: int, *, mask: NDArray[np.bool] | None = None
) -> coo_array:
keep = codes >= 0 if mask is None else (codes >= 0) & mask
obs = np.flatnonzero(keep)
return coo_array((np.ones(obs.size), (obs, codes[keep])), shape=(codes.size, n_categories))
(n_obs, n_categories), the usual one-hot layout; scanpy takes.T(about 3 ms at 5M observations).- float64, as both callers use today.
- Same cost as both copies today: squidpy converts the result to CSR, as it does now.
checked with missing labels, unused categories and mask.
as discussed in squidpy/#1285
- Dominant language
- Python
- Stars
- 15
- Forks
- 5
- Avg merge
- 14h 52m
- Merged PRs (30d)
- 13
Getting set up
We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from scverse/fast-array-utils
-
Difficulty 1/5 1-3 hours Newbie friendliness 74/100
scverse/fast-array-utils#165 ·
Maintainers usually reply within 1 day
-
Difficulty 3/5 1-2 days Newbie friendliness 65/100
scverse/fast-array-utils#213 · 4 comments ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 48/100
scverse/fast-array-utils#179 · 23 comments ·
Maintainers usually reply within 1 day
-
type: numpy/scipy
Difficulty 5/5 Over a week Newbie friendliness 25/100
scverse/fast-array-utils#128 ·
Maintainers usually reply within 1 day
-
good first issue
Difficulty 3/5 Half a day Newbie friendliness 38/100
scverse/fast-array-utils#100 · 4 comments ·
Maintainers usually reply within 1 day
All issues in scverse/fast-array-utils
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 88/100
qgis/QGIS-Plugins-Website#459 ·
-
bug severity:medium
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 2 days
-
bot-found bug priority: P3
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
madenvel/KalinkaPlayer#179 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
ls1intum/edutelligence#1098 ·
Maintainers usually reply within 1 day