Shared sparse one-hot (indicator) helper for squidpy and scanpy
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 72/100
Hướng nghiên cứu
Start with scanpy/get/_aggregated.py and squidpy/gr/_nhood.py, then review the proposed fast_array_utils.conv entry point and its sparse extra. Confirm the shared helper preserves missing labels, unused categories, mask handling, float64 output, and the stated matrix layouts; done means both callers use it without changing their current behavior.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
prettified with ai, basically saying we have two different functions we can unify here. low prio but good to document.
scanpy (sparse_indicator in scanpy/get/_aggregated.py) and squidpy (_onehot in squidpy/gr/_nhood.py) each keep a private helper that turns category codes into a sparse float64 indicator matrix, with a missing label (-1) giving no entry. They differ only in interface: scanpy takes a pd.Categorical and returns (n_categories, n_obs) as a coo_array with an optional mask, while squidpy takes a pd.Series and returns a (n_obs, n_categories) csr_matrix.
Proposal for fast_array_utils.conv (needs the sparse extra; takes codes since fast-array-utils doesn't depend on pandas):
def sparse_indicator(
codes: NDArray[np.integer], n_categories: int, *, mask: NDArray[np.bool] | None = None
) -> coo_array:
keep = codes >= 0 if mask is None else (codes >= 0) & mask
obs = np.flatnonzero(keep)
return coo_array((np.ones(obs.size), (obs, codes[keep])), shape=(codes.size, n_categories))
(n_obs, n_categories), the usual one-hot layout; scanpy takes.T(about 3 ms at 5M observations).- float64, as both callers use today.
- Same cost as both copies today: squidpy converts the result to CSR, as it does now.
checked with missing labels, unused categories and mask.
as discussed in scverse/squidpy#1285
- Ngôn ngữ chính
- Python
- Star
- 15
- Fork
- 5
- Merge trung bình
- 14 giờ 52 phút
- Pull request đã merge (30 ngày)
- 13
Chuẩn bị môi trường
Chúng tôi chưa kiểm tra các tệp thiết lập môi trường của dự án này. Hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của scverse/fast-array-utils
-
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 74/100
scverse/fast-array-utils#165 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 65/100
scverse/fast-array-utils#213 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
type: numpy/scipy
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
scverse/fast-array-utils#128 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
good first issue
Độ khó 3/5 Nửa ngày Mức phù hợp với người mới 38/100
scverse/fast-array-utils#100 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
component: documentation
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 20/100
scverse/fast-array-utils#89 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của scverse/fast-array-utils
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
PedestrianDynamics/pyFDS-Evac#343 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
theskumar/python-dotenv#708 ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 2 ngày
-
Docs Timedelta
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
pandas-dev/pandas#69919 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
API documentation
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
zephyrproject-rtos/west#1009 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày