Question: Simulating gene knockout on novel datasets (zero-shot inference) & Implementation feedback
還沒有人認領這個 Issue。
評估
- 難度
- 5/5
- 預估耗時
- 一週以上
- 新手友好度
- 30/100
- Issue 類型
- 功能
- 描述清晰度
- 需要釐清
- 活躍度
- 活躍
- 技術堆疊
- python
研究方向
先追蹤 dataset_core.py、sampler.py 和 diffusion_core.py 中與推論相關的假設,接著檢查 checkpoint 載入以及所參照的 embedder 對映。使用 finetuned_tahoe100m_fixed.ckpt 和新的 .h5ad 資料集重現回報的失敗。完成的要求是由維護者決定是否支援 zero-shot;如果支援,則需要定義推論進入點或 predict.py 工作流程。
由索引模型根據 Issue 內容生成。
描述
Dear Authors,
First of all, thank you for your outstanding work on PerturbDiff! The concept and the methodology are truly inspiring.
I am currently trying to apply your pre-trained model (specifically the finetuned_tahoe100m_fixed.ckpt) to a novel, independent single-cell dataset (Colorectal Cancer data). My goal is to perform pure inference: simulating the knockout of a specific gene (e.g., TP53) on this unseen dataset, without having any actual ground-truth perturbation data or paired control cells.
My primary question is: Does the current theoretical framework and the pre-trained checkpoint support this kind of "zero-shot" simulation on a completely novel dataset?
While attempting to implement this inference pipeline, I ran into several engineering challenges. It seems that the current codebase is heavily optimized for training and reproducing benchmark metrics, which makes it quite difficult to decouple for pure, out-of-distribution inference. I would love to share a brief summary of the roadblocks I encountered, in hopes it might be helpful for future updates or an inference-only API:
Tight Coupling in DataLoader & Sampler: The dataset_core.py and sampler.py strictly expect paired "control" and "perturbed" cells to calculate metrics. Bypassing this to feed a simple .h5ad file of raw cells requires heavily modifying the dictionary mappings (e.g., grouped_num_cell, data_indices) to prevent KeyErrors and AssertionErrors.
Hardcoded Dataset Names & Metadata: The codebase heavily relies on predefined dataset names (pbmc, tahoe100m, etc.). When feeding novel data, the framework automatically assigns names like dummy_plate_9, which later causes AssertionErrors in functions like get_short_dsname and embedder mapping.
Strict Checkpoint Loading & Embedder Dimensions: When adapting the model to accept my dataset's dimensions (e.g., 2000 HVGs), modifying the nn.Linear layers causes Unexpected key(s) in load_state_dict because the checkpoint contains hardcoded dataset-specific embedders (e.g., x_embedder.pbmc.weight). This required setting strict=False to force initialization.
Shape Assertions in Diffusion Core: During the forward pass, the Transformer blocks often output a 3D tensor [Batch, 1, Dim], but the diffusion_core.py strictly asserts x_t.shape == eps.shape (expecting [Batch, Dim]). This required manual squeeze/reshape operations at the model's output to prevent runtime crashes.
I wanted to ask if you have any plans to release a simplified predict.py script for users who just want to input an .h5ad and a perturbation condition to get the simulated results.
Thank you again for your time, your amazing research, and for making this repository open-source!
- 主要語言
- Python
- 星號
- 64
- 分支
- 10
- PR 合併指標
- 30 天內沒有已合併 PR
環境準備
這個專案沒有提供開發容器、Dockerfile 或貢獻指南,環境需要你自己搭建:先看它的 README,通用步驟見我們的新手貢獻指南。
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
DeepGraphLearning/PerturbDiff 的其他 Issue
-
難度 4/5 3-5 天 新手友好度 45/100
-
難度 4/5 3-5 天 新手友好度 38/100
DeepGraphLearning/PerturbDiff#7 · 1 則留言 ·
-
RuntimeError: mat1 and mat2 shapes cannot be multiplied during sampling with finetuned_replogle.ckpt未關閉
難度 3/5 1-2 天 新手友好度 68/100
DeepGraphLearning/PerturbDiff#4 · 3 則留言 ·
-
Inference未關閉
難度 4/5 3-5 天 新手友好度 45/100
DeepGraphLearning/PerturbDiff#1 · 4 則留言 ·
查看 DeepGraphLearning/PerturbDiff 的全部 Issue
相似的 Issue
-
難度 2/5 1-3 小時 新手友好度 85/100
mozilla/bedrock#17413 · 1 個 reaction ·
維護者通常 2 天內回覆
-
instance instance add
難度 2/5 1-3 小時 新手友好度 68/100
searxng/searx-instances#943 · 1 則留言 ·
-
難度 2/5 1-3 小時 新手友好度 68/100
維護者通常 1 天內回覆
-
bug tools
難度 2/5 1-3 小時 新手友好度 88/100
維護者通常 1 天內回覆
-
bug
難度 2/5 1-3 小時 新手友好度 86/100
lance-format/lance#9655 ·
維護者通常 2 天內回覆