Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

MCMICRO reader does not support new nf-core/mcmicro output formats

Abierto
#409 0 comentarios 1 reacción 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
48/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Tranquilo
Stack tecnológico
python
Área
data

Línea de trabajo

Comienza en spatialdata_io.readers.mcmicro y compara sus suposiciones con los dos diseños de salida enumerados en el issue. Rastrea cómo se descubren el registro, la segmentación, la cuantificación, los marcadores y los núcleos TMA, y define casos de validación para múltiples muestras y segmentadores, la detección automática de la pipeline y la sobrescritura explícita. Se considera terminado cuando las salidas tanto de labsyspharm como de nf-core siguen siendo legibles, incluidos los casos de WSI y TMA.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

We recently completed a port of MCMICRO to nf-core (https://nf-co.re/mcmicro/2.0.0). There are minor changes to expected output formats, and the current spatialdata_io.readers.mcmicro reader — written against the original labsyspharm/mcmicro layout — does not read nf-core output.

What changed between the two pipelines:

Aspect labsyspharm/mcmicro nf-core/mcmicro
Run config qc/params.yml (workflow.tma) pipeline_info/params*.json (tma_dearray, segmentation)
Registration registration/<sample>.ome.tif registration/ashlar/<sample>.ome.tif
Segmentation segmentation/<module>-<sample>/{cell,nuclei}.ome.tif segmentation/<segmenter>/<sample>_<tool-suffix>.tif
Quantification quantification/<module>-<sample>_<comp>.csv quantification/mcquant/<segmenter>/<sample>.csv
TMA cores dearray/ + qc/coreograph/centroidsY-X.txt tma_dearray/ (plus TMA_MAP.tif) + tma_dearray/centroidsY-X.txt
Markers markers.csv at root input sheet (often not copied to outdir); backsub/<sample>_backsub.csv when backsub runs
Samples / segmenters single sample, typically one segmenter multiple samples and multiple segmenters per run

Gotchas found while testing against real nf-core/mcmicro output that will need to be considered:

  • Segmentation mask filenames are tool-specific, not a uniform _mask suffix — e.g. 1_cp_masks.tif (cellpose) vs exemplar-002_1_mask.tif (mesmer).
  • The segmenter directory name differs between segmentation/ and quantification/mcquant/ (e.g. deepcell_mesmer vs mesmer), so tables can't be linked to labels by an exact directory-name match.
  • With background subtraction, the registration image and the quantification tables carry different marker sets (e.g. 40 image channels vs 37 table markers).
  • TMA core numbers must be parsed carefully — the sample name itself can contain digits (exemplar-002_1_mask → core 1, not 002) — and tma_dearray/ also contains TMA_MAP.tif, which is not a core.
  • The prelude/markers_markersheet_mqc.tsv file is a long-format MultiQC validation report, not a usable wide marker sheet.

Proposal

Update the mcmicro reader to auto-detect the pipeline (with an explicit pipeline= override) and handle both layouts, including multiple samples and segmenters, in WSI and TMA modes. Backwards compatibility with labsyspharm output is preserved.

Lenguaje dominante
Python
Estrellas
103
Forks
65
Merge medio
1 h 8 min
PR fusionados (30 d)
3

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de scverse/spatialdata-io

Todos los issues de scverse/spatialdata-io

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.