Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Clarification on high Nfail counts in modkit pileup output despite valid coverage

Đang mở
#577 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
30/100
Loại issue
Tài liệu
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Đình trệ
Công nghệ
rust
Lĩnh vực
bioinformatics, cli

Hướng nghiên cứu

Start with the modkit pileup entry point and the definitions of Nfail, valid_coverage, count_modified, and count_canonical in the documentation. Reproduce the reported command with the stated filters, then document which filters produce Nfail and how min-mod-prob affects interpretation; done means the questions have a clear, supported explanation.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

question

Hi modkit team,

I am analyzing CpG methylation using modkit pileup on Oxford Nanopore data
(mapped BAMs, reference-guided).

Across multiple samples, I observe that the majority of reads contributing to
coverage end up in the Nfail column, even when valid_coverage ≥ 3.

Some details:

  • Genome: Plasmodium falciparum 3D7

  • Command used (example):

    modkit pileup
    --reference PlasmoDB-64_Pfalciparum3D7_Genome.fasta
    --modified-bases C
    --combine-mods
    input.bam output.bed

  • In the resulting bedMethyl files:

    • valid_coverage is often high (≥3 for many CpGs)
    • count_modified and count_canonical are low
    • most reads appear to be classified as Nfail
    • fraction of methylated CpGs among covered sites is very low

This behavior is consistent across samples and across different coverage cutoffs.

My questions:

  1. Is a high Nfail count expected behavior in cases of low-confidence or low-level CpG methylation?
  2. Which filters most commonly cause reads to be classified as Nfail?
    (e.g. modification probability threshold, basecall quality, alignment flags, context mismatch)
  3. Is there a recommended way to summarize or interpret datasets where
    valid coverage exists but most reads fail confidence filters?
  4. Would adjusting parameters like min-mod-prob be appropriate to explore this further?

The behavior seems biologically plausible for this organism, but I would like
to confirm that my interpretation of Nfail is correct.

Thanks for the great tool, and for any clarification!

Ngôn ngữ chính
Rust
Star
274
Fork
33
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của nanoporetech/modkit

Tất cả issue của nanoporetech/modkit

Issue tương tự

Thêm issue về Rust

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.