Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Modkit pileup - segmentation fault and performance issues in high-coverage regions

オープン
#607 コメント 3 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
42/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
静か
技術スタック
rust

調査の方向性

Start by reproducing the two modkit pileup failures with the supplied commands, including --modified-bases inosine, the high-coverage region, and the comma-separated --region value. Use the reported output to investigate the segmentation fault, low CPU utilization, and contig-missing error. Done means each case has a documented resolution or confirmed workaround, with performance behavior checked on the high-coverage region.

索引モデルが issue の本文から書いたものです。

説明

troubleshooting

Dear ONT staff,

I am working with modkit to profile m6A and inosine in a dRNA-seq sample, and I am encountering a couple of issues.

First, when I try to use the --modified-bases parameter, I encounter a segmentation fault error, with no additional information regarding the cause of the issue.

This is the command I am using:

~/Software/dist_modkit_v0.6.1_481e3c9/modkit pileup
/path/to/aligned/bam
/path/to/output/bed.gz
--reference /path/to/fa.gz
--log /path/to/log
--threads 10
--modified-bases inosine
--region 1

This is the output:

parsing region 1
discarded 0 contigs with zero aligned reads
parsed 1 base modification(s). Base modifications other than 'A:17596' will be counted as 'N_other'.
adding single-base motif: 'A 0'
Segmentation fault (core dumped)

Furthermore, I have a couple of genomic regions covered by ~1M reads (due to high expression of specific transcripts), and this results in the code getting stuck, with only 2 out of 10 CPUs (according to top) being actively used.

This is the output:

parsing region 2
discarded 0 contigs with zero aligned reads
attempting to sample 10042 reads
Threshold of 0.69921875 for base A is low. Consider increasing the filter-percentile or specifying a higher threshold.
using general workers
93037146 B written to output: Test.bed.gz [4.78 MB/s]
[00:00:18] ###############------------------------- 88000000/242193529 genome positions 4,732,930.0021/s 33s
1222798 rows written
0 ~records errored

Could you recommend a workaround to avoid this issue and/or improve performance so that all available CPUs are effectively utilized? I also tried with the --high-depth --max-depth 100 options but it did not help.

Finally, I attempted to skip the two chromosomes with very high coverage. However, when I pass a comma-separated list of chromosome names to the --region option (as suggested in the documentation), I encounter a “contig missing” error.

Thank you very much for your support.

Best regards,

Mattia

主要言語
Rust
スター
276
フォーク
33
PR マージ指標
30日以内にマージされた PR はありません

環境構築

このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

nanoporetech/modkit のほかの issue

nanoporetech/modkit の issue をすべて見る

似ている issue

Rust の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。