tseemann/snippy

real reads versus simulated reads

开放

#398 创建于 2020年6月6日

 (2 条评论) (0 个反应) (1 位负责人)Perl (123 个派生)github user discovery
help wanted

仓库指标

星标
 (586 个星标)
PR 合并指标
 (PR 指标待抓取)

描述

Hi Torsten,

Lots of bacterial genomes are lacking SRA data, preventing us from performing several reads-based analysis. I think you mentioned somewhere that "best reads are contigs", and do you reckon simulated reads generated from the assembly (perfect reads, no error, perfectly even coverage) should be used in Snippy even if we do have the real reads.

I've run some tests to see if they are significantly different (simulated reads were generated from Shovill-assembled genome with the same read length and coverage as that of the real reads) and the answer is "yes". I found more variants including SNP were detected by Snippy using simulated reads compared with using the real reads. So it makes me wondering which one is closer to the truth.

Thanks

Yu

贡献者指南