tseemann/snippy

real reads versus simulated reads

Aperta

#398 aperta il 6 giu 2020

 (2 commenti) (0 reazioni) (1 assegnatario)Perl (123 fork)github user discovery
help wanted

Metriche repository

Star
 (586 stelle)
Metriche merge PR
 (Metriche PR in attesa)

Descrizione

Hi Torsten,

Lots of bacterial genomes are lacking SRA data, preventing us from performing several reads-based analysis. I think you mentioned somewhere that "best reads are contigs", and do you reckon simulated reads generated from the assembly (perfect reads, no error, perfectly even coverage) should be used in Snippy even if we do have the real reads.

I've run some tests to see if they are significantly different (simulated reads were generated from Shovill-assembled genome with the same read length and coverage as that of the real reads) and the answer is "yes". I found more variants including SNP were detected by Snippy using simulated reads compared with using the real reads. So it makes me wondering which one is closer to the truth.

Thanks

Yu

Guida contributor