Dataset Information


Employing whole genome mapping for optimal de novo assembly of bacterial genomes.

ABSTRACT: BACKGROUND: De novo genome assembly can be challenging due to inherent properties of the reads, even when using current state-of-the-art assembly tools based on de Bruijn graphs. Often users are not bio-informaticians and, in a black box approach, utilise assembly parameters such as contig length and N50 to generate whole genome sequences, potentially resulting in mis-assemblies. FINDINGS: Utilising several assembly tools based on de Bruijn graphs like Velvet, SPAdes and IDBA, we demonstrate that at the optimal N50, mis-assemblies do occur, even when using the multi-k-mer approaches of SPAdes and IDBA. We demonstrate that whole genome mapping can be used to identify these mis-assemblies and can guide the selection of the best k-mer size which yields the highest N50 without mis-assemblies. CONCLUSIONS: We demonstrate the utility of whole genome mapping (WGM) as a tool to identify mis-assemblies and to guide k-mer selection and higher quality de novo genome assembly of bacterial genomes.


PROVIDER: S-EPMC4118782 | BioStudies | 2014-01-01

REPOSITORIES: biostudies

Similar Datasets

2013-01-01 | S-EPMC3791033 | BioStudies
2015-01-01 | S-EPMC4379979 | BioStudies
2014-01-01 | S-EPMC4290589 | BioStudies
2012-01-01 | S-EPMC3489510 | BioStudies
2008-01-01 | S-EPMC2336801 | BioStudies
2014-01-01 | S-EPMC4058956 | BioStudies
2017-01-01 | S-EPMC5870571 | BioStudies
2012-01-01 | S-EPMC3431216 | BioStudies
2016-01-01 | S-EPMC4876485 | BioStudies
1000-01-01 | S-EPMC5084376 | BioStudies