Unknown

Dataset Information

Biobambam: tools for read pair collation based algorithms on BAM files


ABSTRACT:

Background

Sequence alignment data is often ordered by coordinate (id of the reference sequence plus position on the sequence where the fragment was mapped) when stored in BAM files, as this simplifies the extraction of variants between the mapped data and the reference or of variants within the mapped data. In this order paired reads are usually separated in the file, which complicates some other applications like duplicate marking or conversion to the FastQ format which require to access the full information of the pairs.

Results

In this paper we introduce biobambam, a set of tools based on the efficient collation of alignments in BAM files by read name. The employed collation algorithm avoids time and space consuming sorting of alignments by read name where this is pos

SUBMITTER: Tischler G 

PROVIDER: S-EPMC4075596 | biostudies-literature | 2014 Jan

REPOSITORIES: biostudies-literature

Similar Datasets