Unknown

Dataset Information

FASTAFS: file system virtualisation of random access compressed FASTA files.


ABSTRACT:

Background

The FASTA file format, used to store polymeric sequence data, has become a bioinformatics file standard used for decades. The relatively large files require additional files, beyond the scope of the original format, to identify sequences and to provide random access. Multiple compressors have been developed to archive FASTA files back and forth, but these lack direct access to targeted content or metadata of the archive. Moreover, these solutions are not directly backwards compatible to FASTA files, resulting in limited software integration.

Results

We designed a linux based toolkit that virtualises the content of DNA, RNA and protein FASTA archives into the filesystem by using filesystem in userspace. This guarantees in-sync virtualised metadata files and offers

SUBMITTER: Hoogstrate Y 

PROVIDER: S-EPMC8558547 | biostudies-literature | 2021 Nov

REPOSITORIES: biostudies-literature

altmetric image

Publications

Sorry, this publication's infomation has not been loaded in the Indexer, please go directly to PUBMED or Altmetric.

Similar Datasets