<HashMap><database>biostudies-literature</database><scores/><additional><submitter>Allesoe RL</submitter><funding>Marie Sklodowska-Curie Individual Fellowship</funding><funding>Novo Nordisk Foundation</funding><funding>Medical Research Council</funding><funding>Novo Nordisk Foundation Center for Protein Research</funding><funding>European Union’s Horizon 2020 research and innovation programme</funding><funding>European Union’s Horizon 2020 research and innovation program</funding><pagination>705-710</pagination><full_dataset_link>https://www.ebi.ac.uk/biostudies/studies/S-EPMC8097684</full_dataset_link><repository>biostudies-literature</repository><omics_type>Unknown</omics_type><volume>37(5)</volume><pubmed_abstract>&lt;h4>Summary&lt;/h4>Here, we present an automated pipeline for Download Of NCBI Entries (DONE) and continuous updating of a local sequence database based on user-specified queries. The database can be created with either protein or nucleotide sequences containing all entries or complete genomes only. The pipeline can automatically clean the database by removing entries with matches to a database of user-specified sequence contaminants. The default contamination entries include sequences from the UniVec database of plasmids, marker genes and sequencing adapters from NCBI, an E.coli genome, rRNA sequences, vectors and satellite sequences. Furthermore, duplicates are removed and the database is automatically screened for sequences from green fluorescent protein, luciferase and antibiotic resistan</pubmed_abstract><journal>Bioinformatics (Oxford, England)</journal><pubmed_title>Automated download and clean-up of family-specific databases for kmer-based virus identification.</pubmed_title><pmcid>PMC8097684</pmcid><funding_grant_id>NNF14CC0001</funding_grant_id><funding_grant_id>643476</funding_grant_id><funding_grant_id>PI Simon Rasmussen</funding_grant_id><funding_grant_id>874735</funding_grant_id><funding_grant_id>799417</funding_grant_id><funding_grant_id>MC_UU_12014/12</funding_grant_id><pubmed_authors>Koopmans MPG</pubmed_authors><pubmed_authors>Clausen PTLC</pubmed_authors><pubmed_authors>Lemvigh CK</pubmed_authors><pubmed_authors>Cotten M</pubmed_authors><pubmed_authors>Allesoe RL</pubmed_authors><pubmed_authors>Florensa AF</pubmed_authors><pubmed_authors>Phan MVT</pubmed_authors><pubmed_authors>Lund O</pubmed_authors></additional><is_claimable>false</is_claimable><name>Automated download and clean-up of family-specific databases for kmer-based virus identification.</name><description>&lt;h4>Summary&lt;/h4>Here, we present an automated pipeline for Download Of NCBI Entries (DONE) and continuous updating of a local sequence database based on user-specified queries. The database can be created with either protein or nucleotide sequences containing all entries or complete genomes only. The pipeline can automatically clean the database by removing entries with matches to a database of user-specified sequence contaminants. The default contamination entries include sequences from the UniVec database of plasmids, marker genes and sequencing adapters from NCBI, an E.coli genome, rRNA sequences, vectors and satellite sequences. Furthermore, duplicates are removed and the database is automatically screened for sequences from green fluorescent protein, luciferase and antibiotic resistan</description><dates><release>2021-01-01T00:00:00Z</release><publication>2021 May</publication><modification>2026-05-08T02:59:17.628Z</modification><creation>2022-02-10T09:57:18.722Z</creation></dates><accession>S-EPMC8097684</accession><cross_references><pubmed>33031509</pubmed><doi>10.1093/bioinformatics/btaa857</doi></cross_references></HashMap>