<HashMap><database>biostudies-literature</database><scores/><additional><submitter>Eberhardt RY</submitter><funding>NHGRI NIH HHS</funding><funding>Wellcome Trust</funding><pagination>bas003</pagination><full_dataset_link>https://www.ebi.ac.uk/biostudies/studies/S-EPMC3308159</full_dataset_link><repository>biostudies-literature</repository><omics_type>Unknown</omics_type><volume>2012</volume><pubmed_abstract>As the deluge of genomic DNA sequence grows the fraction of protein sequences that have been manually curated falls. In turn, as the number of laboratories with the ability to sequence genomes in a high-throughput manner grows, the informatics capability of those labs to accurately identify and annotate all genes within a genome may often be lacking. These issues have led to fears about transitive annotation errors making sequence databases less reliable. During the lifetime of the Pfam protein families database a number of protein families have been built, which were later identified as composed solely of spurious open reading frames (ORFs) either on the opposite strand or in a different, overlapping reading frame with respect to the true protein-coding or non-coding RNA gene. These famil</pubmed_abstract><journal>Database : the journal of biological databases and curation</journal><pubmed_title>AntiFam: a tool to help identify spurious ORFs in protein annotation.</pubmed_title><pmcid>PMC3308159</pmcid><funding_grant_id>WT077044/Z/05/Z</funding_grant_id><funding_grant_id>R01 HG004881).</funding_grant_id><pubmed_authors>Haft DH</pubmed_authors><pubmed_authors>O'Donovan C</pubmed_authors><pubmed_authors>Martin M</pubmed_authors><pubmed_authors>Bateman A</pubmed_authors><pubmed_authors>Eberhardt RY</pubmed_authors><pubmed_authors>Punta M</pubmed_authors></additional><is_claimable>false</is_claimable><name>AntiFam: a tool to help identify spurious ORFs in protein annotation.</name><description>As the deluge of genomic DNA sequence grows the fraction of protein sequences that have been manually curated falls. In turn, as the number of laboratories with the ability to sequence genomes in a high-throughput manner grows, the informatics capability of those labs to accurately identify and annotate all genes within a genome may often be lacking. These issues have led to fears about transitive annotation errors making sequence databases less reliable. During the lifetime of the Pfam protein families database a number of protein families have been built, which were later identified as composed solely of spurious open reading frames (ORFs) either on the opposite strand or in a different, overlapping reading frame with respect to the true protein-coding or non-coding RNA gene. These famil</description><dates><release>2012-01-01T00:00:00Z</release><publication>2012</publication><modification>2025-04-19T10:06:33.944Z</modification><creation>2019-03-27T00:51:23Z</creation></dates><accession>S-EPMC3308159</accession><cross_references><pubmed>22434837</pubmed><doi>10.1093/database/bas003</doi></cross_references></HashMap>