A novel algorithm for computational identification of contaminated EST libraries.
Ontology highlight
ABSTRACT: A key goal of the Human Genome Project was to understand the complete set of human proteins, the proteome. Since the genome sequence by itself is not sufficient for predicting new genes and alternative splicing events that lead to new proteins, expressed sequence tags (ESTs) are used as the primary tool for these purposes. The high prevalence of artifacts in dbEST, however, often leads to invalid predictions. Here we describe a novel method for recognizing genomic DNA contamination and other artifacts that cannot be identified using current EST cleaning techniques. Our method uses the alignment of the entire set of ESTs to the human genome to identify highly contaminated EST libraries. We discovered 53 highly contaminated libraries and a subset of 24 766 ESTs from these libraries that prob
SUBMITTER: Sorek R
PROVIDER: S-EPMC149192 | biostudies-literature | 2003 Feb
REPOSITORIES: biostudies-literature
ACCESS DATA