Predicting biomedical metadata in CEDAR: A study of Gene Expression Omnibus (GEO).
Ontology highlight
ABSTRACT: A crucial and limiting factor in data reuse is the lack of accurate, structured, and complete descriptions of data, known as metadata. Towards improving the quantity and quality of metadata, we propose a novel metadata prediction framework to learn associations from existing metadata that can be used to predict metadata values. We evaluate our framework in the context of experimental metadata from the Gene Expression Omnibus (GEO). We applied four rule mining algorithms to the most common structured metadata elements (sample type, molecular type, platform, label type and organism) from over 1.3million GEO records. We examined the quality of well supported rules from each algorithm and visualized the dependencies among metadata elements. Finally, we evaluated the performance of the algorith
SUBMITTER: Panahiazar M
PROVIDER: S-EPMC5643580 | biostudies-literature | 2017 Aug
REPOSITORIES: biostudies-literature
ACCESS DATA