NIH R01 · 2025
Automated methods for standardization and enhancement of metadata in biomedical databases
PROJECT SUMMARY/ABSTRACT Biomedical research data sets are increasingly being deposited in public, centralized databases, such as the Sequence Read Archive (SRA), to which researchers submit sequencing-based data. Large centralized databases greatly enable opportunities for training powerful machine learning models, as well as for reanalysis and cross-study meta-analysis of biomedical data. These analyses can be used to answer questions that were not addressed in the papers first describing the data, including those that could only be answered by aggregating data from multiple studies. Unfortunately, researchers have not been able to fully capitalize on databases of biomedical data sets…
From the public funding record at NIH RePORTER. Describes the funded project, not the reviews below.