Deep metadata annotations and FAIR enhancement of DrugMatrix in vitro bioassay data
摘要
DrugMatrix is a comprehensive molecular toxicology database integrating in vivo gene-expression profiles with conventional toxicology endpoints. It represents one of the largest toxicogenomic resources, encompassing data from over 600 benchmarked chemicals across eight rat tissues together with in vitro gene-expression data from primary hepatocyte cultures. In addition to its rat-based data, DrugMatrix includes in vitro pharmacology data that provide complementary mechanistic insight into compound activity. However, the in vitro pharmacology dataset distributed through ChEMBL was partially incomplete, with missing assay records and experimental details that limited reuse and integration. To address this, we conducted comprehensive curation of these data to improve accessibility and alignment with FAIR (Findable, Accessible, Interoperable, Reusable) principles. Metadata were expanded and previously unverified (“unchecked”) entries were resolved to enhance dataset completeness and consistency. The resulting DrugMatrix in vitro dataset provides a standardized and interoperable resource that supports robust data integration and reuse. The final updated DrugMatrix in vitro dataset is archived on Zenodo and will be integrated in ChEMBL as part of the ChEMBL 38 release.