Improving the utility of the Tox21 dataset by deep metadata annotations and constructing reusable benchmarked chemical reference signatures

Daniel J. Cooper, Stephan C Schuerer

Research output: Contribution to journalArticle

Abstract

The Toxicology in the 21st Century (Tox21) project seeks to develop and test methods for high-throughput examination of the effect certain chemical compounds have on biological systems. Although primary and toxicity assay data were readily available for multiple reporter gene modified cell lines, extensive annotation and curation was required to improve these datasets with respect to how FAIR (Findable, Accessible, Interoperable, and Reusable) they are. In this study, we fully annotated the Tox21 published data with relevant and accepted controlled vocabularies. After removing unreliable data points, we aggregated the results and created three sets of signatures reflecting activity in the reporter gene assays, cytotoxicity, and selective reporter gene activity, respectively. We benchmarked these signatures using the chemical structures of the tested compounds and obtained generally high receiver operating characteristic (ROC) scores, suggesting good quality and utility of these signatures and the underlying data. We analyzed the results to identify promiscuous individual compounds and chemotypes for the three signature categories and interpreted the results to illustrate the utility and re-usability of the datasets. With this study, we aimed to demonstrate the importance of data standards in reporting screening results and high-quality annotations to enable re-use and interpretation of these data. To improve the data with respect to all FAIR criteria, all assay annotations, cleaned and aggregate datasets, and signatures were made available as standardized dataset packages (Aggregated Tox21 bioactivity data, 2019).

Original languageEnglish (US)
Article number1604
JournalMolecules
DOIs
StatePublished - Apr 23 2019

Fingerprint

toxicology
annotations
metadata
Metadata
Toxicology
Assays
Genes
Reporter Genes
signatures
genes
Thesauri
Chemical compounds
Reusability
Controlled Vocabulary
Biological systems
Cytotoxicity
Bioactivity
Toxicity
Screening
chemical compounds

Keywords

  • Benchmarking
  • Data standards
  • FAIR data
  • High-throughput screening
  • Metadata
  • Ontologies
  • Signatures
  • Tox21

ASJC Scopus subject areas

  • Analytical Chemistry
  • Chemistry (miscellaneous)
  • Molecular Medicine
  • Pharmaceutical Science
  • Drug Discovery
  • Physical and Theoretical Chemistry
  • Organic Chemistry

Cite this

@article{aa820ae97708466697ec271c21f9482c,
title = "Improving the utility of the Tox21 dataset by deep metadata annotations and constructing reusable benchmarked chemical reference signatures",
abstract = "The Toxicology in the 21st Century (Tox21) project seeks to develop and test methods for high-throughput examination of the effect certain chemical compounds have on biological systems. Although primary and toxicity assay data were readily available for multiple reporter gene modified cell lines, extensive annotation and curation was required to improve these datasets with respect to how FAIR (Findable, Accessible, Interoperable, and Reusable) they are. In this study, we fully annotated the Tox21 published data with relevant and accepted controlled vocabularies. After removing unreliable data points, we aggregated the results and created three sets of signatures reflecting activity in the reporter gene assays, cytotoxicity, and selective reporter gene activity, respectively. We benchmarked these signatures using the chemical structures of the tested compounds and obtained generally high receiver operating characteristic (ROC) scores, suggesting good quality and utility of these signatures and the underlying data. We analyzed the results to identify promiscuous individual compounds and chemotypes for the three signature categories and interpreted the results to illustrate the utility and re-usability of the datasets. With this study, we aimed to demonstrate the importance of data standards in reporting screening results and high-quality annotations to enable re-use and interpretation of these data. To improve the data with respect to all FAIR criteria, all assay annotations, cleaned and aggregate datasets, and signatures were made available as standardized dataset packages (Aggregated Tox21 bioactivity data, 2019).",
keywords = "Benchmarking, Data standards, FAIR data, High-throughput screening, Metadata, Ontologies, Signatures, Tox21",
author = "Cooper, {Daniel J.} and Schuerer, {Stephan C}",
year = "2019",
month = "4",
day = "23",
doi = "10.3390/molecules24081604",
language = "English (US)",
journal = "Molecules",
issn = "1420-3049",
publisher = "Multidisciplinary Digital Publishing Institute (MDPI)",

}

TY - JOUR

T1 - Improving the utility of the Tox21 dataset by deep metadata annotations and constructing reusable benchmarked chemical reference signatures

AU - Cooper, Daniel J.

AU - Schuerer, Stephan C

PY - 2019/4/23

Y1 - 2019/4/23

N2 - The Toxicology in the 21st Century (Tox21) project seeks to develop and test methods for high-throughput examination of the effect certain chemical compounds have on biological systems. Although primary and toxicity assay data were readily available for multiple reporter gene modified cell lines, extensive annotation and curation was required to improve these datasets with respect to how FAIR (Findable, Accessible, Interoperable, and Reusable) they are. In this study, we fully annotated the Tox21 published data with relevant and accepted controlled vocabularies. After removing unreliable data points, we aggregated the results and created three sets of signatures reflecting activity in the reporter gene assays, cytotoxicity, and selective reporter gene activity, respectively. We benchmarked these signatures using the chemical structures of the tested compounds and obtained generally high receiver operating characteristic (ROC) scores, suggesting good quality and utility of these signatures and the underlying data. We analyzed the results to identify promiscuous individual compounds and chemotypes for the three signature categories and interpreted the results to illustrate the utility and re-usability of the datasets. With this study, we aimed to demonstrate the importance of data standards in reporting screening results and high-quality annotations to enable re-use and interpretation of these data. To improve the data with respect to all FAIR criteria, all assay annotations, cleaned and aggregate datasets, and signatures were made available as standardized dataset packages (Aggregated Tox21 bioactivity data, 2019).

AB - The Toxicology in the 21st Century (Tox21) project seeks to develop and test methods for high-throughput examination of the effect certain chemical compounds have on biological systems. Although primary and toxicity assay data were readily available for multiple reporter gene modified cell lines, extensive annotation and curation was required to improve these datasets with respect to how FAIR (Findable, Accessible, Interoperable, and Reusable) they are. In this study, we fully annotated the Tox21 published data with relevant and accepted controlled vocabularies. After removing unreliable data points, we aggregated the results and created three sets of signatures reflecting activity in the reporter gene assays, cytotoxicity, and selective reporter gene activity, respectively. We benchmarked these signatures using the chemical structures of the tested compounds and obtained generally high receiver operating characteristic (ROC) scores, suggesting good quality and utility of these signatures and the underlying data. We analyzed the results to identify promiscuous individual compounds and chemotypes for the three signature categories and interpreted the results to illustrate the utility and re-usability of the datasets. With this study, we aimed to demonstrate the importance of data standards in reporting screening results and high-quality annotations to enable re-use and interpretation of these data. To improve the data with respect to all FAIR criteria, all assay annotations, cleaned and aggregate datasets, and signatures were made available as standardized dataset packages (Aggregated Tox21 bioactivity data, 2019).

KW - Benchmarking

KW - Data standards

KW - FAIR data

KW - High-throughput screening

KW - Metadata

KW - Ontologies

KW - Signatures

KW - Tox21

UR - http://www.scopus.com/inward/record.url?scp=85064829518&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=85064829518&partnerID=8YFLogxK

U2 - 10.3390/molecules24081604

DO - 10.3390/molecules24081604

M3 - Article

C2 - 31018579

AN - SCOPUS:85064829518

JO - Molecules

JF - Molecules

SN - 1420-3049

M1 - 1604

ER -