The Role of Metadata in Reproducible Computational Research

July 7, 2026 · View on GitHub

This is a supplemental resource to Leipzig et al. "The Role of Metadata in Reproducible Computational Research" now published in Cell Patterns https://www.cell.com/patterns/fulltext/S2666-3899(21)00170-7

Contributions are welcome!

Organization

├───data/
│   ├───examples/                  Examples of metadata standards
│   ├───lens/                      Search exports for scimetric journal analysis
│   └───standards.tsv              Raw standards table
├───src/
│   ├───cwl/tools/                 CWL configuration to produce the timeline plot
│   ├───manuscript/                Manuscript revision document
│   ├───secrets/
│   │   └───api.template.py        Replace this with api.py using your NCBI/NCBO keys
│   ├───ontologies/                Scimetric ontology popularity analysis
│   ├───repotutils/                Scripts for automating management of this repository
│   ├───scimetric/                 Scimetric journal meta/rcr frequency analysis in a Jupyter Notebook
│   ├───timeline/                  R Markdown document to produce the RCR case study timeline in the paper, incl. helper files for execution with CWL (wrapper script, Dockerfile)
│   ├───wget2jsonld.py             Helper script to convert wget output to jsonld
│   └───wordcloud/                 R script to produce word cloud from cited abstracts
├───LICENSE                        The LICENSE file
├───README.md                      What you are looking at
├───environment.osx.yaml           OSX pinned Conda depenencies
├───environment.unpinned.yaml      Unpinned Conda depenencies
└───ro-crate-metadata.jsonld       RO Crate config
└───.binder                        Environment configuration files for usage with Binder (mybinder.org)

Examples of RCR metadata standards

In this table we provide links to the authoritative publications and homepages for these metadata standards, as well as examples we have collected. Schema refers the parent structure this standard conforms to, if any. Encoding refers to the markup format used. Note that for schemas such as OWL, which relies on RDF subject–predicate–object triplets, the encoding could be one of at least seven serialization types (RDF/XML, RDF/JSON, JSON-LD, Turtle, N-Triples, N-Quads, N3), so the listed encoding is somewhat arbitrary. For other standards, such as DICOM, the encoding is a custom binary although there are numerous export format and even attempts to serialize JSON within DICOM.

[📚] Publication [🏠] Homepage [📋] Example

StandardLayerDomainEncodingSchemaDescription
CellML 📚 🏠 📋InputBiologyXMLRDFmathematical models for biology
CIF2 📚 🏠 InputCrystallographyCustomatomic structure
DATS 📚 🏠 InputBiomedicalJSONdesc metadata (people, org, repo) for data pubs
DICOM 📚 🏠 📋InputImagesCustomKey-Valuestandard for all medical imaging
EML 📚 🏠 InputEcologyXMLeco support for geo, species, pubs used in KNB
FAANG  🏠 InputSpecimensTabularsample and experiment metadata for farm animal genomes
GBIF 📚 🏠 InputBiodiversityJSONspecies occurrence and biodiversity records
GO 📚 🏠 InputGenesXMLcontrolled vocabulary of gene and gene product functions
ISO/TC 276  🏠 InputBiotechnologyISO committee defining biotechnology data standards
MIAME 📚 🏠 InputMicroarraysXMLminimum information to interpret a microarray experiment
NetCDF 📚 🏠 InputArraysself-describing, array-oriented scientific data
OGC  🏠 InputGeospatialopen standards for geospatial content and services
ThermoML 📚 🏠 InputCompoundsXMLthermodynamic and thermophysical property data
CRAN  🏠 ToolsR packagesR package DESCRIPTION metadata (dependencies, authors)
Conda  🏠 ToolsDependenciespackage and environment dependency specifications
pip setup.cfg  🏠 ToolsPython modulesCFGKey-ValuePython cfg files have headers and key-value pairs similar to Windows INI files
EDAM 📚 🏠 ToolsBfx dataontology of bioinformatics operations, data, formats, and topics
CodeMeta  🏠 ToolsSource codeschema.org-based crosswalk for software metadata
Biotoolsxsd 📚 🏠 ToolsBfx softwareXMLXML schema behind the bio.tools software registry
DOAP  🏠 ToolsSoftwareXMLDescription of a Project vocabulary for software projects
ontosoft  🏠 ToolsGeo softwareontology for describing and sharing scientific software
SWO 📚 🏠 ToolsBfx Softwareontology of software, its inputs, outputs, and tasks
OBCS 📚 🏠 ReportsBiostatisticsontology of biological and clinical statistics
STATO  🏠 ReportsStatisticsontology of statistical methods and tests
SDMX  🏠 ReportsStatisticsJSONstandard for exchanging statistical data and metadata
DDI  🏠 ReportsStudiesXMLdocumentation for social, behavioral, and economic data
MEX 📚 🏠 ReportsMLXMLlightweight vocabulary for interchanging machine learning experiments
MLSchema  🏠 ReportsMLschema for describing machine learning algorithms and experiments
MLFlow  🏠 ReportsMLtracking of ML experiments, runs, parameters, and models
Rmd  🏠 ReportsDocsYAMLKey-ValueYAML front matter for reproducible R Markdown reports
CWL 📚 🏠 Tools, PipelinesYAMLSchema SaladCommon Workflow Language specifies how to invoke a command line tool or a pipeline of such tools
CWLProv 📚 🏠 PipelinesYAML, JSON, XMLBagIt of Research Object folder containing manifest (JSON-LD), CWL (YAML), PROV (JSON, XML, RDF)
RO-Crate  🏠 Input, Pipelines, PublicationJSON-LDRDF, schema.orgRO-Crate is a profile of using schema.org to annotate any collections of research data and their real-life origins
RO  🏠 PipelinesTurtle, JSON-LD, XMLOWLbundles data, methods, and provenance into a single research object
WICUS  🏠 Pipelinesontology for conserving the computational infrastructure of workflows
OPM  🏠 Pipelinesmodel for representing provenance as a directed graph
PROV-O  🏠 PipelinesOWLSeveral PROV serializations exists; PROV-O is in OWL, which again has many serializations including the RDF syntaxes
ReproZip  🏠 PipelinesThe config for ReproZip is YAML. The actual recorded sessions are more like coredumps.
ProvOne  🏠 Pipelinesextends PROV to capture scientific workflow provenance
WES   PipelinesGA4GH API for submitting and monitoring workflow runs
BagIt  🏠 Input, PipelinesTextKey-ValueFor long-term perservation and availability BagIt specifies a fixed folder structure of payload files, their checksums and other metadata tag files. Bags can be archived as zip, tar, etc or remain folders
BCO   PipelinesBioCompute Object records regulatory bioinformatics pipelines
ERC 📚 🏠 PipelinesResearch CompendiaYAMLKey-ValueExecutable Research Compendium packaging code, data, and environment
BEL   PublicationBiological Expression Language for causal biological relationships
DC   PublicationDublin Core general-purpose metadata element set
JATS  🏠 PublicationArticlesXMLTags DTDXML tag set for marking up scholarly journal articles
ONIX   Publicationmetadata standard for the book and publishing supply chain
MeSH   Publicationcontrolled vocabulary for indexing biomedical literature
LCSH   PublicationLibrary of Congress subject headings for bibliographic records
MP 📚  PublicationMicropublicationsOWLmicropublications model linking scientific claims to evidence
Open PHACTS 📚 🏠 PublicationDrugsRDFsemantic interoperability across drug discovery datasets
SWAN 📚  PublicationNeuromedicineontology for scientific discourse in neuromedicine
SPAR  🏠 PublicationPublishingOWLsuite of ontologies for semantic publishing and referencing
PWO 📚  PublicationPublishingontology describing the steps of a publishing workflow
PAV 📚  PublicationAuthorshipOWLprovenance, authoring, and versioning ontology
Manubot   📋PublicationPublishingYAMLmetadata for collaborative, automated manuscript writing
ReScience   📋PublicationPublishingYAMLmetadata for peer-reviewed replications of computational research
PandocScholar   📋PublicationPublishingYAMLscholarly article metadata authored through Pandoc

RDF vs OWL https://stackoverflow.com/questions/1740341/what-is-the-difference-between-rdf-and-owl

How to generate the timeline for this article

Install cwltool

pip install cwltool
cwltool src/cwl/tools/timeline.cwl --reportfile timeline.html

Note that the tools requires Docker for runningthe computing environment, see the file timeline/Dockerfile for the definition of the image used in the .cwl file.

Run on Binder

MyBinder is a tool for creating executable computing environments based on standard and widely used dependency management files. You can easily run important parts of the analysis for the manuscript by clicking on the badges below. Binder will create a container using the environment configuration from the directory .binder/ and provide you with an interactive environment to execute notebooks or scripts.

  • Scimetric journal frequency analysis of RCR and metadata terms (opens a Jupyter Notebook) Binder
  • Create Figure 2 from the paper (R Markdown notebook, open the file src/timeline/timeline.Rmd manually in RStudio) Binder
  • Create word cloud from cited abstracts (run R script src/wordcloud/wordcloud.R) Binder

For development purposes, you can also run repo2docker locally in the directory of the repository.

repo2docker --editable .

License

CC0