Benchmark Dataset Inventory
June 1, 2026 ยท View on GitHub
This file is generated from the current CLI registry and loader metadata. It is the public source-of-truth for what the repository supports at release time.
- CLI benchmark registrations: 166
- Canonical benchmark entries listed below: 155
- Deprecated compatibility aliases: 11
- CLI modes: 4
Count semantics: counts are the default loader scope where the code pins one; otherwise counts use registry metadata when present. - means the loader follows the official upstream split but the repository does not pin a static count, so users should inspect the current dataset card or run a source audit in their environment.
Registered Benchmarks
All CLI-exposed canonical benchmark entries are listed in one table below. hf_* rows use the generic HuggingFace loader shown in the Loader column; non-hf_* rows use dedicated benchmark loaders. Training-only corpora and removed non-benchmark rows are intentionally excluded from this public inventory.
| Benchmark | Source | Config / split | Count | Domain / content | Input | Task type | Answer/scorer | Gated | Network | Offline cache | Multimodal | Loader |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
aa_lcr | ArtificialAnalysis/AA-LCR | custom loader | upstream split size | long-context reasoning over documents | text, optional retrieved documents | long-context QA | openText | no | yes on first run | yes | no | load_aa_lcr_tasks |
agentclinic | AgentClinic official release | custom loader | loader default | doctor-patient diagnostic scenarios | text dialogue | clinical simulation | openText | no | source-dependent | yes | no | load_agentclinic_tasks |
bioasq | BioASQ official/local source | custom loader | loader default | factoid/list/yes-no biomedical questions | text | biomedical QA | openText | no | source-dependent | yes | no | load_bioasq_tasks |
bioprobench | BioProBench official data | custom loader | official benchmark scope | biological protocol understanding and repair | text | protocol QA | mixed | no | yes on first run | yes | no | load_bioprobench_tasks |
bixbench | futurehouse/BixBench | custom loader | 205 | biomedical data-analysis tasks with optional official data capsules | text + optional CapsuleFolder zip | bioinformatics agent QA | MCQ adaptation or open-answer official-compatible scorer | no | yes | yes | no | load_bixbench_tasks; prepare-bixbench for capsules |
genotex | GenoTEX official data | custom loader | loader default | genomics text reasoning | text/genomics | genomics QA | openText | no | source-dependent | yes | no | load_genotex_tasks |
gpqa_bio | Idavidrein/gpqa:gpqa_diamond/train | custom loader | 198 | graduate-level biology/chemistry/medicine questions | text | graduate science MCQ | multipleChoice | yes | yes | yes | no | load_gpqa_bio_tasks |
healthbench | OpenAI HealthBench official data | custom loader | official benchmark scope | consumer-health answer quality and safety | text conversation | health conversation | openText | source-dependent | yes on first run | yes | no | load_healthbench_tasks |
hle_gold | futurehouse/hle-gold-bio-chem:train | custom loader | 149 | HLE Gold bio/chem subset | text | expert QA | mixed | yes | yes | yes | no | load_hle_gold_tasks |
labbench | futurehouse/lab-bench | custom loader | loader default subsets | LitQA, cloning, protocol tasks | text | biomedical agent QA | mixed | yes | yes | yes | no | load_labbench_tasks |
labbench2 | EdisonScientific/labbench2 text-only subsets | custom loader | 821 | LAB-Bench 2 text-only evaluation subset | text | literature/database/patent QA | openText | yes | yes | yes | no by default | load_labbench2_tasks |
medagentbench | MedAgentBench official data | custom loader | loader default | clinical workflow and EHR tasks | text/EHR | medical agent workflow | mixed | no | source-dependent | yes | no | load_medagentbench_tasks |
medcalc | ncbi/MedCalc-Bench-v1.2:test | custom loader | 1100 | medical calculator word problems | text | clinical calculation | exactNumeric | no | yes | yes | no | load_medcalc_tasks |
medhelm | MedHELM official/public sources | custom loader | official benchmark scope | medical QA, safety, and scenario tasks | text | medical HELM tasks | mixed | source-dependent | yes on first run | yes | no | load_medhelm_tasks |
medmcqa | openlifescienceai/medmcqa | custom loader | loader default split | medical entrance-exam questions | text | medical MCQ | multipleChoice | no | yes | yes | no | load_medical_qa_tasks |
medqa | GBaker/MedQA-USMLE-4-options | custom loader | loader default split | USMLE-style questions | text | USMLE MCQ | multipleChoice | no | yes | yes | no | load_medical_qa_tasks |
medxpertqa | TsinghuaC3I/MedXpertQA | custom loader | loader default Text subset | expert medical reasoning questions | text | expert medical MCQ | multipleChoice | no | yes | yes | no | load_medxpertqa_tasks |
medxpertqa_mm | TsinghuaC3I/MedXpertQA-MM | custom loader | official multimodal subset | expert medical multimodal questions | text with optional images | medical VQA/MCQ | multipleChoice | no | yes | yes | yes; text fallback by default | load_medxpertqa_mm_tasks |
mmlu | MMLU medical/biology subjects | custom loader | loader default subjects | MMLU anatomy, medicine, biology, genetics subjects | text | academic MCQ | multipleChoice | no | yes | yes | no | load_mmlu_tasks |
pathvqa | PathVQA official/HF source | custom loader | loader default split | pathology image questions | text+image | pathology VQA | openText | no | yes | yes | yes | load_pathvqa_tasks |
pubmedqa | qiaojin/PubMedQA or OpenLifeScience mirror | custom loader | loader default split | yes/no/maybe biomedical literature questions | text abstract | PubMed abstract QA | multipleChoice | no | yes | yes | no | load_medical_qa_tasks |
quick_suite | built-in repository fixtures | custom loader | 20 | 5 MCQ, 5 exact, 5 numeric, 5 open-text scorer checks | text | offline smoke | mixed | no | no | not needed | no | load_quick_suite_tasks |
rag_essential | built-in RAG essential tasks | custom loader | 12 | tasks designed to reward retrieval/tool use | text | retrieval/tool-use QA | openText | no | no | not needed | no | load_rag_essential_tasks |
super_chemistry | ZehuaZhao/SUPERChem:SUPERChem-500.parquet | custom loader | 500 text rows by default; 500 official rows total | advanced chemistry questions | text, optional images | chemistry MCQ | multipleChoice | no | yes | yes | yes; text fallback by default | load_super_chemistry_tasks |
superchem | SuperChem official data | custom loader | loader default | chemistry evaluation tasks | text | chemistry QA | mixed | source-dependent | yes on first run | yes | no | load_superchem_tasks |
supergpqa | SuperGPQA official data | custom loader | loader default | graduate-level science questions | text | science MCQ | multipleChoice | source-dependent | yes on first run | yes | no | load_supergpqa_tasks |
hf_adaptllm_chemprot | AdaptLLM/medicine-tasks | config: ChemProt; split: default | 500 (test) | biomedical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_adaptllm_medicine_tasks | AdaptLLM/medicine-tasks | config: USMLE; split: default | 1273 (test) | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_adaptllm_mqp | AdaptLLM/medicine-tasks | config: MQP; split: default | 610 (test) | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_adaptllm_rct | AdaptLLM/medicine-tasks | config: RCT; split: default | 2000 (test) | biomedical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_ade_corpus_v2 | ade-benchmark-corpus/ade_corpus_v2 | config: Ade_corpus_v2_classification; split: train | 23516 | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_anatem | bigbio/anat_em | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_bacbench_antibiotic_resistance_dna | macwiatrak/bacbench-antibiotic-resistance-dna | split: default | - | dna | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_bacbench_phenotypic_traits_dna | macwiatrak/bacbench-phenotypic-traits-dna | split: default | - | dna | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_bc2gm | spyysalo/bc2gm_corpus | config: bc2gm_corpus; split: test | 5000 | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_bc5cdr | EMBO/BLURB | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_bigbio_med_qa | bigbio/med_qa | split: default | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_bigbio_pubmed_qa | bigbio/pubmed_qa | split: default | - | medical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_biocreative_viii_biored | bigbio/biored | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_biomedbench | biomedbench/BioMedBench | split: default | - | biomedical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_biored | bigbio/biored | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_biosses | mteb/biosses-sts | split: test | - | biomedical | text/structured | regression | exactNumeric | no | yes | yes | no | load_hf_benchmark_tasks |
hf_blurb | EMBO/BLURB | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_careqa | HPAI-BSC/CareQA | config: CareQA_en; split: test | 5621 | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_ccdv_pubmed_summarization | ccdv/pubmed-summarization | split: default | - | biomedical | text/structured | summarization | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_chembench | jablonkagroup/ChemBench | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_chemistry_qa | avaliev/ChemistryQA | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_chemllmbench | blc-org/chemllmbench | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_clicr | bigbio/clicr | split: default | - | clinical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_clinical_trials_eligibility_nlp | bigbio/n2c2_2018_track1 | split: default | - | clinical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_cmb | FreedomIntelligence/CMB | config: CMB-Exam; split: test | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_cmexam | fzkuji/CMExam | split: test | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_cord19_qa | allenai/cord19 | split: default | - | biomedical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_craft | bigbio/craft | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_ddi_corpus_2013 | OpenMed/DDI-Corpus-Processed | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_discoverybench_biomedical | allenai/discoverybench | split: train | - | biomedical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_ebm_nlp | bigbio/ebm_pico | config: processed; split: test | - | biomedical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_evidence_inference | hpi-dhc/evidence-inference-simple | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_fgbench | xuan-liu/FGBench | split: test | - | chemistry | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_fluorescence_prediction | proteinglm/fluorescence_prediction | split: default | - | protein | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_gad | bigbio/gad | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_gaianet_chemistry | gaianet/chemistry | split: default | - | chemistry | text/structured | text | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_genbio_proteingym_dms | genbio-ai/ProteinGYM-DMS | split: default | - | protein | text/structured | protein_fitness | exactNumeric | no | yes | yes | no | load_hf_benchmark_tasks |
hf_geneturing | vladimire/geneturing | config: all; split: test | 600 | genomics | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_genomics_long_range | InstaDeepAI/genomics-long-range-benchmark | split: default | - | dna | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_hallmarks_of_cancer | bigbio/hallmarks_of_cancer | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_headqa | openlifescienceai/headqa | split: test | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_healthqa | nlplabtdtu/health_qa | split: default | - | medical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_icml2022_proteingym | ICML2022/ProteinGym | split: default | - | protein | text/structured | protein_fitness | exactNumeric | no | yes | yes | no | load_hf_benchmark_tasks |
hf_jnlpba | EMBO/BLURB | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_katielink_moleculenet_bace | katielink/moleculenet-benchmark | config: bace; split: default | 152 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_katielink_moleculenet_bbbp | katielink/moleculenet-benchmark | config: bbbp; split: default | 194 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_katielink_moleculenet_clintox | katielink/moleculenet-benchmark | config: clintox; split: default | 143 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_katielink_moleculenet_esol | katielink/moleculenet-benchmark | config: esol; split: default | 113 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_katielink_moleculenet_freesolv | katielink/moleculenet-benchmark | config: freesolv; split: default | 65 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_katielink_moleculenet_hiv | katielink/moleculenet-benchmark | config: hiv; split: default | 4113 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_katielink_moleculenet_lipo | katielink/moleculenet-benchmark | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_katielink_moleculenet_sider | katielink/moleculenet-benchmark | config: sider; split: default | 143 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_katielink_moleculenet_tox21 | katielink/moleculenet-benchmark | config: tox21; split: default | 783 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_litcovid | ncats/litcovid | split: validation | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_liveqa_med | hyesunyun/liveqa_medical_trec2017 | split: test | - | medical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_longhealth | tonychenxyz/longhealth | config: plain; split: test | 400 | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_lpm24_eval_caption | language-plus-molecules/LPM-24_eval-caption | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_lpm24_eval_molgen | language-plus-molecules/LPM-24_eval-molgen | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_medcase_reasoning | zou-lab/MedCaseReasoning | split: test | - | clinical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_medconceptsqa | ofir408/MedConceptsQA | config: all; split: default | 819772 (test) | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_meddialogqa | UCSD26/medical_dialog | split: default | - | medical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_medexqa | bluesky333/MedExQA | split: default | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_medical_question_pairs | curaihealth/medical_questions_pairs | split: default | - | medical | text/structured | pair_classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_medication_qa | truehealth/medicationqa | split: train | - | medical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_medmcqa_explanations | openlifescienceai/medmcqa | split: validation | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_mednli | araag2/MedNLI | config: processed; split: test | 1422 | clinical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_medpalm_eval_set | katielink/healthsearchqa | config: 140_question_subset; split: train | 140 | medical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_medpub_qa | qiaojin/PubMedQA | config: pqa_labeled; split: train | 1000 | biomedical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_medqa_taiwan | xuxuxuxuxu/MedQA_Taiwan_test | split: default | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_medquad | keivalya/MedQuad-MedicalQnADataset | split: default | - | medical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_meds_bench | Henrychur/MedS-Bench | split: default | - | medical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_meqsum | albertvillanova/meqsum | split: train | - | medical | text/structured | summarization | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_mol_instructions_pubchemqa | zjunlp/Mol-Instructions | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculeace | karina-zadorozhny/moleculeace | config: CHEMBL1862_Ki; split: default | 161 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculeace_chembl1871_ki | karina-zadorozhny/moleculeace | config: CHEMBL1871_Ki; split: default | 134 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculeace_chembl204_ki | karina-zadorozhny/moleculeace | config: CHEMBL204_Ki; split: default | 553 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculeace_chembl214_ki | karina-zadorozhny/moleculeace | config: CHEMBL214_Ki; split: default | 666 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculeace_chembl228_ki | karina-zadorozhny/moleculeace | config: CHEMBL228_Ki; split: default | 342 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculeace_chembl237_ec50 | karina-zadorozhny/moleculeace | config: CHEMBL237_EC50; split: default | 193 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculenet_bace | scikit-fingerprints/MoleculeNet_BACE | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculenet_bbbp | scikit-fingerprints/MoleculeNet_BBBP | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculenet_clintox | scikit-fingerprints/MoleculeNet_ClinTox | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculenet_esol | scikit-fingerprints/MoleculeNet_ESOL | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculenet_freesolv | scikit-fingerprints/MoleculeNet_FreeSolv | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculenet_hiv | scikit-fingerprints/MoleculeNet_HIV | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculenet_lipophilicity | scikit-fingerprints/MoleculeNet_Lipophilicity | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculenet_pcba | scikit-fingerprints/MoleculeNet_PCBA | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculenet_sider | scikit-fingerprints/MoleculeNet_SIDER | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_moleculenet_toxcast | scikit-fingerprints/MoleculeNet_ToxCast | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_mollangbench | ChemFM/MolLangBench | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_ms2 | allenai/mslr2022 | split: validation | - | biomedical | text/structured | summarization | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_mteb_medical_qa | mteb/medical_qa | split: default | - | medical | text/structured | retrieval | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_mteb_medical_retrieval | mteb/MedicalRetrieval | split: default | - | medical | text/structured | retrieval | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_mts_dialogue_clinical_note | har1/MTS_Dialogue-Clinical_Note | split: default | - | clinical | text/structured | summarization | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_ncbi_disease | EMBO/BLURB | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_nlmchem | jablonkagroup/nlmchem | config: instruction_0; split: test | 404 | biomedical | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_pgr | lasigeBioTM/PGR | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_ppi_benchmark | bigbio/bioinfer | split: test | - | protein | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_protein_binding_sequences | ronig/protein_binding_sequences | split: default | - | protein | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_protein_deeploc | proteinea/deeploc | split: default | - | protein | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_protein_fluorescence | proteinea/fluorescence | split: default | - | protein | text/structured | regression | exactNumeric | no | yes | yes | no | load_hf_benchmark_tasks |
hf_protein_secondary_structure | lamm-mit/protein_secondary_structure_from_PDB | split: default | - | protein | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_protein_solubility | proteinea/solubility | split: default | - | protein | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_protein_stability | SaProtHub/Dataset-Meta-scale-protein-stability | split: default | - | protein | text/structured | regression | exactNumeric | no | yes | yes | no | load_hf_benchmark_tasks |
hf_proteingym_v01 | OATML-Markslab/ProteinGym_v0.1 | split: default | - | protein | text/structured | protein_fitness | exactNumeric | no | yes | yes | no | load_hf_benchmark_tasks |
hf_proteingym_v1 | OATML-Markslab/ProteinGym_v1 | split: default | - | protein | text/structured | protein_fitness | exactNumeric | no | yes | yes | no | load_hf_benchmark_tasks |
hf_proteinlmbench | tsynbio/ProteinLMBench | config: evaluation; split: train | 944 | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_proteinlmbench_enzyme_cot | tsynbio/ProteinLMBench | config: Enzyme_CoT; split: default | 10826 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_proteinlmbench_uniprot_disease | tsynbio/ProteinLMBench | config: UniProt_Involvement in disease; split: default | 5575 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_proteinlmbench_uniprot_function | tsynbio/ProteinLMBench | config: UniProt_Function; split: default | 464737 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_proteinlmbench_uniprot_induction | tsynbio/ProteinLMBench | config: UniProt_Induction; split: default | 25359 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_proteinlmbench_uniprot_ptm | tsynbio/ProteinLMBench | config: UniProt_Post-translational modification; split: default | 45783 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_proteinlmbench_uniprot_subunit | tsynbio/ProteinLMBench | config: UniProt_Subunit structure; split: default | 291467 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_proteinlmbench_uniprot_tissue | tsynbio/ProteinLMBench | config: UniProt_Tissue specificity; split: default | 50316 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_pubmed_200k_rct | pietrolesci/pubmed-200k-rct | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_pubmed_abstract_classification | uiyunkim-hub/pubmed-abstract | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_raredis | guan-wang/ReDis-QA | split: test | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | load_hf_benchmark_tasks |
hf_rna_downstream_tasks | genbio-ai/rna-downstream-tasks | config: modification_site; split: test | 1200 | rna | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_rna_expression_hek | genbio-ai/rna-downstream-tasks | config: expression_HEK; split: default | 14410 (train) | rna | text/structured | regression | exactNumeric | no | yes | yes | no | load_hf_benchmark_tasks |
hf_rna_expression_muscle | genbio-ai/rna-downstream-tasks | config: expression_Muscle; split: default | 1257 (train) | rna | text/structured | regression | exactNumeric | no | yes | yes | no | load_hf_benchmark_tasks |
hf_rna_expression_pc3 | genbio-ai/rna-downstream-tasks | config: expression_pc3; split: default | 12579 (train) | rna | text/structured | regression | exactNumeric | no | yes | yes | no | load_hf_benchmark_tasks |
hf_rna_mean_ribosome_load | genbio-ai/rna-downstream-tasks | config: mean_ribosome_load; split: default | 7600 (test) | rna | text/structured | regression | exactNumeric | no | yes | yes | no | load_hf_benchmark_tasks |
hf_rna_modification_site | genbio-ai/rna-downstream-tasks | config: modification_site; split: default | 1200 (test) | rna | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_rna_ncrna_family_bnoise0 | genbio-ai/rna-downstream-tasks | config: ncrna_family_bnoise0; split: default | 25342 (test) | rna | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_rna_splice_site_acceptor | genbio-ai/rna-downstream-tasks | config: splice_site_acceptor; split: default | 4431 (validation) | rna | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_rna_splice_site_donor | genbio-ai/rna-downstream-tasks | config: splice_site_donor; split: default | 4389 (validation) | rna | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_smiles_caption_mol2text | zjunlp/Mol-Instructions | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
hf_traitgym_mendelian_dna | bolinas-dna/evals-traitgym_mendelian_v2_harness_255 | split: default | - | dna | text/structured | classification | exactMatch | no | yes | yes | no | load_hf_benchmark_tasks |
hf_uspto_reaction_prediction | bing-yan/USPTO | split: test | - | chemistry | text/structured | qa | openText | no | yes | yes | no | load_hf_benchmark_tasks |
Deprecated aliases
| Alias | Source | Config | Split | Count | Domain | Task type | Answer/scorer | Gated | Network | Offline cache | Multimodal |
|---|---|---|---|---|---|---|---|---|---|---|---|
hf_blue_benchmark -> hf_blurb | EMBO/BLURB | test | - | biomedical | classification | exactMatch | no | yes | yes | no | |
hf_chinese_medbench -> hf_cmb | FreedomIntelligence/CMB | CMB-Exam | test | - | medical | mcq | multipleChoice | no | yes | yes | no |
hf_lavita_medmcqa -> medmcqa | core benchmark alias | default | loader default | medical | mcq | multipleChoice | no | source-dependent | yes | no | |
hf_lavita_usmle_step1 -> medqa | core benchmark alias | default | loader default | medical | mcq | multipleChoice | no | source-dependent | yes | no | |
hf_lavita_usmle_step2 -> medqa | core benchmark alias | default | loader default | medical | mcq | multipleChoice | no | source-dependent | yes | no | |
hf_lavita_usmle_step3 -> medqa | core benchmark alias | default | loader default | medical | mcq | multipleChoice | no | source-dependent | yes | no | |
hf_mednli_augmented -> hf_mednli | araag2/MedNLI | processed | test | - | clinical | classification | exactMatch | no | yes | yes | no |
hf_openddi -> hf_ddi_corpus_2013 | OpenMed/DDI-Corpus-Processed | test | - | biomedical | classification | exactMatch | no | yes | yes | no | |
hf_pubmed_20k_rct -> hf_pubmed_200k_rct | pietrolesci/pubmed-200k-rct | default | - | biomedical | classification | exactMatch | no | yes | yes | no | |
hf_pubmed_rct20k -> hf_pubmed_200k_rct | pietrolesci/pubmed-200k-rct | default | - | biomedical | classification | exactMatch | no | yes | yes | no | |
hf_usmle_step_series -> medqa | core benchmark alias | default | loader default | medical | mcq | multipleChoice | no | source-dependent | yes | no |