llm-benchmarkJanuary 17, 2024 ยท View on GitHubA list of comprehensive LLM evaluation frameworks. Contributions welcome! BenchmarkRelease DateRepositoryPaper/BlogDataset NumberAspectLicenceHELM---https://github.com/stanford-crfm/helmHolistic Evaluation of Language Models42------BIG-bench---https://github.com/google/BIG-benchBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language models214------BigBIO---https://github.com/bigscience-workshop/biomedicalBigBio: A Framework for Data-Centric Biomedical Natural Language Processing126------BigScience Evaluation---https://github.com/bigscience-workshop/evaluation---28------Language Model Evaluation Harness---https://github.com/EleutherAI/lm-evaluation-harnessEvaluating Large Language Models (LLMs) with Eleuther AI Evaluating LLMs56------Scholar Evals---https://github.com/scholar-org/scholar-evals------------Code Generation LM Evaluation Harness---https://github.com/bigcode-project/bigcode-evaluation-harness---13------Chatbot Arena---https://github.com/lm-sys/FastChat------------GLUE---https://github.com/nyu-mll/jiant---11------SuperGLUE---https://github.com/nyu-mll/jiant---10------CLUE---https://github.com/CLUEbenchmark/CLUE---9------CodeXGLUE---https://github.com/microsoft/CodeXGLUE---10------