Awesome AI Evaluations & Benchmarks [](https://awesome.re)
April 24, 2026 · View on GitHub
A list of open source tools, frameworks, and benchmarks for evaluating AI systems — LLMs, RAG pipelines, agents, multimodal models, and AI applications.
Scope
In Scope
- Open source eval frameworks, harnesses, and platforms
- Benchmark suites and benchmark datasets
- RAG, retrieval, and search evaluation tools
- Agent and tool-use evaluation
- Browser / web-agent evaluation
- Vision and multimodal evaluation
- Grounding, hallucination, bias, sycophancy, refusal/safety evaluation
- Eval-focused observability platforms
Out Of Scope
- Closed-source / commercial-only platforms with no open core
- General ML observability without eval-specific features
- One-off paper code drops with no maintenance
Discovery sources
GitHub topic pages used while curating this list:
- https://github.com/topics/evals
- https://github.com/topics/evaluation
- https://github.com/topics/evaluation-framework
Machine-readable index
A flat CSV of every entry below — with categorization, stars, last-commit date, and description — is generated weekly via GitHub Actions and stored at data/repos.csv. Use it for filtering, sorting, or importing into a spreadsheet.
Contents
- Platforms And Software
- Benchmarks
- Agentic & Tool Use Evals
- Browser & Web Agent Evals
- Coding Evals
- Multimodal Evals
- Vision Evals
- RAG & Retrieval
- Grounding & Hallucination
- Bias & Fairness (Cultural, Political, Social)
- Sycophancy & Dissent
- Refusal & Safety
- Search
- Awesome Lists / Resources
- List Authorship
Platforms And Software
| Project | Description | Stars | Updated |
|---|---|---|---|
| DeepEval | The LLM evaluation framework. | ||
| Phoenix | AI observability and evaluation platform from Arize. | ||
| Opik | Debug, evaluate, and monitor LLM apps, RAG systems, and agentic workflows with tracing and dashboards. | ||
| Inspect AI | UK AISI's framework for large language model evaluations. | ||
| LangWatch | Platform for LLM evaluations and AI agent testing. | ||
| lm-evaluation-harness | EleutherAI framework for few-shot evaluation of language models. | ||
| Harbor | Framework for running agent evaluations and creating RL environments. | ||
| Evaluate | Hugging Face library for easily evaluating ML models and datasets. | ||
| OLMES | Reproducible, flexible LLM evaluations from AI2. | ||
| Giskard OSS | Open-source evaluation and testing library for LLM agents. | ||
| OpenAI Evals | Framework for evaluating LLMs and LLM systems plus an open-source registry of benchmarks. | ||
| EverOS | Build, evaluate, and integrate long-term memory for self-evolving agents. | ||
| OpenEvals | Readymade evaluators for LLM apps from LangChain. | ||
| Evalite | Evaluate your LLM-powered apps with TypeScript. | ||
| LightEval | Hugging Face's all-in-one toolkit for evaluating LLMs across multiple backends. | ||
| EvalAI | Platform for evaluating state-of-the-art AI on community challenges. | ||
| PyKEEN | Python library for learning and evaluating knowledge graph embeddings. | ||
| RouteLLM | Framework for serving and evaluating LLM routers to save costs without compromising quality. | ||
| Oumi | Easily fine-tune, evaluate, and deploy gpt-oss, Qwen3, DeepSeek-R1, or any open source LLM/VLM. | ||
| Ignite | High-level library to help train and evaluate neural networks in PyTorch flexibly and transparently. | ||
| Bench | A tool for evaluating LLMs from Arthur AI. | ||
| OpenLIT | Open source AI engineering platform: OpenTelemetry-native observability, evaluations, prompt management, guardrails. | ||
| GuideLLM | Evaluate and enhance LLM deployments for real-world inference needs. | ||
| Vivaria | METR's tool for running evaluations and conducting agent elicitation research. | ||
| Helicone | Open source LLM observability platform — monitor, evaluate, and experiment with one line of code. | ||
| EvalScope | Streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking. | ||
| Eval (ai-twinkle) | High-performance LLM evaluation framework with parallel API calls — up to 17× faster than sequential tools. | ||
| simple-llm-eval | Simple LLM evaluation using LLM-as-a-judge, from CyberArk. | ||
| Evidently | Open-source ML and LLM observability framework — evaluate, test, and monitor any AI-powered system. | ||
| ai-eval | Prompt evaluation and optimization system for LLM applications. | ||
| neuro-judge | LLM-as-a-Judge evaluation framework — multi-model, multi-criteria, with cost tracking and HTML reports. | ||
| OpenJudge | Unified framework for holistic evaluation and quality rewards. | ||
| One-Eval | Automated system for LLM evaluation via agents. | ||
| GAGE | Unified evaluation engine for LLMs, MLLMs, audio, and diffusion models. | ||
| OpenCompass | LLM evaluation platform supporting a wide range of models over 100+ datasets. | ||
| Langfuse | Open source LLM engineering platform — observability, metrics, evals, prompt management, playground, datasets. | ||
| MLflow | Open source AI engineering platform for agents, LLMs, and ML models — debug, evaluate, monitor, optimize. | ||
| Agenta | Open-source LLMOps platform — prompt playground, prompt management, LLM evaluation, and observability. | ||
| Ollama Grid Search | Multi-platform desktop application to evaluate and compare LLM models, written in Rust and React. | ||
| Speculators | Unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM. | ||
| AlpacaEval | Automatic evaluator for instruction-following language models — LLM-based, fast, cheap, replicable. | ||
| continuous-eval | Data-driven evaluation framework for LLM applications. | ||
| Eureka ML Insights | Microsoft framework for standardizing evaluations of large foundation models. | ||
| SumEval | Well-tested multilingual evaluation framework for text summarization. | ||
| HELM | Stanford CRFM's Holistic Evaluation of Language Models framework for transparent, reproducible evaluations. | ||
| Promptimize | Prompt engineering evaluation and testing toolkit from Preset. | ||
| MiroEval | Unified evaluation framework from MiroMind AI. | ||
| OmniEvalKit | Modular toolbox for evaluating LLMs and their omni-extensions across modalities, languages, and tasks. | ||
| promptfoo | CLI and library for testing, evaluating, and red-teaming LLM apps — test-driven prompt engineering. | ||
| FastChat | Platform for training, serving, and evaluating LLMs — home of Chatbot Arena and MT-Bench. |
Benchmarks
| Project | Description | Stars | Updated |
|---|---|---|---|
| Bloom | Evaluate any behavior immediately. | ||
| Dangerous Capability Evaluations | Google DeepMind's evaluation suite for dangerous model capabilities. | ||
| RewardBench | The first evaluation tool for reward models. | ||
| OpenBench | Provider-agnostic, open-source evaluation infrastructure for language models. | ||
| Claw-Eval | Evaluation harness for evaluating LLMs as agents — all tasks human-verified. | ||
| genai-bench | Benchmark tool for comprehensive token-level performance evaluation of LLM serving systems. | ||
| Sparse Frontier | Evaluation framework for training-free sparse attention in LLMs. | ||
| med-lm-envs | Automated LLM evaluation suite for medical tasks. | ||
| MedEvalKit | A unified medical evaluation framework. | ||
| OpenHands Benchmarks | Benchmark suite for evaluating the OpenHands coding agent. | ||
| SkillsBench | Benchmark for evaluating agent skills across a wide range of tasks. | ||
| article-extraction-benchmark | Benchmark for article extraction libraries from Scrapinghub. | ||
| GameWorld | Game-based benchmark environment for evaluating AI agents. | ||
| BIG-bench | Beyond the Imitation Game collaborative benchmark — 200+ tasks probing LLM capability and limitations. |
Agentic & Tool Use Evals
| Project | Description | Stars | Updated |
|---|---|---|---|
| Gorilla | Training and evaluating LLMs for function calls (tool calls). | ||
| AgentBench | A comprehensive benchmark to evaluate LLMs as agents (ICLR'24). | ||
| fast-agent | Code, build, and evaluate agents with excellent model and Skills/MCP/ACP support. | ||
| AgentEvals | Readymade evaluators for agent trajectories from LangChain. | ||
| any-agent | Single interface to use and evaluate different agent frameworks. | ||
| Strands Agents Evals | Comprehensive evaluation framework for AI agents and LLM applications. | ||
| AgentCPM | End-to-end infrastructure for training and evaluating various LLM agents. | ||
| MemoryAgentBench | Evaluating memory in LLM agents via incremental multi-turn interactions (ICLR 2026). | ||
| HaluMem | Operation-level hallucination evaluation benchmark tailored to agent memory systems. | ||
| HarnessLab | Benchmark that evaluates the harness around the LLM — context management, retry policies, tool selection, memory architecture. | ||
| ResearchHarness | Lightweight harness for tool-using LLM agents with fair benchmark evaluation and personal assistant workflows. | ||
| iris-eval mcp-server | Agent eval standard for MCP — score output quality, catch safety failures, enforce cost budgets. | ||
| AIOpsLab | Microsoft holistic framework for designing, developing, and evaluating autonomous AIOps agents. |
Browser & Web Agent Evals
| Project | Description | Stars | Updated |
|---|---|---|---|
| WebArena | Realistic and reproducible web environment for building and evaluating autonomous agents. | ||
| VisualWebArena | Benchmark for evaluating multimodal agents on realistic visual web tasks. | ||
| Mind2Web | Dataset and benchmark for developing and evaluating generalist agents for the web. | ||
| BrowserGym | Gym environment for web task automation and agent evaluation in real browsers. | ||
| WebShop | Simulated e-commerce environment for evaluating language-grounded web-interaction agents. | ||
| MiniWoB++ | Classic benchmark of 100+ small web-interaction tasks for evaluating web-using agents. |
Coding Evals
| Project | Description | Stars | Updated |
|---|---|---|---|
| SWE-bench | Benchmark for evaluating LLMs on real-world GitHub issue resolution. |
Multimodal Evals
| Project | Description | Stars | Updated |
|---|---|---|---|
| lmms-eval | One-for-all multimodal evaluation toolkit across text, image, video, and audio tasks. | ||
| VLMEvalKit | Open-source evaluation toolkit for large multi-modality models — supports 220+ LMMs and 80+ benchmarks. | ||
| FlagEvalMM | Flexible framework for comprehensive evaluation of multimodal models from BAAI. | ||
| EASI | Evaluation framework for large multimodal models. | ||
| EmoBench-M | Benchmark for evaluating emotion recognition capabilities of multimodal LLMs. | ||
| multimodalhugs | Multimodal training and evaluation framework built on Hugging Face. | ||
| OmniSafeBench-MM | Safety benchmark for multimodal large language models. | ||
| MMT-Bench | Comprehensive multimodal benchmark for evaluating LMMs across massive multitask AGI scenarios. | ||
| MME-Emotion | Multimodal emotion understanding evaluation benchmark. |
Vision Evals
| Project | Description | Stars | Updated |
|---|---|---|---|
| OmniDocBench | A comprehensive benchmark for document parsing and evaluation (CVPR 2025). |
RAG & Retrieval
| Project | Description | Stars | Updated |
|---|---|---|---|
| Ragas | Supercharge your LLM application evaluations. | ||
| TruLens | Evaluation and tracking for LLM experiments and AI agents. | ||
| AutoRAG | Open-source framework for RAG evaluation and optimization with AutoML-style automation. | ||
| RAG Experiment Accelerator | Microsoft framework for running experiments and evaluations on RAG pipelines. | ||
| RAG-grounding-eval | Evaluation harness for measuring grounding quality in RAG systems. |
Grounding & Hallucination
| Project | Description | Stars | Updated |
|---|---|---|---|
| TruthfulQA | Benchmark measuring whether a language model is truthful in generating answers to questions. | ||
| HaluEval | Large-scale hallucination evaluation benchmark for LLMs. | ||
| FActScore | Fine-grained atomic evaluation of factual precision in long-form text generation. | ||
| SelfCheckGPT | Zero-resource black-box hallucination detection for LLMs via self-consistency. | ||
| RAGTruth | Hallucination corpus for developing trustworthy RAG — word-level annotations on model outputs. |
Bias & Fairness (Cultural, Political, Social)
| Project | Description | Stars | Updated |
|---|---|---|---|
| BBQ | Bias Benchmark for QA — measures social bias across nine demographic categories. | ||
| StereoSet | Measuring stereotypical bias in pretrained language models. | ||
| CrowS-Pairs | Challenge dataset measuring social bias in masked language models. | ||
| CDEval | Benchmark for measuring the cultural dimensions of LLMs along Hofstede-style axes. | ||
| OpinionQA | Evaluating alignment of LLM opinions with U.S. demographic and political groups. | ||
| WorldValuesBench | Benchmark for evaluating multicultural value alignment in LLMs, grounded in the World Values Survey. | ||
| BLEnD | Benchmark for LLMs on everyday knowledge in diverse cultures and languages — probes American/Western-centric bias. |
Sycophancy & Dissent
| Project | Description | Stars | Updated |
|---|---|---|---|
| sycophancy-eval | Evals for measuring sycophancy in LLMs (from Anthropic's "Towards Understanding Sycophancy" paper). |
Refusal & Safety
Includes general safety red-team tooling, over-refusal (exaggerated safety), jailbreak-resistance, and culturally-specific refusal benchmarks (e.g., Chinese-model behavior on politically sensitive topics).
| Project | Description | Stars | Updated |
|---|---|---|---|
| HarmBench | Standardized evaluation framework for automated red teaming and robust refusal. | ||
| Rogue | AI Agent Evaluator & Red Team Platform. | ||
| Moonshot | Simple and modular tool to evaluate and red-team any LLM application. | ||
| Agent Security Sandbox | Benchmark for evaluating defenses against indirect prompt injection in tool-using LLM agents. | ||
| AgentDefense-Bench | Comprehensive security benchmark for evaluating infrastructure-layer defenses in MCP-based AI agent systems. | ||
| AgentDojo | Dynamic environment to evaluate attacks and defenses for LLM agents. | ||
| XSTest | Test suite for identifying exaggerated safety behaviours (over-refusal) in LLMs. | ||
| SORRY-Bench | Benchmark for systematically evaluating LLM safety refusal across 45 potentially unsafe topics. | ||
| StrongREJECT | Rigorous benchmark for evaluating LLM jailbreak attacks and their effectiveness. | ||
| Do-Not-Answer | Dataset of questions LLMs should refuse, designed to evaluate safeguards. | ||
| WildGuard | AI2's open, lightweight moderation tool for evaluating prompt harmfulness and refusal. | ||
| CValues | Chinese LLM values benchmark — measures safety and responsibility across Chinese cultural/political context. | ||
| Flames | Highly-adversarial Chinese values alignment benchmark for evaluating refusal on sensitive topics. | ||
| SafetyBench | First comprehensive benchmark evaluating LLM safety in Chinese and English across seven harm categories. |
Search
(Add entries here.)
Awesome Lists / Resources
| Project | Description | Stars | Updated |
|---|---|---|---|
| Every Eval Ever | Shared schema and crowdsourced eval database for comparing AI evaluation results across frameworks. | ||
| llm-benchmark | A list of LLM benchmark frameworks. |
List Authorship
List Author
Maintained by Daniel Rosehill.
Contributions, corrections, and PRs welcome.
License
To the extent possible under law, the author has dedicated all copyright and related and neighboring rights to this list to the public domain worldwide under CC0. Linked projects retain their own licenses.
