Overview

February 9, 2026 ยท View on GitHub

Onyx

Onyx Deep Research Benchmark Results

Overview

This repository contains the Onyx submission to the Deep Research Benchmark.

The latest leaderboard can be accessed here: Deep Research Benchmark Leaderboard.

Benchmark Performance

RankModelOverallComprehensivenessInsightInstruction FollowingReadability
1Onyx54.92 ๐Ÿฅ‡55.07 ๐Ÿฅ‡57.10 ๐Ÿฅ‡52.99 ๐Ÿฅ‡52.41 ๐Ÿฅˆ
2Cellcog54.54 ๐Ÿฅˆ54.43 ๐Ÿฅˆ56.22 ๐Ÿฅˆ52.76 ๐Ÿฅˆ53.14 ๐Ÿฅ‡
3Qianfan-DeepResearch Pro54.2255.07 ๐Ÿฅ‡56.0951.7752.12
4Qianfan-DeepResearch53.0752.6555.4451.6151.21
5Tavily Research52.4452.8453.5951.9249.21
6Thinkdepthai DeepResearch52.4352.0253.8852.0450.12
7Salesforce AIR50.6550.0051.0950.7750.32
8LangChain Open Deep Research (GPT-5, with Gensee search)50.6050.0650.7651.3149.72
9Gemini 2.5 Pro Deep Research49.7149.5149.4550.1250.00
10LangChain Open Deep Research (GPT-5 with Tavily)49.3349.8047.3451.0548.99

Legend: ๐Ÿฅ‡ = best in column ยท ๐Ÿฅˆ = second best

Other Notable Mentions:

RankModelOverallComprehensivenessInsightInstruction FollowingReadability
11OpenAI Deep Research46.4546.4643.7349.3947.22
12Claude Research45.0045.3442.7947.5844.66
17Perplexity Deep Research40.4639.135.6546.1143.08

Onyx Deep Research

Input Plan
Agent Answer

Additional Info

Onyx is a production system with an emphasis on user experience. Additional product/UX constraints that we feel are important which were taken into account for the generation of the answers for the benchmark are as follows:

  • All research + answer generations for any given question are capped to 30 minutes maximum.
  • Outputs were tuned for human usability and conciseness - the reports are typically around 10,000 tokens (though the longest can be upwards of 20,000 tokens).
  • The biggest allowed latency between any user facing output is 2 minutes.

Reproducing Results

The reports submitted for the benchmark were produced using this nightly build of Onyx. External dependencies used for this particular set of results were:

The same Deep Research flow can be run with other LLMs and Web Search/Crawler APIs (or Onyx's built in), but results will vary based on the setup.

Files Included

All of the standard files for the benchmark in their standard format are included under eval_results:

  • onyx_raw_results.jsonl contains the raw final outputs from the Onyx system.
  • onyx_cleaned_results.jsonl contains the same contents as above but citations are stripped.
  • raw_results.jsonl contains the scores across all of the metrics for each of the reports individually.
  • race_result.txt contains the final aggregate scores.

The report_logs directory contains all of the reports along with all of the research agent tasks. The research agent tasks are unique to the Onyx system, they are not required for the benchmark and simply included for the reader's interest.

Acknowledgements

Thank you to the DeepResearch Bench team for creating the benchmark and maintaining the leaderboard!

Du, M., Xu, B., Zhu, C., Wang, X., & Mao, Z. (2025). DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents. arXiv preprint arXiv:2506.11763.

BibTeX:

@article{du2025deepresearch,
  title={DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents},
  author={Du, Mingxuan and Xu, Benfeng and Zhu, Chiwei and Wang, Xiaorui and Mao, Zhendong},
  journal={arXiv preprint arXiv:2506.11763},
  year={2025},
  url={https://arxiv.org/abs/2506.11763}
}