S1-Bench: A Simple Benchmark for Evaluating System 1 Thinking Capability of Large Reasoning Models

May 1, 2026 ยท View on GitHub

HuggingFace Arxiv modelscope opencompass

News

  • [2025/05/08] ๐ŸŽ‰ The paper was accepted by IJCAI 2026 (Main Track)!
  • [2025/05/08] ๐Ÿ”ฅ The evaluation results of Qwen3 have been added!
  • [2025/04/24] ๐Ÿ“ข We released S1-Bench dataset hosted on ModelScope and OpenCompass.
  • [2025/04/15] ๐Ÿš€ The paper is now publicly available on arXiv, alphaxiv and Huggingface Daily Papers.
  • [2025/04/13] ๐Ÿ“ข We released S1-Bench dataset hosted on Huggingface.
  • [2025/04/13] We released our code source.

How to use our project?

Before running our code, download the open-source LRMs.

Model IDAbbreviationURL
Qwen3-1.7BQwen3-1.7Bhttps://huggingface.co/Qwen/Qwen3-1.7B
Qwen3-8BQwen3-8Bhttps://huggingface.co/Qwen/Qwen3-8B
Qwen3-14BQwen3-14Bhttps://huggingface.co/Qwen/Qwen3-14B
Qwen3-32BQwen3-32Bhttps://huggingface.co/Qwen/Qwen3-32B
Qwen3-30B-A3BQwen3-30B-A3Bhttps://huggingface.co/Qwen/Qwen3-30B-A3B
Qwen3-235B-A22BQwen3-235B-A22Bhttps://huggingface.co/Qwen/Qwen3-235B-A22B
Qwen3-32BQwen3-32Bhttps://huggingface.co/Qwen/Qwen3-32B
DeepSeek-R1-Distill-Qwen-1.5BDS-R1-1.5Bhttps://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
DeepSeek-R1-Distill-Qwen-7BDS-R1-7Bhttps://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B
DeepSeek-R1-Distill-Llama-8BDS-R1-8Bhttps://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-8B
DeepSeek-R1-Distill-Qwen-14BDS-R1-14Bhttps://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
DeepSeek-R1-Distill-Qwen-32BDS-R1-32Bhttps://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
DeepSeek-R1-Distill-Llama-70BDS-R1-70Bhttps://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B
DeepSeek-R1DS-R1https://huggingface.co/deepseek-ai/DeepSeek-R1
Light-R1-7B-DSL-R1-7B-DShttps://huggingface.co/qihoo360/Light-R1-7B-DS
Light-R1-14B-DSL-R1-14B-DShttps://huggingface.co/qihoo360/Light-R1-14B-DS
Light-R1-32B-DSL-R1-32B-DS{https://huggingface.co/qihoo360/Light-R1-32B-DS
Light-R1-32BL-R1-32Bhttps://huggingface.co/qihoo360/Light-R1-32B
s1.1-7Bs1.1-7Bhttps://huggingface.co/simplescaling/s1.1-7B
s1.1-14Bs1.1-14Bhttps://huggingface.co/simplescaling/s1.1-14B
s1.1-32Bs1.1-32Bhttps://huggingface.co/simplescaling/s1.1-32B
EXAONE-Deep-2.4BEXAONE-2.4Bhttps://huggingface.co/LGAI-EXAONE/EXAONE-Deep-2.4B
EXAONE-Deep-7.8BEXAONE-7.8Bhttps://huggingface.co/LGAI-EXAONE/EXAONE-Deep-7.8B
EXAONE-Deep-32BEXAONE-32Bhttps://huggingface.co/LGAI-EXAONE/EXAONE-Deep-32B
Llama-3.1-Nemotron-Nano-8B-v1Nemotron-8Bhttps://huggingface.co/nvidia/Llama-3.1-Nemotron-Nano-8B-v1
Llama-3.3-Nemotron-Super-49B-v1Nemotron-49Bhttps://huggingface.co/nvidia/Llama-3.3-Nemotron-Super-49B-v1
Sky-T1-32B-FlashSky-T1-32Bhttps://huggingface.co/NovaSky-AI/Sky-T1-32B-Flash

Fill in the path of the open-source model in the local_model_list of get_LRM_vllm_response.py.

Execute get_LRM_vllm_response.py and run all LRMs by switching model_list[i].

python get_LRM_vllm_response.py

Next, run split_think_answer.py to obtain the several format types of the LRMs' responses.

python split_think_answer.py

Run the evaluation script get_LRM_eval.py to invoke GPT-4o for evaluating the final answers of the LRMs.

python get_LRM_eval.py

Finally, run get_acc_scores.py to obtain the evaluation results.

python get_acc_scores.py

Experiment Results

Citation

If you find our work useful, please consider citing our paper:

@misc{zhang2025s1benchsimplebenchmarkevaluating,
      title={S1-Bench: A Simple Benchmark for Evaluating System 1 Thinking Capability of Large Reasoning Models}, 
      author={Wenyuan Zhang and Shuaiyi Nie and Xinghua Zhang and Zefeng Zhang and Tingwen Liu},
      year={2025},
      eprint={2504.10368},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2504.10368}, 
}