Efficient-CoT-LRMs

April 1, 2025 · View on GitHub

Chain of Thoughts (CoT) is hot now. But we can observe that CoT is also getting longer and sometimes overthinks.

Here, we track the latest news of how to keep CoT/LRMs more efficient.

Topics included:

  • Prompting-based
  • Budget Control
  • Compress in Latent Space
  • Summarization
  • Skip Something
  • KV Cache Management
  • Reinforment Learning
  • General Topic

Here we go.

Prompting-based

TimeTitlecodepublished
07/03/2025Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching link/
25/02/2025Chain of Draft: Thinking Faster by Writing Lesslink/
17/02/2025SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs//
04/06/2024Break the Chain: Large Language Models Can be Shortcut Reasoners //
28/07/2023Skeleton-of-Thought: Prompting LLMs for Efficient Parallel GenerationlinkICLR2024

Budget Control

TimeTitlecodepublished
06/03/2025DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models//
24/12/2024Token-budget-aware llm reasoning link/
29/07/2024Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost//

Compress in Latent Space

TimeTitlecodepublished
24/02/2025Reasoning with Latent Thoughts: On the Power of Looped Transformers/ICLR2025
13/02/2025CoT-Valve: Length-Compressible Chain-of-Thought Tuninglink/
31/01/2025Efficient Reasoning with Hidden Thinkinglink/
17/12/2024Compressed Chain of Thought: Efficient Reasoning Through Dense Representations//
09/12/2024Training Large Language Models to Reason in a Continuous Latent Spacelink/
23/03/2024From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step//

Summarization

TimeTitlecodepublished
09/03/2025InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models//
21/02/2025LightThinker: Thinking Step-by-Step Compression)//
16/12/2024C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness/AAAI2025
12/02/2024Anchor-based Large Language Models (related work)linkACL2024

Skip Something

TimeTitlecodepublished
18/02/2025Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models//
17/02/2025TokenSkip: Controllable Chain-of-Thought Compression in LLMslink/
04/11/2024Can Language Models Learn to Skip Steps?/NeurIPS2024)

KV Cache Management

TimeTitlecodepublished
21/02/2025LightThinker: Thinking Step-by-Step Compression)//
16/02/2025Efficient Long-Decoding Inference with Reasoning-Aware Attention Sparsity//

Reinforment Learning

TimeTitlecodepublished
06/03/2025L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learninglink/
22/01/2025O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruninglink/
22/01/2025Kimi k1.5: Scaling Reinforcement Learning with LLMs//

General Topic

TimeTitlecodepublished
19/02/2025AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence//
30/12/2024Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs//
10/01/2024The Impact of Reasoning Step Length on Large Language ModelslinkACL2024findings

About us:

PEILab, Hong Kong University of Science and Technology (HKUST)