๐ŸŽญ HER: Hierarchical Emotion Reasoning

January 30, 2026 ยท View on GitHub

๐ŸŽญ HER: Hierarchical Emotion Reasoning

HER: Human-like Reasoning and Reinforcement Learning for LLM Role-playing

Paper Dataset HER-RL HER-RM GitHub

HER Framework

HER introduces dual-layer thinking that distinguishes characters' first-person thinking from LLMs' third-person thinking for cognitive-level persona simulation.


๐Ÿ“– Overview

LLM role-playingโ€”using LLMs to simulate specific personasโ€”has emerged as a key capability in companionship, content creation, and digital games. While current models effectively capture character tones and knowledge, simulating the inner thoughts behind their behaviors remains a challenge.

HER (Hierarchical Emotion Reasoning) is a unified framework for cognitive-level persona simulation that introduces:

  • ๐Ÿง  Dual-Layer Thinking: Distinguishes characters' first-person thinking from LLMs' third-person thinking
  • ๐Ÿ“š Reasoning-Augmented Data: High-quality role-playing data via reverse engineering
  • ๐ŸŽฏ Human-Aligned Rewards: Principle-based reward models aligned with human preferences

๐Ÿ† Results

RankModelCoSER AvgCoSER SCCoSER ANCoSER CFCoSER SQMiniMax AvgMiniMax Worlds (50%)MiniMax Stories (25%)MiniMax Pref (25%)95% CI
1Claude-4.5-Opus62.4363.7464.2858.4563.2476.6267.2382.1089.90[75.5, 77.7]
2Gemini-3-Pro61.8065.9560.4258.3462.4975.6062.7283.8793.08[74.5, 76.7]
3GPT-5.161.1064.9553.9960.1365.3580.6376.6272.2197.05[79.6, 81.6]
4Gemini-2.5-Pro60.6861.0560.8057.4863.4068.2352.3682.1186.08[67.1, 69.3]
5DeepSeek-v3.258.6855.8557.0757.4464.3560.2745.8166.6482.83[59.2, 61.4]
6MiniMax-M2-her57.3060.0350.1149.3069.7784.6580.5579.9797.51[83.6, 85.7]
7DeepSeek-v3.153.5050.1553.1853.9356.7264.2251.1166.4588.21[62.9, 65.5]
8HER-RL (this model)53.1254.3347.2652.7858.1265.7359.1357.7486.90[63.0, 68.4]
9HER-SFT50.9250.5245.9949.7857.3758.4447.2952.7886.40[56.5, 60.4]
10Grok-4.1-Fast47.4049.2147.5742.6450.1748.4729.8747.5186.64[47.4, 49.5]
11Claude-4.5-Sonnet45.2147.1836.0247.5550.0969.3555.7275.6690.28[68.2, 70.5]
12Claude-3.7-Think39.7344.8431.0042.4540.6561.2550.6659.5384.15[58.5, 64.0]
13CoSER-70B35.9535.0531.1632.2845.3345.3834.3230.3282.58[43.5, 47.2]
14GPT-5-Mini32.9738.1024.6027.2042.0057.6343.3250.1193.78[55.9, 59.3]
15GPT-4o-24080627.6934.0014.9022.9038.9066.3964.9646.2389.40[64.1, 68.7]
16GPT-OSS-120B26.1232.8014.8021.5035.4060.7247.2756.6591.71[58.0, 63.4]
17Qwen3-32B22.8630.5619.6115.5230.5650.7640.3832.8289.48[48.4, 53.2]

HER achieves +30.26% improvement on CoSER and +14.97% on MiniMax Role-Play Bench.

๐Ÿš€ Quick Start

Installation

git clone https://github.com/xxx/HER.git
cd HER
pip install -r requirements.txt

Try the Demo

Chat with AI characters from classic literature:

cd chat_demo

# Interactive chat with 200 classic book scenarios
python chat_demo.py

# Show the model's thinking process
python chat_demo.py --show-think --show-rolethink
๐Ÿ’ฌ Example Conversation
๐Ÿ“– Pride and Prejudice
๐Ÿ‘ฅ AI plays: Elizabeth Bennet | You play: Mr. Darcy

โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
๐ŸŽญ ใ€Elizabeth Bennet's Responseใ€‘
โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
[His tone is light, but the air feels heavy. I cannot let him see how much
Lady Catherine's intrusion still stings.]
(takes a steadying breath, smoothing the folds of her dress)
I believe I can manage, Father. Though I must admit, I am curious about
what this letter contains.
โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•

๐Ÿ“ Project Structure

HER/
โ”œโ”€โ”€ ๐Ÿ“‚ chat_demo/              # ๐ŸŽฎ Interactive demo (try this first!)
โ”‚   โ””โ”€โ”€ chat_demo.py           # Chat with book characters
โ”‚
โ”œโ”€โ”€ ๐Ÿ“‚ data_process_code/      # ๐Ÿ”ง Data synthesis pipeline
โ”‚   โ”œโ”€โ”€ step1_data_process/    # Data cleaning & format conversion
โ”‚   โ”œโ”€โ”€ step2_gen_rolethinking/# Role thinking enhancement
โ”‚   โ”œโ”€โ”€ step3_gen_systhinking/ # System thinking generation
โ”‚   โ””โ”€โ”€ step4_setting_completion/ # Character profile enrichment
โ”‚
โ”œโ”€โ”€ ๐Ÿ“‚ training_code/          # ๐ŸŽฏ Model training pipeline
โ”‚   โ”œโ”€โ”€ step1_roleplay_sft/    # Supervised fine-tuning
โ”‚   โ”œโ”€โ”€ step2_reward_sft/      # Reward model data
โ”‚   โ”œโ”€โ”€ step3_reward_rl/       # RM training
โ”‚   โ””โ”€โ”€ step4_roleplay_rl/     # RL training
โ”‚
โ””โ”€โ”€ ๐Ÿ“‚ eval_code/              # ๐Ÿ“Š Evaluation framework
    โ”œโ”€โ”€ benchmarks/coser/      # CoSER multi-turn benchmark
    โ”œโ”€โ”€ models/                # Model adapters (vLLM, API)
    โ””โ”€โ”€ configs/               # Evaluation configs

๐Ÿง  Dual-Layer Thinking Architecture

HER introduces a hierarchical thinking mechanism:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    HER Response Structure                        โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”‚
โ”‚  โ”‚ ๐Ÿ” System Thinking (Third-Person Analysis)              โ”‚   โ”‚
โ”‚  โ”‚ "Elizabeth should respond with wit but maintain dignity. โ”‚   โ”‚
โ”‚  โ”‚  Consider her pride and her evolving feelings..."       โ”‚   โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ”‚
โ”‚                              โ†“                                   โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”‚
โ”‚  โ”‚ ๐Ÿ’ญ Role Thinking (First-Person Inner Thoughts)          โ”‚   โ”‚
โ”‚  โ”‚ [His manner is so different now. Can I trust this       โ”‚   โ”‚
โ”‚  โ”‚  change, or is it merely another game?]                 โ”‚   โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ”‚
โ”‚                              โ†“                                   โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”‚
โ”‚  โ”‚ ๐ŸŽญ Role Response (Speech + Action)                      โ”‚   โ”‚
โ”‚  โ”‚ (lifts her chin slightly) "Mr. Darcy, I find your       โ”‚   โ”‚
โ”‚  โ”‚ sudden interest in conversation rather remarkable."     โ”‚   โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
LayerPerspectivePurposeVisibility
System ThinkingThird-personHow to portray the characterHidden
Role ThinkingFirst-personCharacter's inner thoughtsSelf only
Role ResponseFirst-personSpeech and actionsAll

๐Ÿ“Š Training Pipeline

Raw Dialogues (760 books)
       โ”‚
       โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ Data Processing  โ”‚โ”€โ”€โ–บ Enhanced with dual-layer thinking
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
       โ”‚
       โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ Roleplay SFT     โ”‚โ”€โ”€โ–บ Base roleplay capability
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
       โ”‚
       โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ Reward Model     โ”‚โ”€โ”€โ–บ Principle-based quality assessment
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
       โ”‚
       โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ Roleplay RL      โ”‚โ”€โ”€โ–บ Aligned with human preferences
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
       โ”‚
       โ–ผ
   Final Model (HER-RL)

Run Training

# Step 1: Data Processing
cd data_process_code
# See DATA_PIPELINE.md for detailed workflow

# Step 2: SFT Training
cd training_code/step1_roleplay_sft
python convert_to_sft.py

# Step 3-4: Reward Model & RL
# See training_code/PIPELINE.md

๐Ÿ“ˆ Evaluation

CoSER Benchmark

Multi-agent group conversation evaluation with 4 dimensions:

MetricDescription
SC (Storyline Consistency)Character consistency across turns
AN (Anthropomorphism)Human-like behavior and emotions
CF (Character Fidelity)Adherence to character settings
SQ (Storyline Quality)Overall dialogue quality
cd eval_code

# Run evaluation
python run_coser.py \
    --actor your-model \
    --judge qwen \
    --max-rounds 20

๐Ÿ“š Documentation

DocumentDescription
Data Processing PipelineComplete data synthesis workflow
Training PipelineMulti-stage training guide
Evaluation GuideCoSER benchmark details
Chat DemoInteractive demo guide

Contact

For questions or feedback, please open an issue in the repository.

๐ŸŽ“ Citation

@article{her2025,
  title={HER: Human-like Reasoning and Reinforcement Learning for LLM Role-playing},
  author={Chengyu Du, Xintao Wang, Aili Chen, Weiyuan Li, Rui Xu, Junteng Liu, Zishan Huang, Rong Tian, Zijun Sun, Yuhao Li, Liheng Feng, Deming Ding, Pengyu Zhao, Yanghua Xiao},
  journal={arXiv preprint arXiv:2026.21459},
  year={2026}
}

๐Ÿ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.

๐Ÿค Acknowledgments

  • CoSER for the evaluation benchmark
  • MiniMax for the evaluation benchmark

Paper | Model | Data

Made with โค๏ธ for better AI role-playing