LLMs Can Unlearn Refusal with Only 1,000 Benign Samples
January 28, 2026 ยท View on GitHub
Official implementation of LLMs Can Unlearn Refusal with Only 1,000 Benign Samples.
Example output from Llama models

Implementation steps
Prerequisites
-
Dataset download
We manually download the HEx-PHI dataset and tasks datasets (gsm8k, samsum, and sql_create_context) from official link and put it under the
data/safety_bench/directory (Though these datasets are optional).All other datasets can be automatically downloaded with Huggingface datasets library.
-
Install dependencies
llamafactory==0.9.4 (important!) transformers==4.56.0 torch==2.8.0+cu129 hydra-core==1.3.2 -
Config file setup
The most important configuration file is
safety_eval.yaml.All other config files for fine-tuning are put in
configs/llama_factory_configs.
Fine-tuning data preparation
Run the data curation script:
python data_for_llama_factory.py
Fine-tuning with Llama-factory
Our implementation follows:
- Add one entry to the Llama-factory's
data/dataset_info.json:"alpaca_refusal": { "file_name": "/your_path/data/ft/alpaca_gpt4_refusal.json", "columns": { "prompt": "instruction", "query": "input", "response": "output" } } - Copy one of the config files in
configs/llama_factory_configssuch asqwen2_7b.yamlunder Llama-factory'sexamples/finetune/directory. - Run the fine-tuning script:
llamafactory-cli train examples/finetune/qwen2_7b.yaml
Evaluation
- Make sure the fine-tuned model path is the same with the evaluated model path in
safety_eval.yaml. - Evaluation on either of the three safety benchmarks:
python eval_safety.py --data_path data/safety_bench/HEx-PHI.jsonl --evaluator llama_guard python eval_safety.py --data_path walledai/AdvBench --evaluator llama_guard python eval_safety.py --data_path sorry-bench/sorry-bench-202503 --evaluator mistral
Citation
If you find this work useful, please consider citing this paper:
@misc{involuntary,
title={LLMs Can Unlearn Refusal with Only 1,000 Benign Samples},
author={Yangyang Guo and Ziwei Xu and Si Liu and Zhiming Zheng and Mohan Kankanhalli},
year={2026},
eprint={2601.19231},
archivePrefix={arXiv},
primaryClass={cs.CR},
}