GenEval
February 13, 2026 ยท View on GitHub
Evaluation
Directly run scripts/eval/run_geneval.sh to evaluate GenEVAL. The output will be saved in $output_path.
- Set
$model_pathand$output_pathfor the path for checkpoint and log. - we Set
--thinkfor native prompt enhancing. - See GenEval for original GenEval prompts.
DPGBench
Evaluation
Directly run scripts/eval/run_dpgbench.sh to evaluate DPGBench. The output will be saved in $output_path.
- Set
$model_pathand$output_pathfor the path for checkpoint and log. - See DPGBench for original DPGBench prompts.
WISE
We modify the code in WISE for faster evaluation and download prompt json files.
Evaluation
Directly run scripts/eval/run_wise.sh to evaluate WISE. The output will be saved in $output_path.
- Set
$model_pathand$output_pathfor the path for checkpoint and log. - Use
thinkfor World Knowledge-Enhanced Textual Reasoning. scripts/eval/run_wise_refine.shsupport both World Knowledge-Enhanced Textual Reasoning and Fine-grained Editing-like Visual Refinement through twice thinking and correct initial image.
GEdit-Bench
We adopt the code in GEdit-Bench for evaluation.
Evaluation
Modify the model path, the output path in scripts/eval/run_gedit.sh. Then, run the following command:
bash script/eval/run_gedit.sh
KRIS
We modify the code in KRIS-Bench for faster evaluation.
Data prepration
Please download the benchmark data from KRIS-Bench and and place it in the KRIS_Bench directory.
Evaluation
Directly run scripts/eval/run_kris.sh to evaluate KRIS-Bench. The output will be saved in $output_path.
- Set
$model_pathand$output_pathfor the path for checkpoint and log. - Use
thinkfor World Knowledge-Enhanced Textual Reasoning. scripts/eval/run_kris_refine.shsupport both World Knowledge-Enhanced Textual Reasoning and Fine-grained Editing-like Visual Refinement through twice thinking and correct initial image.
ImgEdit
We modify the code in ImgEdit for faster evaluation.
Data prepration
Please download the benchmark data from ImgEdit-Bench and and place it in the Benchmark directory.
Evaluation
Directly run scripts/eval/run_imgedit.sh to evaluate ImgEdit-Bench. The output will be saved in $output_path.
UniGenBench
We create the eval code of UniGenBench for faster evaluation.
Data prepration
Please download the benchmark data from UniGenBench and place it in the Benchmark directory.
Evaluation
Directly run scripts/eval/run_unigenbench.sh to evaluate UniGenBench. The output will be saved in $output_path.
- Set
$model_pathand$output_pathfor the path for checkpoint and log. scripts/eval/run_unigenbench_refine.shsupport Fine-grained Editing-like Visual Refinement improve initial image quality from object presence, attribute accuracy, style consistency, and aesthetic quality
T2I-CoreBench
We create the eval code of T2I-CoreBench for faster evaluation.
Data prepration
Please download the benchmark data from T2I-CoreBench and and place it in the Benchmark directory.
Evaluation
Directly run scripts/eval/run_corebench.sh to evaluate T2I-CoreBench. The output will be saved in $output_path.
- Set
$model_pathand$output_pathfor the path for checkpoint and log. scripts/eval/run_corebench_refine.shsupport Fine-grained Editing-like Visual Refinement improve initial image quality from object presence, attribute accuracy, style consistency, and aesthetic quality
UniREditBench
We merge the eval code of UniREditBench for faster evaluation.
Data prepration
Please download the benchmark data from UniREditBench and and place it in the Benchmark directory.
Evaluation
Directly run scripts/eval/run_unireditbench.sh to evaluate UniREditBench. The output will be saved in $output_path.
- Set
$model_pathand$output_pathfor the path for checkpoint and log. - Use
thinkfor World Knowledge-Enhanced Textual Reasoning.