SlideChat: A Multimodal Generative AI Assistant for Whole-Slide Computational Pathology

September 2, 2026 · View on GitHub

Note

This repository provides the official code for SlideChat. The expanded study, published in Nature Cancer, substantially extends our original CVPR 2025 work with a larger instruction dataset, broader expert-reviewed evaluation, and additional pathology tasks across cancer types.

Release

We release SlideChat, SlideInstruction, and SlideBench as open-source resources, hoping to facilitate research and development in computational pathology.

  • SlideChat: The first large vision-language assistant for whole-slide pathology image analysis, capable of generating comprehensive descriptions and contextually relevant responses.
  • SlideInstruction: The largest comprehensive WSI instruction-following dataset, derived from pathology reports..
  • SlideBench: A WSI multimodal benchmark including SlideBench-Caption/Report (TCGA, CPTAC, HISTAI) and SlideBench-Closed(VQA-TCGA, VQA-BCNB, VQA-CPTAC, VQA-HISTAI). Before open‑sourcing, SlideBench‑VQA‑TCGA underwent a second round of expert review with pathologists, further enhancing data quality. The initial version covers 10 cancer types with 1,494 samples. We later expanded it to 31 additional cancer types, rigorously validated by experts, yielding 3,176 samples (SlideBench‑VQA‑TCGA.csv). SlideChat achieved an accuracy of 75.23% on the initial version and 74.12% on the expanded version.

The results below correspond to the expanded journal version of SlideChat.

Closed-Ended Question Answering

DatasetTaskSlideChatGPT-4oMedDrQuilt-LLaVALLaVA-Med
TCGAHistopathological Changes0.859 ± 0.0190.481 ± 0.0270.533 ± 0.0270.311 ± 0.0250.292 ± 0.025
TCGACytomorphological Characteristics0.843 ± 0.0400.584 ± 0.0520.690 ± 0.0500.180 ± 0.0420.216 ± 0.045
TCGATissue Architecture0.836 ± 0.0190.614 ± 0.0260.601 ± 0.0250.385 ± 0.0260.465 ± 0.026
TCGATumor Characteristics0.717 ± 0.0350.607 ± 0.0390.612 ± 0.0390.259 ± 0.0340.306 ± 0.035
TCGADisease Classification0.796 ± 0.0150.558 ± 0.0170.489 ± 0.0180.271 ± 0.0150.299 ± 0.016
TCGADisease Detection0.765 ± 0.0670.371 ± 0.0760.746 ± 0.0660.444 ± 0.0730.372 ± 0.073
TCGADifferential Diagnosis0.749 ± 0.0390.532 ± 0.0440.547 ± 0.0440.260 ± 0.0400.317 ± 0.043
TCGAStaging0.656 ± 0.0220.606 ± 0.0220.309 ± 0.0220.185 ± 0.0180.253 ± 0.020
TCGAGrading0.604 ± 0.0200.494 ± 0.0200.406 ± 0.0200.140 ± 0.0140.171 ± 0.016
TCGATreatment Guidance0.774 ± 0.0510.751 ± 0.0540.761 ± 0.0530.397 ± 0.0580.514 ± 0.061
TCGABiomarker Analysis0.738 ± 0.0830.591 ± 0.0950.594 ± 0.0940.588 ± 0.0990.589 ± 0.095
TCGARisk Factors0.713 ± 0.0840.678 ± 0.0870.640 ± 0.0920.504 ± 0.0970.713 ± 0.086
TCGAPrognostic Assessment0.695 ± 0.0750.614 ± 0.0760.587 ± 0.0780.230 ± 0.0670.362 ± 0.078
TCGAOverall0.741 ± 0.0080.557 ± 0.0090.490 ± 0.0090.256 ± 0.0080.298 ± 0.008
BCNBTumor Subtype Classification0.894 ± 0.0090.368 ± 0.0150.387 ± 0.0150.560 ± 0.0150.437 ± 0.015
BCNBER Status Classification0.791 ± 0.0130.819 ± 0.0120.309 ± 0.0140.480 ± 0.0150.481 ± 0.015
BCNBPR Status Classification0.703 ± 0.0140.533 ± 0.0150.230 ± 0.0130.445 ± 0.0150.438 ± 0.015
BCNBHER2 Status Classification0.768 ± 0.0130.687 ± 0.0140.625 ± 0.0150.295 ± 0.0150.287 ± 0.015
BCNBOverall0.789 ± 0.0060.602 ± 0.0070.388 ± 0.0070.445 ± 0.0080.411 ± 0.008
CPTACCM Subtype0.487 ± 0.0670.218 ± 0.0530.418 ± 0.0630.181 ± 0.0510.331 ± 0.062
CPTACLSCC Subtype0.681 ± 0.0600.500 ± 0.0660.414 ± 0.0630.085 ± 0.0360.068 ± 0.033
CPTACLUAD Subtype0.418 ± 0.0660.370 ± 0.0640.236 ± 0.0550.215 ± 0.0530.134 ± 0.045
CPTACUCEC Subtype0.482 ± 0.0670.234 ± 0.0540.364 ± 0.0630.200 ± 0.0510.200 ± 0.051
CPTACOverall0.517 ± 0.0330.330 ± 0.0300.358 ± 0.0310.170 ± 0.0240.183 ± 0.025
HISTAIBreast Disease Classification0.795 ± 0.0200.484 ± 0.0240.444 ± 0.0240.336 ± 0.0230.314 ± 0.022
HISTAISkin Disease Classification0.736 ± 0.0310.449 ± 0.0360.272 ± 0.0310.243 ± 0.0310.116 ± 0.023
HISTAIColorectum Disease Classification0.695 ± 0.0360.634 ± 0.0380.570 ± 0.0400.506 ± 0.0400.434 ± 0.041
HISTAIOverall0.761 ± 0.0160.505 ± 0.0180.387 ± 0.0180.354 ± 0.0170.320 ± 0.017

Report Generation

METEOR scores on the SlideBench-Report:

DatasetCancer TypeSlideChatHistoGPTPRISM
TCGAACC0.153 ± 0.0500.079 ± 0.0090.007 ± 0.000
TCGABLCA0.254 ± 0.0040.095 ± 0.0030.028 ± 0.001
TCGABRCA0.250 ± 0.0030.096 ± 0.0020.024 ± 0.001
TCGACESC0.147 ± 0.0040.104 ± 0.0040.024 ± 0.002
TCGACHOL0.246 ± 0.0170.111 ± 0.0110.021 ± 0.003
TCGACOAD0.258 ± 0.0050.099 ± 0.0040.032 ± 0.002
TCGADLBC0.109 ± 0.0140.082 ± 0.0080.031 ± 0.006
TCGAESCA0.143 ± 0.0060.118 ± 0.0070.039 ± 0.004
TCGAGBM0.235 ± 0.0100.093 ± 0.0070.012 ± 0.001
TCGAHNSC0.224 ± 0.0050.097 ± 0.0030.039 ± 0.002
TCGAKICH0.187 ± 0.0070.084 ± 0.0040.013 ± 0.001
TCGAKIRC0.200 ± 0.0040.089 ± 0.0030.025 ± 0.001
TCGAKIRP0.202 ± 0.0050.088 ± 0.0040.017 ± 0.001
TCGALGG0.231 ± 0.0050.088 ± 0.0030.015 ± 0.001
TCGALIHC0.198 ± 0.0070.108 ± 0.0030.023 ± 0.001
TCGALUAD0.245 ± 0.0050.100 ± 0.0030.020 ± 0.001
TCGALUSC0.269 ± 0.0050.106 ± 0.0030.033 ± 0.002
TCGAMESO0.163 ± 0.0070.098 ± 0.0090.016 ± 0.002
TCGAOV0.121 ± 0.0110.087 ± 0.0080.020 ± 0.003
TCGAPAAD0.143 ± 0.0070.090 ± 0.0050.029 ± 0.003
TCGAPCPG0.198 ± 0.0080.116 ± 0.0050.015 ± 0.001
TCGAPRAD0.196 ± 0.0050.093 ± 0.0030.026 ± 0.001
TCGAREAD0.265 ± 0.0070.106 ± 0.0060.027 ± 0.002
TCGASARC0.167 ± 0.0090.088 ± 0.0050.011 ± 0.001
TCGASKCM0.200 ± 0.0140.113 ± 0.0100.016 ± 0.002
TCGASTAD0.151 ± 0.0060.096 ± 0.0030.032 ± 0.002
TCGATGCT0.163 ± 0.0120.115 ± 0.0060.014 ± 0.003
TCGATHCA0.191 ± 0.0040.101 ± 0.0030.025 ± 0.001
TCGATHYM0.241 ± 0.0070.109 ± 0.0050.013 ± 0.002
TCGAUCEC0.198 ± 0.0060.092 ± 0.0020.022 ± 0.001
TCGAUCS0.206 ± 0.0080.075 ± 0.0060.012 ± 0.001
TCGAOverall0.214 ± 0.0010.097 ± 0.0120.025 ± 0.000
CPTACCCRCC0.170 ± 0.0070.114 ± 0.0090.096 ± 0.006
CPTACLSCC0.151 ± 0.0130.156 ± 0.0180.061 ± 0.013
CPTACLUAD0.114 ± 0.0080.144 ± 0.0130.050 ± 0.008
CPTACUCEC0.118 ± 0.0040.140 ± 0.0110.024 ± 0.008
CPTACOverall0.140 ± 0.0050.139 ± 0.0200.059 ± 0.005
HISTAIBreast0.119 ± 0.0020.058 ± 0.0010.012 ± 0.001
HISTAISkin0.084 ± 0.0020.079 ± 0.0020.022 ± 0.001
HISTAIColorectum0.107 ± 0.0020.063 ± 0.0010.022 ± 0.001
HISTAIOverall0.106 ± 0.0010.065 ± 0.0010.018 ± 0.001

Installation

Environment Setup

This project is built upon Xtuner. To get started:

conda create --name xtuner-env python=3.10 -y
conda activate xtuner-env
git clone https://github.com/uni-medical/SlideChat.git
cd SlideChat
pip install -e .

Dependencies

Please refer to requirements.txt and environment.yaml for the complete list of dependencies.

Pre-requisites

Download the JSON file containing WSI IDs (TCGA) and conversation data from the Dataset. The input image file is in CSV format and contains 512-dimensional feature representations for all patches within the WSI. Example files are provided in the ./dataset/ folder. For slide downloading and processing, please refer to CLAM and DSMIL.

Training

SlideChat serializes each input WSI into a sequence of patches, converting each into visual embeddings with a patch-level encoder CONCH. A slide-level encoder then interacts with these features to generate contextual embeddings. Then, a multimodal projector maps the visual features from the slide-level encoder into a unified space, aligned seamlessly with the LLM. SlideChat was trained for two stages: (1) Cross-Domain Alignment: SlideChat is trained to generate descriptive captions using WSI-caption pairs from SlideInstruction. Specifically, only the slide-level encoder and projection are updated, while the patch-level encoder and LLM weights remain fixed; (2) Visual Instruction Learning: we utilize WSI VQAs from SlideInstruction, allowing the slide encoder, projection layer, and large language model components to be fully trainable to ensure comprehensive adaptability.


Config files are in configs/.

NPROC_PER_NODE=${GPU_NUM} xtuner train \
 <your config file path>  \
  --deepspeed <deepspeed config file path> \
  --work-dir <workdir path>

# stage1 example
NPROC_PER_NODE=${GPU_NUM} xtuner train \
 configs/slidechat/stage_1.py \
  --deepspeed configs/deepspeed/deepspeed_zero2.json \
  --work-dir work_dirs/stage1

# stage2 example
NPROC_PER_NODE=${GPU_NUM} xtuner train \
 configs/slidechat/stage_2.py \
  --deepspeed configs/deepspeed/deepspeed_zero2.json \
  --work-dir work_dirs/stage2

Where ${GPU_NUM} is the number of GPUs

For a detailed explanation of the configuration file, please refer here.

  • llm_name_or_path: The parameter llm_name_or_path corresponds to the Hugging Face LLM path, such as internlm/internlm2-chat-7b or Qwen/Qwen2.5-7B-Instruct and so on.
  • data_path: Training data (.json) path.
  • evaluation_images: Evaluation data path.

LLAVAModel Hyperparameters:

  • freeze_llm: Freeze the parameters of the LLM.
  • pretrained_pth: If it is the stage 2 training , it refers to the checkpoint file from stage 1 training; otherwise, it is set to None.
  • train_stage: train_stage indicates the training phase, either Stage '1' or Stage '2'.

Inference

xtuner test <your config file path> \
--checkpoint <your checkpoint path> \
--test_slide_csv  <your test file> \
--test_output_csv <result file> \
--local_rank 0

# example
xtuner test configs/slidechat/stage_2.py \
--checkpoint stage2_pth \
--test_slide_csv  SlideBench-VQA(TCGA).csv \
--test_output_csv output_my_test.csv \
--local_rank 0

Computational Resources

Training

We trained SlideChat using 8 x NVIDIA A100 (80GB) GPUs.

ParameterStage-1Stage-2
Training GPUs8 x A100 (80 GB)8 x A100 (80 GB)
Total Training Time~3 hours~24 hours
Batch Size11
Epochs31

Inference

SlideChat is deployment-friendly. NVIDIA A100 (80GB) is recommended for processing large slides, while a single NVIDIA RTX 4090 (24GB) is sufficient for WSIs with < 20,480 patches.

Contact

Citation

BibTeX:

@inproceedings{chen2025slidechat,
  title={Slidechat: A large vision-language assistant for whole-slide pathology image understanding},
  author={Chen, Ying and Wang, Guoan and Ji, Yuanfeng and Li, Yanjun and Ye, Jin and Li, Tianbin and Hu, Ming and Yu, Rongshan and Qiao, Yu and He, Junjun},
  booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  pages={5134--5143},
  year={2025},
  organization={IEEE}
}