De Novo Peptide Binder Design, Docking, and In Silico Screening using DiffPepBuilder and DiffPepDock

October 19, 2025 · View on GitHub

This is the official repository for the DiffPepBuilder and DiffPepDock tools.

plot

For any questions, please open an issue or contact wangyuzhe_ccme@pku.edu.cn for more information.

News

  • [2025/10/14] Our research article for DiffPepDock has been published in Protein Science! Explore the full paper here. We have also comprehensively refactored the codebase to enhance usability and maintainability.
  • [2025/5/18] A Colab demo for DiffPepDock is now available. Feel free to give it a try!
  • [2025/5/14] We've extended our method to protein–peptide docking with a derivative tool, DiffPepDock. The initial implementation and model weights for the docking functionality have been publicly released.
  • [2024/9/12] Our research article for DiffPepBuilder has been published in JCIM! Dive into the details by checking out the full paper here or on arXiv.
  • [2024/9/11] We released the PepPC-F and PepPC datasets for DiffPepBuilder on Zenodo. The training protocol has also been released. Please refer to the Training section for more details.
  • [2024/7/22] The initial code, model weights, and a Colab demo for DiffPepBuilder are now available.

Quick Start

We provide a Google Colab notebook to facilitate the use of DiffPepBuilder. Please click the following link to open the notebook in Google Colab:

Open In Colab

Similarly, a Colab notebook demonstrating the functionality of DiffPepDock is available at:

Open In Colab

Installation

We recommend using a conda environment to install the required packages. Please clone this repository and navigate to the root directory:

git clone https://github.com/YuzheWangPKU/DiffPepBuilder.git
cd DiffPepBuilder

Then run the following commands to create a new conda environment and install the required packages:

conda env create -f environment.yml
conda activate diffpepbuilder

Before running de novo design protocols, please unzip the SSBLIB data in the SSbuilder directory:

cd SSbuilder
tar -xvf SSBLIB.tar.gz

The post-processing procedure requires PyRosetta to be installed. We recommend installing the pre-built wheel using the following commands:

wget https://west.rosettacommons.org/pyrosetta/release/release/PyRosetta4.MinSizeRel.python39.linux.wheel/pyrosetta-2024.39+release.59628fb-cp39-cp39-linux_x86_64.whl
pip install pyrosetta-2024.39+release.59628fb-cp39-cp39-linux_x86_64.whl

Please adjust the version or build to match your Python environment and system architecture.

De Novo Design

To de novo generate peptide binders for a given target protein, please first download the model weights into experiments/checkpoints/ from Zenodo. You can use the following command to download the model weights:

wget https://zenodo.org/records/12794439/files/diffpepbuilder_v1.pth
mv diffpepbuilder_v1.pth experiments/checkpoints/

We provide an example of the target ALK1 (Activin Receptor-like Kinase 1, PDB ID: 6SF1) to demonstrate the procedures of generating peptide binders. Binding hotspots or motifs of the target protein can be specified in JSON format as comma-separated residue IDs, as showcased in the example file examples/receptor_data/de_novo_cases.json. The following input formats are supported:

SyntaxDescriptionExample stringConceptual expansion
Single residueOne residue IDB40["B40"]
Comma-separated listMultiple discrete residuesB40, B58, B71["B40","B58","B71"]
Dash-connected rangeInclusive sequence within the same chain/prefixB58-60["B58","B59","B60"]
Mixed listCommas separate items; items may be singles or rangesB40, B58-59, B71-72["B40","B58","B59","B71","B72"]

Note: please remove any underscores (_) from the receptor PDB file name, as the script interprets _ as a delimiter.

To preprocess the receptor, run the experiments/process_receptor.py script:

python experiments/process_receptor.py --pdb_dir examples/receptor_data --write_dir data/receptor_data --receptor_info_path examples/receptor_data/de_novo_cases.json

This script will generate the receptor data in the data/receptor_data directory. To generate peptide binders for the target protein, please specify the root directory of DiffPepBuilder repository and then run the experiments/run_inference.py script (modify the nproc-per-node flag accordingly based on the number of GPUs available):

export BASE_PATH="your/path/to/DiffPepBuilder"
torchrun --nproc-per-node=8 experiments/run_inference.py data.val_csv_path=data/receptor_data/metadata_test.csv

The config file config/inference.yaml contains the hyperparameters for the inference process. Below is a brief explanation of the key hyperparameters:

ParameterDescriptionDefault Value
use_ddpIndicates whether Distributed Data Parallel (DDP) training is usedTrue
use_gpuSpecifies whether to use GPU for computationTrue
num_gpusNumber of GPUs to use for computation8
num_tNumber of denoising steps200
noise_scaleScaling factor for noise, analogous to sampling temperature1.0
samples_per_lengthNumber of peptide backbone samples per sequence length8
min_lengthMinimum sequence length to sample8
max_lengthMaximum sequence length to sample30
seq_temperatureSampling temperature of the residue types0.1
build_ss_bondIndicates whether to build disulfide bondsTrue
max_ss_bondMaximum number of disulfide (SS) bonds to build2

You can modify these hyperparameters to customize the inference process. For more details on the hyperparameters, please refer to our paper.

After running the inference script, the generated peptide binders will be saved in the runs/inference/. To run the side chain reconstruction and optimization, please run the following script subsequently:

export BASE_PATH="your/path/to/DiffPepBuilder"
python experiments/run_postprocess.py --in_pdbs runs/inference --ori_pdbs examples/receptor_data --amber_relax --rosetta_relax

The script will generate the final peptide binders and calculate the binding ddG values of the generated peptide binders. The results will be summarized in the runs/inference/postprocess_results.csv file.

Docking

To run peptide docking protocols for a given target protein, please first download the model weights into experiments/checkpoints/ from Zenodo. You can use the following command to download the model weights:

wget https://zenodo.org/records/15398020/files/diffpepdock_v1.pth
mv diffpepdock_v1.pth experiments/checkpoints/

DiffPepDock provides user-friendly, automated scripts for docking batches of peptide sequences to a specific target protein. Here we provide an example of the redocking task of the substrate-binding protein YejA in complex with its native peptide fragment (PDB ID: 7Z6F) to demonstrate the procedures of docking process. You may modify the file examples/docking_data/peptide_seq.fasta include custom peptide sequences for docking. Prior binding information including reference ligands and binding motifs can be specified in JSON format, as demonstrated in examples/docking_data/docking_cases.json, using the same syntax as in de novo design.

To preprocess the target and the peptide sequences, run the experiments/process_batch_dock.py script:

python experiments/process_batch_dock.py --pdb_dir examples/docking_data --write_dir data/docking_data --receptor_info_path examples/docking_data/docking_cases.json --peptide_seq_path examples/docking_data/peptide_seq.fasta

The preprocessed data will be placed in the data/docking_data directory. Redocking of existing protein–peptide complexes can be performed by omitting the peptide_seq_path argument and specify the peptide ligand chain in the docking_cases.json file.

To perform protein-peptide docking, please specify the root directory of DiffPepBuilder repository and then run the experiments/run_docking.py script (please modify the nproc-per-node flag accordingly based on the number of GPUs available):

export BASE_PATH="your/path/to/DiffPepBuilder"
torchrun --nproc-per-node=8 experiments/run_docking.py data.val_csv_path=data/docking_data/metadata_test.csv

The config file config/docking.yaml contains the hyperparameters for both the docking and postprocessing procedures. You can modify these hyperparameters to customize the overall workflow. Upon completion of the inference script, the generated protein-peptide complexes will be saved in the runs/docking/ directory.

To improve efficiency, the automatic postprocessing step can be skipped by setting postprocess.run_postprocess to False in config/docking.yaml. Postprocessing can then be executed separately with improved parallelization using the following command:

export BASE_PATH="your/path/to/DiffPepBuilder"
python experiments/run_postprocess.py --in_pdbs runs/docking --ori_pdbs examples/docking_data --amber_relax --rosetta_relax

This script generates the final protein-peptide complexes and computes the binding ddG values as a pose scoring metric. The results are summarized in the runs/docking/postprocess_results.csv file. Optionally, you can add the --save_best flag to record only the top-ranked poses based on binding ddG in the summary file.

Training

To train the DiffPepBuilder model from scratch, please download the training data from Zenodo and unzip the data in the data/ directory:

wget https://zenodo.org/records/13744959/files/PepPC-F_raw_data.tar.gz
mkdir data/PepPC-F_raw_data
tar -xvf PepPC-F_raw_data.tar.gz --strip-components=1 -C data/PepPC-F_raw_data

To preprocess the training data, run the experiments/process_dataset.py script:

python experiments/process_dataset.py --pdb_dir data/PepPC-F_raw_data --write_dir data/complex_dataset

This script will generate the training data in the data/complex_dataset directory. You can add max_batch_size flag to specify the maximum batch size for ESM embedding to avoid out-of-memory errors. Then split the data into training and validation sets:

python experiments/split_dataset.py --input_path data/complex_dataset/metadata.csv --output_path data/complex_dataset --num_val 200

You can modify the num_val flag to specify the number of validation samples. To train the DiffPepBuilder model, please specify the root directory of the DiffPepBuilder repository and then run the experiments/train.py script (modify the nproc-per-node flag accordingly based on the number of GPUs available):

export BASE_PATH="your/path/to/DiffPepBuilder"
torchrun --nproc-per-node=8 experiments/train.py

The config file config/base.yaml contains the hyperparameters for the training process. You can modify these hyperparameters to customize the training process. Checkpoints will be saved every 10,000 steps after validation in the runs/ckpt/ directory by default. Training logs will be saved every 2,500 steps.

To run subsequent finetuning for docking tasks, please prepare the PepPC dataset, update the config file config/finetune.yaml accordingly, and run the following command:

export BASE_PATH="your/path/to/DiffPepBuilder"
torchrun --nproc-per-node=8 experiments/train.py --config-name=finetune

License

This project is licensed under the MIT License - see the LICENSE file for details.

Citation

Please cite the following paper(s) if you use this code in your research:

@article{wang2024target,
  title={Target-Specific De Novo Peptide Binder Design with DiffPepBuilder},
  author={Wang, Fanhao and Wang, Yuzhe and Feng, Laiyi and Zhang, Changsheng and Lai, Luhua},
  journal={Journal of Chemical Information and Modeling},
  volume={64},
  number={24},
  pages={9135-9149},
  year={2024},
  publisher={ACS Publications},
  doi = {10.1021/acs.jcim.4c00975}
}
@article{wang2025diffpepdock,
  title={DiffPepDock: Efficient Protein--Peptide Docking and Binder Screening via SE (3)-Equivariant Diffusion},
  author={Wang, Yuzhe and Wang, Fanhao and Feng, Laiyi and Zhang, Changsheng and Lai, Luhua},
  journal={Protein Science},
  volume={34},
  number={11},
  pages={e70338},
  year={2025},
  publisher={Wiley Online Library}
}

Acknowledgments

We would like to thank the authors of FrameDiff and OpenFold, whose codebases we used as references for our implementation.