PerceptionGPT: Effectively Fusing Visual Perception into LLM
August 7, 2024 · View on GitHub
This repository contains the code for the paper titled "PerceptionGPT: Effectively Fusing Visual Perception into LLM". [Link to our paper]
Install Packages
conda create -n perceptiongpt python=3.10
conda activate perceptiongpt
pip install -r requirements.txt
Model training with LoRA
Before conducting training, organize the training dataset following Shikra.
The running config located in config/train_configs/shikra3_rec3_mask_box_cls_refcoco_all.py could be modified accordingly
bash script/run.sh
Acknowledgement
The project is built on top of the amazing multimodal large language model LLaVA and Shikra. Thanks for these great work!
If you find our work useful for your research or applications, please cite using this BibTeX:
@misc{pi2023perceptiongpt,
title={PerceptionGPT: Effectively Fusing Visual Perception into LLM},
author={Renjie Pi and Lewei Yao and Jiahui Gao and Jipeng Zhang and Tong Zhang},
year={2023},
eprint={2311.06612},
archivePrefix={arXiv},
primaryClass={cs.CV}
}