Awesome-MLLM-Tuning

July 13, 2025 ยท View on GitHub

MLLM-Tuning

Awesome-MLLM-Tuning

Curated list of Multimodal Large Language Model (MLLM) Tuning resources, aligned with our work:
Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model

arXiv Badge Custom Badge License Badge GitHub stars

๐Ÿ™Œ Abstract

Multi-modal Large Language Models (MLLMs) integrate visual and linguistic reasoning to address complex tasks such as image captioning and visual question answering. While MLLMs demonstrate remarkable versatility, MLLMs appears limited performance on special application. But tuning MLLMs for downstream tasks encounters two key challenges: Task-Expert Specialization, where distribution shifts between pre-training and target datasets constrain target performance, and Open-World Stabilization, where catastrophic forgetting erases the model general knowledge. In this work, we systematically review recent advancements in MLLM tuning methodologies, classifying them into three paradigms: (I) Selective Tuning, (II) Additive Tuning, and (III) Reparameterization Tuning. Furthermore, we benchmark these tuning strategies across popular MLLM architectures and diverse downstream tasks to establish standardized evaluation analysis and systematic tuning principles. Finally, we highlight several open challenges in this domain and propose future research directions.

๐Ÿ“– Paper

Selective Tuning

Iterative Selective Tuning

TimeTitleVenuePaperCode
2024.10AlphaEdit: Null-Space Constrained Knowledge Editing for Language ModelsICLR'24linklink
2023.12Sparse is Enough in Fine-tuning Pre-trained Large Language ModelsICML'24linklink
2023.12Gradient-based Parameter Selection for Efficient Fine-TuningCVPR'24linklink
2023.11Unified Low-Resource Sequence Labeling by Sample-Aware Dynamic Sparse FinetuningEMNLP'23linklink
2023.08Overcoming Generic Knowledge Loss with Selective Parameter UpdateCVPR'24linklink
2023.06LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse ApproximationICML'23linklink
2022.10ROSE: Robust Selective Fine-tuning for Pre-trained Language ModelsIJCAI'22linklink
2022.05Parameter-Efficient Sparsity for Large Language Models Fine-TuningIJCAI'22linklink
2021.10Composable Sparse Fine-Tuning for Cross-Lingual TransferACL'22linklink
2021.09Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuningEMNLP'21linklink
2015.06Learning both Weights and Connections for Efficient Neural NetworksNeurIPS'15link-

Posterior Selective Tuning

TimeTitleVenuePaperCode
2024.12Revisiting Weight Averaging for Model MergingarXiv'24linklink
2024.10Parameter Competition Balancing for Model MergingNeurIPS'24linklink
2024.06Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingNeurIPS'24linklink
2024.05TEMR-Merging: Tuning-Free High-Performance Model MergingNeurIPS'24linklink
2024.02Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language ModelsICML'24linklink
2023.11Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchICML'24linklink
2023.06TIES-Merging: Resolving Interference When Merging ModelsNeurIPS'23linklink
2021.11Merging Models with Fisher-Weighted AveragingNeurIPS'22linklink

Additive Tuning

Adapter Tuning

TimeTitleVenuePaperCode
2024.04Conditional Prototype Rectification Prompt LearningTCSVT'25linklink
2023.11Meta-Adapter: An Online Few-shot Learner for Vision-Language ModelNeurIPS'23linklink
2023.09GraphAdapter: Tuning Vision-Language Models With Dual Knowledge GraphNeurIPS'23linklink
2023.04Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior RefinementICCV'23linklink
2023.03Revisiting Multimodal Representation in Contrastive Learning: From Patch and Token Embeddings to Finite Discrete TokensCVPR'23linklink
2023.02Side Adapter Network for Open-Vocabulary Semantic SegmentationCVPR'23linklink
2022.11Task Residual for Tuning Vision-Language ModelsCVPR'23linklink
2022.06LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer LearningNeurIPS'22linklink
2021.11Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language ModelingECCV'22linklink
2021.10CLIP-Adapter: Better Vision-Language Models with Feature AdaptersIJCV'23linklink
2019.02Parameter-Efficient Transfer Learning for NLPICML'19linklink

Prompt Tuning

TimeTitleVenuePaperCode
2024.03PromptKD: Unsupervised Prompt Distillation for Vision-Language ModelsCVPR'24linklink
2024.03Domain-agnostic mutual prompting for unsupervised domain adaptationCVPR'24linklink
2024.01Learning to prompt with text only supervision for vision-language modelsAAAI'25linklink
2023.11ArGue: Attribute-guided prompt tuning for vision-language modelsCVPR'24linklink
2023.09Dept: Decoupled prompt tuningCVPR'24linklink
2023.09Distribution-aware prompt tuning for vision-language modelsICCV'23linklink
2023.08Knowledge-aware prompt tuning for generalizable vision-language modelsICCV'23link-
2023.07Self-regulating prompts: Foundational model adaptation without forgettingICCV'23linklink
2023.03Visual-language prompt tuning with knowledge-guided context optimizationCVPR'23linklink
2022.10MaPLe: Multi-modal prompt learningCVPR'23linklink
2022.10Prompt learning with optimal transport for vision-language modelsICLR'23linklink
2022.06Dualcoop: Fast adaptation to multi-label recognition with limited annotationsNeurIPS'22linklink
2022.05Prompt-aligned Gradient for Prompt TuningICCV'23linklink
2022.03Visual prompt tuningECCV'22linklink
2022.03Conditional prompt learning for vision-language modelsCVPR'22linklink
2021.09Learning to prompt for vision-language modelsIJCV'22linklink
2021.01Prefix-Tuning: Optimizing Continuous Prompts for GenerationACL'21linklink
2020.10AUTOPROMPT: Eliciting Knowledge from Language Models with Automatically Generated PromptsEMNLP'20linklink
2019.11How Can We Know What Language Models Know?TACL'20linklink

Reparameterization Tuning

Structure Reparameterization Tuning

TimeTitleVenuePaperCode
2025.02REMEDY: Recipe merging dynamics in large vision-language modelsICLR'25link-
2024.12Lora.rar: Learning to merge loras via hypernetworks for subject-style conditioned image generationICCV'25linklink
2024.08Teamlora: Boosting low-rank adaptation with expert collaboration and competitionarXiv'24linklink
2024.06Twinmerging: Dynamic integration of modular expertise in model mergingNeurIPS'24linklink
2024.06Mixture-of-subspaces in low-rank adaptationEMNLP'24linklink
2024.06Sharelora: Parameter efficient and robust large language model fine-tuning via shared low-rank adaptationarXiv'24linklink
2024.05Parameter-Efficient Fine-Tuning with Discrete Fourier TransformICML'24linklink
2024.03Mtlora: Low-rank adaptation approach for efficient multi-task learningCVPR'24linklink
2024.02Multimodal instruction tuning with conditional mixture of loraACL'24linklink
2023.12Loramoe: Alleviating world knowledge forgetting in large language models via moe-style pluginACL'24linklink
2023.10Vera: Vectorbased random matrix adaptationICLR'24linklink
2023.07Lorahub: Efficient cross-task generalization via dynamic lora compositionCOLM'24linklink

Calibration Reparameterization Tuning

TimeTitleVenuePaperCode
2025.03Lorasculpt: Sculpting lora for harmonizing general and specialized knowledge in multimodal large language modelsCVPR'25linklink
2024.10Controlled low-rank adaptation with subspace regularization for continued training on large language modelsarXiv'24link-
2024.07Learn to preserve and diversify: Parameter-efficient group with orthogonal regularization for domain generalizationECCV'24linklink
2024.06Corda: Context-oriented decomposition adaptation of large language models for task-aware parameter-efficient fine-tuningNeurIPS'24linklink
2024.06Milora: Harnessing minor singular components for parameter-efficient llm finetuningNAACL'25linklink
2024.04Pissa: Principal singular values and singular vectors adaptation of large language modelsNeurIPS'24linklink
2024.03Lora meets dropout under a unified frameworkACL'24link-
2024.03Bilora: A bi-level optimization framework for overfitting-resilient low-rank adaptation of large pre-trained modelsarXiv'24link-
2024.02Melora: mini-ensemble low-rank adapters for parameter-efficient fine-tuningACL'24linklink
2024.02Prolora: Partial rotation empowers more parameter-efficient loraACL'24linklink
2024.02Lora+: Efficient low rank adaptation of large modelsICML'24linklink
2024.02Dora: Weight-decomposed low-rank adaptationICML'24linklink
2024.02Flora: Low-rank adapters are secretly gradient compressorsICML'24linklink
2023.08Bayesian low-rank adaptation for large language modelsICLR'24linklink

๐Ÿ‘‹ Contact

This repository is currently maintained by Wenke Huang ๐Ÿ‘จโ€๐Ÿ’ป.
If you have any questions, concerns, or suggestions regarding the contents of this repository or the resources shared here, feel free to reach out! I'm more than happy to assist you with any inquiries or help you navigate through the materials.
Please don't hesitate to send an email to me at wenkehuang@whu.edu.cn ๐Ÿ“ง or Wechat ๐Ÿค—.

๐Ÿฅณ Citation

If you find this repository helpful for your research, we would greatly appreciate it if you could cite our papers. โœจ

@misc{MLLMTuning_arXiv25,
      title={Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model}, 
      author={Wenke Huang, Jian Liang, Xianda Guo, Yiyang Fang, Guancheng Wan, Xuankun Rong, Chi Wen, Zekun Shi,  Qingyun Li, Didi Zhu, Yanbiao Ma, Ke Liang, Bin Yang, He Li, Jiawei Shao, Mang Ye, Bo Du},
      year={2025},
      eprint={2503.04543},
      archivePrefix={arXiv},
      primaryClass={cs.CR}
}

@inproceedings{LiangLoRASculpt_CVPR2025,
    author    = {Liang, Jian and Huang, Wenke and Wan, Guancheng and Yang, Qu and Ye, Mang},
    title     = {LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models},
    booktitle = {CVPR},
    year      = {2025},
}

@inproceedings{FangSEPM_ICML2025,
  title     = {Catch Your Emotion: Sharpening Emotion Perception in Multimodal Large Language Models},
  author    = {Fang, Yiyang and Liang, Jian and Huang, Wenke and Li, He and Su, Kehua and Ye, Mang},
  booktitle = {ICML},
  year      = {2025},
}

@misc{ye2025surveysafetylargevisionlanguage,
      title={A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations}, 
      author={Mang Ye and Xuankun Rong and Wenke Huang and Bo Du and Nenghai Yu and Dacheng Tao},
      year={2025},
      eprint={2502.14881},
      archivePrefix={arXiv},
      primaryClass={cs.CR}
}

๐Ÿ” Relevant Projects

[1] LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models - CVPR 2025 [Link][Code]

[2] Catch Your Emotion: Sharpening Emotion Perception in Multimodal Large Language Models - ICML 2025 [Link][Code]

[3] A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations - arXiv 2025 [Link][Code]

You Only Live Once.

I hope that all players have fun.