SSR2-GCD: Semi-Supervised Rate Reduction for Multimodal GCD
February 24, 2026 ยท View on GitHub
"Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery"
Abstract
State-of-the-art approaches for GCD task are usually built on multi-modality representation learning, which is heavily dependent upon inter-modality alignment. However, few of them cast a proper intra-modality alignment to generate a desired underlying structure of representation distributions. In this paper, we propose a novel and effective multi-modal representation learning framework for GCD via Semi-Supervised Rate Reduction, called SSR2-GCD, to learn cross-modality representations with desired structural properties based on emphasizing to properly align intra-modality relationships. Moreover, to boost knowledge transfer, we integrate prompt candidates by leveraging the inter-modal alignment offered by Vision Language Models. We conduct extensive experiments on generic and fine-grained benchmark datasets demonstrating superior performance of our approach.
Running
Dependencies
pip install -r requirements.txt
Scripts
Train the model:
For a quick start, as an example, run:
python prompt_search.py --dataset 'cifar10' --num_attributes 4 --num_tags 4
python main.py --dataset 'cifar10' --num_attributes 4 --num_tags 4 --base_lr 0.001
or:
python prompt_search.py --dataset 'flowers102' --num_attributes 4 --num_tags 4
python main.py --dataset 'cifar10' --num_attributes 4 --num_tags 4 --base_lr 0.001
Acknowledgements
We follow TextGCD (https://github.com/HaiyangZheng/TextGCD) to generate the candidate pool and search prompt candidates we acknowledge their great contribution for the multi-modal GCD community.
Remark
Additionally, we remove prompts corresponding to UNKNOWN categories in the {CUB, Flowers102, Oxford Pets, Stanford Cars} datasets from the candidate pool, as these prompts may lead to semantic leakage (see the deduplicated lexicon in "./Lexicon/Lexicon_tags_deduplicated.csv"). Specifically, prompts from 233 categories are removed (see the ImageNet_Matching_Report.csv).
License
This project is licensed under the MIT License - see the LICENSE file for details.