⭐️ Star History

June 30, 2026 · View on GitHub

Awesome PR's Welcome

Multimodal Referring Segmentation: A Survey

Henghui Ding · Song Tang · Shuting He · Chang Liu · Zuxuan Wu · Yu-Gang Jiang

arXiv PDF

An Illustration of Multimodal Referring Segmentation Tasks.

An Illustration of Multimodal Referring Segmentation Representative Works.

🔥 Add Your Paper in our Repo and Survey!

  • We welcome contributions to enhance the comprehensiveness of this survey. If you identify any omitted works or wish to suggest additional papers, implementations, or resources, please submit a pull request. All relevant contributions will be promptly reviewed and incorporated.

  • Note: While we strive for comprehensive coverage, the vast number of papers on arXiv makes it impractical to include every work in our survey. Nevertheless, we encourage researchers to submit pull requests with their work for potential inclusion in future survey versions.

🔥 New

🔥 Highlight!!

Table of contents

  1. Referring Expression Segmentation (RES)
  2. Referring Video-Object Segmentation (RVOS)
  3. Referring Audio-Visual Segmentation (RAVS)
  4. 3D Referring Expression Segmentation (3D-RES)
  5. Generalized Referring Expression x (GREx)
  6. Application

1. Referring Expression Segmentation (RES)

TitleSourceCode / Homepage
[ORES]Refer to Any Segmentation Mask Group With Vision-Language PromptsICCV 2025Homepage
[Falcon Perception]Falcon PerceptionarXiv 2026Code
[DR2^2Seg]DR2^2Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language ModelsarXiv 2026
[CroBIM-U]CroBIM-U: Uncertainty-Driven Referring Remote Sensing Image SegmentationarXiv 2026
[IBISAgent]IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for UniversalarXiv 2026
[EVOL-SAM3]Evolving, Not Training: Zero-Shot Reasoning Segmentation via Evolutionary PromptingarXiv 2025Code
[SSA]Spatial-aware Symmetric Alignment for Text-guided Medical Image SegmentationarXiv 2025
[Think2Seg-RS]Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Remote SensingarXiv 2025Code
[OmniRIS]Omni-Referring Image SegmentationarXiv 2025Code
[UniGeoSeg]UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial ScenesarXiv 2025Code
[SaFiRe]SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image SegmentationNeurIPS 2025Homepage
[UniPixel]UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual ReasoningNeurIPS 2025Code
[LRB-WREL]Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce ExpressionsarXiv 2025
[CSINet]Referring Remote Sensing Image Segmentation with Cross-view Semantics Interaction NetworkarXiv 2025
[DGL-RSIS]DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image SegmentationarXiv 2025
[RIS-FUSION]RIS-FUSION: Rethinking Text-Driven Infrared and Visible Image Fusion from the Perspective of Referring Image SegmentationarXiv 2025
[SVP]Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image SegmentationarXiv 2025
[UGround]UGround: Towards Unified Visual Grounding with Unrolled TransformersarXiv 2025Code
[CoT Referring]CoT Referring: Improving Referring Expression Tasks with Grounded ReasoningarXiv 2025
[MedSeg EarlyFusion]A Text-Image Fusion Method with Data Augmentation Capabilities for Referring Medical Image SegmentationarXiv 2025Code
[PixelRefer]PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary GranularityarXiv 2025Code
[RefAM]RefAM: Attention Magnets for Zero-Shot Referral SegmentationarXiv 2025Homepage
[CoPatch]CoPatch: Zero-Shot Referring Image Segmentation by Leveraging Untapped Spatial Knowledge in CLIParXiv 2025
[Latent-VG]Latent Expression Generation for Referring Image Segmentation and GroundingICCV 2025
[RIS-LAD]RIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image SegmentationarXiv 2025Code
[X-SAM]X-SAM: From Segment Anything to Any SegmentationarXiv 2025Code
[Seg-R1]Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement LearningarXiv 2025
[ELBO-T2IAlign]ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-level Text-Image Alignment in Diffusion ModelsarXiv 2025
[MBA]MBA: Multimodal Bidirectional Attack for Referring Expression Segmentation ModelsarXiv 2025
[FOCUS]FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential SegmentationarXiv 2025
[NWPU-Refer]A Large-Scale Referring Remote Sensing Image Segmentation Dataset and BenchmarkarXiv 2025Code
[SegVLM]Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language ModelarXiv 2025
[RemoteSAM] RemoteSAM: Towards Segment Anything for Earth ObservationarXiv 2025Code
[LISAT] LISAT: Language-Instructed Segmentation Assistant for Satellite ImageryarXiv 2025Homepage
[PRS-Med] PRS-Med: Position Reasoning Segmentation with Vision-Language Model in Medical ImagingarXiv 2025
[RVTBench] RVTBench: A Benchmark for Visual Reasoning TasksarXiv 2025Code
[SAM-R1] SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement LearningarXiv 2025
[VisionReasoner] VisionReasoner: Unified Visual Perception and Reasoning via Reinforcement LearningarXiv 2025Code
[ReasoningSeg Survey] Reasoning Segmentation for Images and Videos: A SurveyarXiv 2025
[PixelThink] PixelThink: Towards Efficient Chain-of-Pixel ReasoningarXiv 2025
[SynRES] SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic DataarXiv 2025Code
[RESAnything]RESAnything: Attribute Prompting for Arbitrary Referring SegmentationNeurIPS 2025Homepage
[Pixel-SAIL] Pixel-SAIL: Single Transformer For Pixel-Grounded UnderstandingarXiv 2025Code
[LVLM_CSP] LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning SegmentationarXiv 2025
[SegEarth-R1] SegEarth-R1: Geospatial Pixel Reasoning via Large Language ModelarXiv 2025Code
[UniRES++] Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target GranularitiesarXiv 2025Code
[MediSee] MediSee: Reasoning-based Pixel-level Perception in Medical ImagesarXiv 2025Code
[PLVL] Progressive Language-guided Visual Learning for Multi-Task Visual GroundingarXiv 2025Code
[UFO] UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language InterfacearXiv 2025Code
[GroundingSuite] GroundingSuite: Measuring Complex Multi-Granular Pixel GroundingarXiv 2025Code
[UniVG] UniVG: A Generalist Diffusion Model for Unified Image Generation and EditingarXiv 2025
[AURA] Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURAarXiv 2025
[Seg-Zero] Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive ReinforcementarXiv 2025Code
[AeroReformer] AeroReformer: Aerial Referring Transformer for UAV-based Referring Image SegmentationarXiv 2025Code
[PixFoundation] PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?arXiv 2025Code
[MIRAS] Pixel-Level Reasoning Segmentation via Multi-turn ConversationsarXiv 2025Code
[NegRefCOCOg & NegationCLIP] Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIParXiv 2025
[DETRIS] Densely Connected Parameter-Efficient Tuning for Referring Image SegmentationarXiv 2025Code
[MVP-LM]Advancing Visual Large Language Model for Multi-granular Versatile PerceptionICCVCode
[DeRIS]DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback SynergyICCVCode
[RGA3]Object-centric Video Question Answering with Visual Grounding and ReferringICCVHomepage
[SegAgent] SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator TrajectoriesCVPR 2025Code
[HybridGL] Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image SegmentationCVPR 2025Code
[READ] Reasoning to Attend: Try to Understand How Token WorksCVPR 2025Code
[WeakMCN] WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and SegmentationCVPR 2025Code
[MIMO] MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal OutputCVPR 2025Code
[DViN] DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression ComprehensionCVPR 2025Code
[POPEN] POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning SegmentationCVPR 2025Homepage
[ADDP] Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual PerceptionICLR 2025Code
[MMR] MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning SegmentationICLR 2025Code
[SegLLM] SegLLM: Multi-round Reasoning SegmentationICLR 2025Homepage
[Text4Seg] Text4Seg: Reimagining Image Segmentation as Text GenerationICLR 2025Homepage
[Segment Anyword] Segment Anyword: Mask Prompt Inversion for Open-Set Grounded SegmentationICML 2025Homepage
[IteRPrimE] IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word EmphasisAAAI 2025Code
[PRIMA] PRIMA: Multi-Image Vision-Language Models for Reasoning SegmentationAAAI 2025
[MaTTR]Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal RetrievalICMR 2025
[RSVP] RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-ThoughtACL 2025
[RESMatch] RESMatch: Referring Expression Segmentation in a Semi-Supervised MannerIS 2025
[FIANet] Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image SegmentationTGRS 2025Code
[InstructSeg] InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language ModelsarXiv 2024Code
[MaskRIS] MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image SegmentationarXiv 2024Code
[CroBIM] Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image SegmentationarXiv 2024Code
[PVP] How Well Can Vision Language Models See Image Details?arXiv 2024
[EAVL] EAVL: Explicitly Align Vision and Language for Referring Image SegmentationarXiv 2024
[Shared-RIS] A Simple Baseline with Single-encoder for Referring Image SegmentationarXiv 2024Code
[EVF-SAM] EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything ModelarXiv 2024Code
[S2RM] Spatial Semantic Recurrent Mining for Referring Image SegmentationarXiv 2024
[LaSagnA] LaSagnA: Language-based Segmentation Assistant for Complex QueriesarXiv 2024Code
[LLaVASeg] Empowering Segmentation Ability to Multi-modal Large Language ModelsarXiv 2024
[BSAP] Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot LearningarXiv 2024
[GELLA] Generalizable Entity Grounding via Assistance of Large Language ModelarXiv 2024
[CPRN] Collaborative Position Reasoning Network for Referring Image SegmentationarXiv 2024
[Grounded SAM] Grounded SAM: Assembling Open-World Models for Diverse Visual TasksarXiv 2024Code
[F-LMM] F-LMM: Grounding Frozen Large Multimodal ModelsCVPR 2024Code
[SESAME] See Say and Segment: Teaching LMMs to Overcome False PremisesCVPR 2024Code
[LQMFormer] LQMFormer: Language-aware Query Mask Transformer for Referring Image SegmentationCVPR 2024
[PerceptionGPT] PerceptionGPT: Effectively Fusing Visual Perception into LLMCVPR 2024Code
[PixelLM] PixelLM: Pixel Reasoning with Large Multimodal ModelCVPR 2024Homepage
[Osprey] Osprey: Pixel Understanding with Visual Instruction TuningCVPR 2024Code
[GLaMM] GLaMM: Pixel Grounding Large Multimodal ModelCVPR 2024Code
[RMSIN] Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image SegmentationCVPR 2024Code
[CaR] CLIP as RNN: Segment Countless Visual Concepts without Training EndeavorCVPR 2024Code
[GeoChat] GeoChat: Grounded Large Vision-Language Model for Remote SensingCVPR 2024Code
[PPT] Curriculum Point Prompting for Weakly-Supervised Referring Image SegmentationCVPR 2024
[Prompt-RIS] Prompt-Driven Referring Image Segmentation with Instance ContrastingCVPR 2024
[AnyRef] Multi-modal Instruction Tuned LLMs with Fine-grained Visual PerceptionCVPR 2024
[MRES & UniRES & Refcocom] Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression SegmentationCVPR 2024Code
[GSVA] GSVA: Generalized Segmentation via Multimodal Large Language ModelsCVPR 2024Code
[Barleria] Barleria: An Efficient Tuning Framework for Referring Image SegmentationICLR 2024Code
[OneRef] OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring ModelingNeurIPS 2024Code
[PCNet] Boosting Weakly-Supervised Referring Image Segmentation via Progressive ComprehensionNeurIPS 2024
[OMGLLaVA]OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and UnderstandingNeurIPS 2024
[VRSBench] VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image UnderstandingNeurIPS 2024Code
[PSALM] PSALM: Pixelwise SegmentAtion with Large Multi-Modal ModelECCV 2024Code
[NeMo] Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image SegmentationECCV 2024Homepage
[CoReS] CoReS: Orchestrating the Dance of Reasoning and SegmentationECCV 2024Code
[SAM4MLLM] SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression SegmentationECCV 2024Code
[Pseudo-RIS] Pseudo-RIS: Distinctive Pseudo-supervision Generation for Referring Image SegmentationECCV 2024Code
[GTMS] GTMS: A Gradient-driven Tree-guided Mask-free Referring Image Segmentation MethodECCV 2024Code
[ReMamber] ReMamber: Referring Image Segmentation with Mamba TwisterECCV 2024Code
[SafaRi] SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression SegmentationECCV 2024Homepage
[SegVG] SegVG: Transferring Object Bounding Box to Segmentation for Visual GroundingECCV 2024Code
[SemiRES] SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression SegmentationICML 2024
[NExT-Chat] NExT-Chat: An LMM for Chat, Detection and SegmentationICML 2024Code
[LAVT-RS] Language-Aware Vision Transformer for Referring SegmentationTPAMI 2024Code
[Segment-Select-Correct] Segment, Select, Correct: A Framework for Weakly-Supervised Referring SegmentationECCVW2024Code
[yan2024fuse] Fuse & Calibrate: A bi-directional Vision-Language Guided Framework for Referring Image SegmentationICIC 2024
[VATEX] Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context UnderstandingWACV 2024
[HARIS] HARIS: Human-Like Attention for Reference Image SegmentationICME 2024
[yan2024calibration] Calibration & Reconstruction: Deep Integrated Language for Referring Image SegmentationICMR 2024
[DIT-SAM] Deep Instruction Tuning for Segment Anything ModelACMMM 2024Code
[ASDA] Adaptive Selection based Referring Image SegmentationACMMM 2024Code
[DANet] Rethinking the Implicit Optimization Paradigm with Dual Alignments for Referring Remote Sensing Image SegmentationACMMM 2024
[PTQ4RIS] PTQ4RIS: Post-Training Quantization for Referring Image SegmentationICRA 2024Code
[CLIPU2Net] Robot Manipulation in Salient Vision through Referring Image Segmentation and Geometric ConstraintsICRA 2024
[ETRG] A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot GraspingICRA 2024Homepage
[CLIPUNetr] CLIPUNetr: Assisting Human-robot Interface for Uncalibrated Visual Servoing Control with CLIP-driven Referring Expression SegmentationICRA 2024
[OPT-RSVG] Language-Guided Progressive Attention for Visual Grounding in Remote Sensing ImagesTGRS 2024Code
[LQVG] Language Query-Based Transformer With Multiscale Cross-Modal Alignment for Visual Grounding on Remote Sensing ImagesTGRS 2024Code
[RRSIS] RRSIS: Referring Remote Sensing Image SegmentationTGRS 2024Code
[FAN] Fully Aligned Network for Referring Image SegmentationVCIP 2024
[CrossVLT] Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image SegmentationTMM 2024Code
[SkyScapes] Referring Image Segmentation for Remote Sensing DataIGARSS 2024
[PVD] Parallel Vertex Diffusion for Unified Visual GroundingAAAI 2024
[LLM-Seg] LLM-Seg: Bridging Image Segmentation and Large Language Model ReasoningCVPRW 2024Code
[RIS-CQ] Towards Complex-query Referring Image Segmentation: A Novel BenchmarkTOMM 2024
[RISCLIP] Extending CLIP's Image-Text Alignment to Referring Image SegmentationNAACL-HLT 2024
[LISA++] LISA++: An Improved Baseline for Reasoning Segmentation with Large Language ModelarXiv 2023Code
[ViLaM] Enhancing Visual Grounding and Generalization: A Multi-Task Cycle Training Approach for Vision-Language ModelsarXiv 2023
[BTMAE] Synchronizing Vision and Language: Bidirectional Token-Masking AutoEncoder for Referring Image SegmentationarXiv 2023
[MARIS] MARIS: Referring Image Segmentation via Mutual-Aware Attention FeaturesarXiv 2023
[Ref-Diff] Ref-Diff: Zero-shot Referring Image Segmentation with Generative ModelsarXiv 2023Code
[MMNet] MMNet: Multi-Mask Network for Referring Image SegmentationarXiv 2023
[LGFormer] Linguistic Query-Guided Mask Generation for Referring Image SegmentationarXiv 2023
[rewatbowornwong2023zero] Zero-guidance Segmentation Using Zero Segment LabelsICCV 2023Code
[VPD] Unleashing Text-to-Image Diffusion Models for Visual PerceptionICCV 2023Code
[ETRIS] Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image SegmentationICCV 2023Code
[SaG] Shatter and Gather: Learning Referring Image Segmentation with Text SupervisionICCV 2023Code
[weakly-ris] Weakly Supervised Referring Image Segmentation with Intra-Chunk and Inter-Chunk ConsistencyICCV 2023
[TRIS] Referring Image Segmentation Using Text SupervisionICCV 2023Code
[Partial-RES] Learning To Segment Every Referring Object Point by PointCVPR 2023Code
[MCRES] Meta Compositional Referring Expression SegmentationCVPR 2023
[CGFormer] Contrastive Grouping with Transformer for Referring Image SegmentationCVPR 2023Code
[MagNet] Mask Grounding for Referring Image SegmentationCVPR 2023Homepage
[VG-LAW] Language Adaptive Weight Generation for Multi-task Visual GroundingCVPR 2023Code
[Peekaboo] Peekaboo: Text to Image Diffusion Models are Zero-Shot SegmentorsCVPR 2023Code
[GLEE] General Object Foundation Model for Images and Videos at ScaleCVPR 2023Code
[UNINEXT] Universal Instance Perception as Object Discovery and RetrievalCVPR 2023Code
[X-Decoder] Generalized Decoding for Pixel, Image, and LanguageCVPR 2023Code
[PolyFormer] PolyFormer: Referring Image Segmentation as Sequential Polygon GenerationCVPR 2023Code
[LISA] LISA: Reasoning Segmentation via Large Language ModelCVPR 2023Code
[Global-Local CLIP] Zero-shot Referring Image Segmentation with Global-Local Context FeaturesCVPR 2023Code
[SEEM] Segment Everything Everywhere All at OnceNeurIPS 2023Code
[BKINet] Bilateral Knowledge Interaction Network for Referring Image SegmentationTMM 2023Code
[CM-MaskSD] CM-MaskSD: Cross-Modality Masked Self-Distillation for Referring Image SegmentationTMM 2023
[CARIS] CARIS: Context-Aware Referring Image SegmentationCMMM 2023Code
[CVMN] Unsupervised Domain Adaptation for Referring Semantic SegmentationACMMM 2023Code
[DIOR & MGVLF] RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing DataTGRS 2023Code
[WiCo] WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image SegmentationIJCAI 2023
[SLViT] SLViT: Scale-Wise Language-Guided Vision Transformer for Referring Image SegmentationIJCAI 2023Code
[JMCELN] Referring Image Segmentation via Joint Mask Contextual Embedding Learning and Progressive Alignment NetworkEMNLP 2023
[TAS] Text Augmented Spatial-aware Zero-shot Referring Image SegmentationEMNLP 2023
[SADLR] Semantics-Aware Dynamic Localization and Refinement for Referring Image SegmentationAAAI 2023
[MDSM] Multimodal Diffusion Segmentation Model for Object Segmentation from Manipulation InstructionsIROS 2023
[M3Att] Multi-Modal Mutual Attention and Iterative Interaction for Referring Image SegmentationTIP 2023
[GraspNet-RIS] Towards Generalizable Referring Image Segmentation via Target Prompt and Visual CoherenceICIP 2023
[Omni-RES] Towards Omni-supervised Referring Expression SegmentationICME 2023Code
[PCAN] Position-Aware Contrastive Alignment for Referring Image SegmentationarXiv 2022
[TSEG] Weakly-supervised segmentation of referring expressionsarXiv 2022
[SHNet] Comprehensive Multi-Modal Interactions for Referring Image SegmentationarXiv 2022Code
[GroupViT] GroupViT: Semantic Segmentation Emerges from Text SupervisionCVPR 2022Code
[CRIS] CRIS: CLIP-Driven Referring Image SegmentationCVPR 2022Code
[ReSTR] ReSTR: Convolution-free Referring Image Segmentation Using TransformersCVPR 2022Code
[LAVT] LAVT: Language-Aware Vision Transformer for Referring Image SegmentationCVPR 2022Code
[CLIPSeg] Image Segmentation Using Text and Image PromptsCVPR 2022Code
[CoupAlign] CoupAlign: Coupling Word-Pixel with Sentence-Mask Alignments for Referring Image SegmentationNeurIPS 2022
[SeqTR] SeqTR: A Simple yet Universal Network for Visual GroundingECCV 2022Code
[kesen2022modulating] Modulating Bottom-Up and Top-Down Visual Processing via Language-Conditional FiltersCVPRW 2022Code
[GeoVG] Visual Grounding in Remote Sensing ImagesACMMM 2022Homepage
[feng2022learning] Learning From Box Annotations for Referring Image SegmentationTNNLS 2022Code
[ISF] Instance-Specific Feature Propagation for Referring SegmentationTMM 2022
[PKS] Fully and Weakly Supervised Referring Expression Segmentation with End-to-End LearningTCSVT 2022
[MaIL] MaIL: A Unified Mask-Image-Language Trimodal Network for Referring Image SegmentationarXiv 2021
[BUSNet] Bottom-Up Shift and Reasoning for Referring Image SegmentationCVPR 2021Code
[LTS] Locate then Segment: A Strong Pipeline for Referring Image SegmentationCVPR 2021
[EFN] Encoder Fusion Network with Co-Attention Embedding for Referring Image SegmentationCVPR 2021Code
[Referring Transformer] Referring Transformer: A One-step Approach to Multi-task Visual GroundingNeurIPS 2021Code
[GbS] Detector-Free Weakly Supervised Grounding by SeparationICCV 2021Code
[MDETR] MDETR -- Modulated Detection for End-to-End Multi-Modal UnderstandingICCV 2021Code
[VLT] Vision-Language Transformer and Query Generation for Referring SegmentationTPAMI 2021Code
[CMPC] Cross-Modal Progressive Comprehension for Referring SegmentationTPAMI 2021Code
[TV-Net] Two-stage Visual Cues Enhancement Network for Referring Image SegmentationACMMM 2021Code
[CMPC] Referring Image Segmentation via Cross-Modal Progressive ComprehensionCVPR 2020Code
[BRINet] Bi-directional Relationship Inferring Network for Referring Image SegmentationCVPR 2020Code
[MCN] Multi-Task Collaborative Network for Joint Referring Expression Comprehension and SegmentationCVPR 2020Code
[PhraseCut] PhraseCut: Language-based Image Segmentation in the WildCVPR 2020Code
[LSCM] Linguistic Structure Guided Context Modeling for Referring Image SegmentationECCV 2020Code
[ConvLSTM] Dual Convolutional LSTM Network for Referring Image SegmentationTMM 2020
[CGAN] Cascade Grouped Attention Network for Referring Expression SegmentationACMMM 2020
[CMSA] Cross-Modal Self-Attention Network for Referring Image SegmentationCVPR 2019Code
[CLEVR-Ref+] CLEVR-Ref+: Diagnosing Visual Reasoning with Referring ExpressionsCVPR 2019Code
[NMTree] Learning to Assemble Neural Module Tree Networks for Visual GroundingICCV 2019Code
[STEP] See-Through-Text Grouping for Referring Image SegmentationICCV 2019
[Lang2Seg] Referring Expression Object Segmentation with Caption-Aware ConsistencyBMVC 2019Code
[MAttNet] MAttNet: Modular Attention Network for Referring Expression ComprehensionCVPR 2018Code
[RRN] Referring Image Segmentation via Recurrent Refinement Networks/document/8578700/)CVPR 2018Code
[DMN] Dynamic Multimodal Instance Segmentation Guided by Natural Language QueriesECCV 2018Code
[KWA] Key-Word-Aware Network for Referring Expression Image SegmentationECCV 2018Code
[CMN] Modeling Relationships in Referential Expressions with Compositional Modular NetworksCVPR 2017Code
[RMI] Recurrent Multimodal Interaction for Referring Image SegmentationICCV 2017Code
[MMI & G-Ref] Generation and Comprehension of Unambiguous Object DescriptionsCVPR 2016Code
[LSTM-CNN] Segmentation from Natural Language ExpressionsECCV 2016Code
[RefCOCO&RefCOCO+&RefCOCO(g)] Modeling Context in Referring ExpressionsECCV 2016Code
[ReferItGame] ReferItGame: Referring to Objects in Photographs of Natural ScenesEMNLP 2014Code

2. Referring Video-Object Segmentation (RVOS)

TitleSourceCode / Homepage
[CroBIM-V]CroBIM-V: Memory-Quality Controlled Remote Sensing Referring Video Object SegmentationarXiv 2026
[VideoLoom]VideoLoom: A Video Large Language Model for Joint Spatial-Temporal UnderstandingarXiv 2026
[CERES]Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal InterventionNeurIPS 2025
[ReVSeg]ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement LearningarXiv 2025Homepage
[ProxyFormer]Referring Video Object Segmentation with Cross-Modality Proxy QueriesarXiv 2025
[HCD]Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion ModelarXiv 2025
[PARSE-VOS]Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object SegmentationarXiv 2025
[TQF]Mitigating Query Selection Bias in Referring Video Object SegmentationarXiv 2025
Enhancing Sa2VA for Referent Video Object Segmentation: 2nd Solution for 7th LSVOS RVOS TrackarXiv 2025
[SaSaSa2VA]The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VAarXiv 2025Code
[SVAC]SVAC: Scaling Is All You Need For Referring Video Object SegmentationBMVC 2025Code
[VideoSeg-R1]VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement LearningarXiv 2025Code
[Sa2VA-i]Sa2VA-i: Improving Sa2VA Results with Consistent Training and InferencearXiv 2025Code
[FlowRVS]Deforming Videos to Masks: Flow Matching for Referring Video SegmentationarXiv 2025Code
[RFMNet]Referring Camouflaged Object Detection With Multi-Context Overlapped Windows Cross-AttentionarXiv 2025
[EventRR]EventRR: Event Referential Reasoning for Referring Video Object SegmentationarXiv 2025Code
[Planner-Refiner]Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in VideosarXiv 2025
[LTCA]LTCA: Long-range Temporal Context Attention for Referring Video Object SegmentationTCSVT 2025Code
[MomentSeg]MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel UnderstandingarXiv 2025Code
[MeViSv2] MeViS: A Multi-Modal Dataset for Referring Motion Expression Video SegmentationTPAMI 2025Code
[VoCap]VoCap: Video Object Captioning and Segmentation from Any PromptarXiv 2025Code
[PixFoundation 2.0]PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?arXiv 2025Code
[VideoMolmo]VideoMolmo: Spatio-Temporal Grounding Meets PointingarXiv 2025Code
[CS3]Segmenting Collision Sound Sources in Egocentric VideosarXiv 2025Homepage
[SAMA] SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language ModelsarXiv 2025
[ReSurgSAM2] ReSurgSAM2: Referring Segment Anything in Surgical Video via Credible Long-term TrackingarXiv 2025Code
[Long-RVOS] Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object SegmentationarXiv 2025Homepage
[VEGGIE] VEGGIE: Instructional Editing and Reasoning of Video Concepts with Grounded GenerationarXiv 2025Code
[FindTrack] Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object SegmentationarXiv 2025Code
[JiT] Online Reasoning Video Segmentation with Just-in-Time Digital TwinsarXiv 2025
[ORDiRS] Operating Room Workflow Analysis via Reasoning Segmentation over Digital TwinsarXiv 2025
[TPP] Text-Promptable Propagation for Referring Medical Image Sequence SegmentationarXiv 2025
[Sa2VA] Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and VideosarXiv 2025Code
[MPG-SAM 2] MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationarXiv 2025
[ReferDINO] ReferDINO: Referring Video Object Segmentation with Visual Grounding FoundationsarXiv 2025Code
[VRS-HQ] The Devil is in Temporal Token: High Quality Video Reasoning SegmentationCVPR 2025Code
[GLUS] GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video SegmentationCVPR 2025Code
[MoRA] Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel LevelCVPR 2025Homepage
[ViCaS] ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded SegmentationCVPR 2025Homepage
[SSA] Semantic and Sequential Alignment for Referring Video Object SegmentationCVPR 2025
[DMVS] Decoupled Motion Expression Video SegmentationCVPR 2025Code
[MTCM] Multi-Context Temporal Consistent Modeling for Referring Video Object SegmentationICASSP 2025Code
[FS-RVMOS] Few-Shot Referring Video Single- and Multi-Object Segmentation Via Cross-Modal Affinity with Instance Sequence MatchingIJCV 2025Code
[SOLA] Referring Video Object Segmentation via Language-aligned Track SelectionarXiv 2024Homepage
[InstructSeg] InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language ModelsarXiv 2024Code
[REM] ReferEverything: Towards Segmenting Everything We Can Speak of in VideosarXiv 2024Homepage
[ViLLa] ViLLa: Video Reasoning Segmentation with Large Language ModelarXiv 2024Code
[VLP-RVOS] Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object SegmentationarXiv 2024
[DsHmp] Decoupling Static and Hierarchical Motion Perception for Referring Video SegmentationCVPR 2024Code
[SAMWISE] SAMWISE: Infusing wisdom in SAM2 for Text-Driven Video SegmentationCVPR 2024Code
[OMG-Seg] OMG-Seg: Is One Model Good Enough for all Segmentation?CVPR 2024Homepage
[UniVS] UniVS: Unified and Universal Video Segmentation with Prompts as QueriesCVPR 2024Code
[VideoGLaMM] VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in VideosCVPR 2024Code
[LoSh] LoSh: Long-Short Text Joint Prediction Network for Referring Video Object SegmentationCVPR 2024Code
[VideoLISA] One Token to Seg Them All: Language Instructed Reasoning Segmentation in VideosNeurIPS 2024Code
[SOC] SOC: Semantic-Assisted Object Cluster for Referring Video Object SegmentationNeurIPS 2024Code
[ActionVOS] ActionVOS: Actions as Prompts for Video Object SegmentationECCV 2024Code
[VD-IT] Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object SegmentationECCV 2024Code
[VISA] VISA: Reasoning Video Object Segmentation via Large Language ModelsECCV 2024Code
[GroPrompt] GroPrompt: Efficient Grounded Prompting and Adaptation for Referring Video Object SegmentationCVPRW 2024Homepage
[AL-Ref-SAM 2] Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object SegmentationAAAI 2024Code
[MUTR] Referred by Multi-Modality: A Unified Temporal Transformer for Video Object SegmentationAAAI 2024Code
[TF2] Tracking-forced Referring Video Object SegmentationACMMM 2024
[SLVP] SLVP: Self-Supervised Language-Video Pre-Training for Referring Video Object SegmentationWACV 2024
[TCE-RVOS] Temporal Context Enhanced Referring Video Object SegmentationWACV 2024Code
[HTR] Temporally Consistent Referring Video Object Segmentation with Hybrid MemoryTCSVT 2024Code
[TrackGPT] Tracking with Human-Intent ReasoningarXiv 2023Code
[SimRVOS] Learning Referring Video Object Segmentation from Weak AnnotationarXiv 2023
[RefSAM] RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object SegmentationarXiv 2023Code
[SgMg] Spectrum-guided Multi-granularity Referring Video Object SegmentationICCV 2023Code
[OnlineRefer] OnlineRefer: A Simple Online Baseline for Referring Video Object SegmentationICCV 2023Code
[HTML] HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object SegmentationICCV 2023Homepage
[TempCD] Temporal Collection and Distribution for Referring Video Object SegmentationICCV 2023Homepage
[MeViS] MeViS: A Large-scale Benchmark for Video Segmentation with Motion ExpressionsICCV 2023Code
[UniRef++] Segment Every Reference Object in Spatial and Temporal SpacesICCV 2023Code
[FS-RVOS] Learning Cross-Modal Affinity for Referring Video Object Segmentation Targeting Limited SamplesICCV 2023Code
[DMFormer] Decoupling Multimodal Transformers for Referring Video Object SegmentationTCSVT 2023Code
[UniMM] Unified Multi-Modality Video Object Segmentation Using Reinforcement LearningTCSVT 2023
[EPCFormer] EPCFormer: Expression Prompt Collaboration Transformer for Universal Referring Video Object SegmentationKS 2023
[STBridge] Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and TextEMNLP 2023
[Locater] Local-Global Context Aware Transformer for Language-Guided Video SegmentationTPAMI 2023Code
[LASTC] Language-Aware Spatial-Temporal Collaboration for Referring Video SegmentationTPAMI 2023
[CLUE] CLUE: Contrastive language-guided learning for referring video object segmentationPRL 2023
[BIFIT] Bidirectional Correlation-Driven Inter-Frame Interaction Transformer for Referring Video Object SegmentationPR 2023
[FTEA] Fully Transformer-Equipped Architecture for end-to-end Referring Video Object SegmentationIPM 2023
[MTTR] End-to-End Referring Video Object Segmentation with Multimodal TransformersCVPR 2022Code
[ReferFormer] Language as Queries for Referring Video Object SegmentationCVPR 2022Code
[LBDT] Language-Bridged Spatial-Temporal Interaction for Referring Video Object SegmentationCVPR 2022Code
[MLRL] Multi-Level Representation Learning with Semantic Alignment for Referring Video Object SegmentationCVPR 2022
[zhao2022modeling] Modeling Motion with Multi-Modal Features for Text-Based Video SegmentationCVPR 2022Code
[RefVOS] A closer look at referring expressions for video object segmentationMTA 2022Code
[OATNet] Object-Agnostic Transformers for Video Referring SegmentationTIP 2022
[EFCMA] Referring Segmentation via Encoder-Fused Cross-Modal Attention NetworkTPAMI 2022
[MANet] Multi-Attention Network for Compressed Video Referring Object SegmentationACMMM 2022Code
[YOFO] You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object SegmentationAAAI 2022Code
[CITD] Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object SegmentationarXiv 2021
[ClawCraneNet] ClawCraneNet: Leveraging Object-level Relation for Text-based Video SegmentationarXiv 2021
[hui2021collaborative] Collaborative Spatial-Temporal Modeling for Language-Queried Video Actor SegmentationCVPR 2021Code
[CMSA] Referring Segmentation in Images and Videos With Cross-Modal Self-Attention NetworkTPAMI 2021Code
[CMPC] Cross-Modal Progressive Comprehension for Referring SegmentationTPAMI 2021Code
[mcintosh2020visual] Visual-Textual Capsule Routing for Text-Based Video SegmentationCVPR 2020
[Refer-Youtube-VOS] URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale BenchmarkECCV 2020Code
[wang2020context] Context Modulated Dynamic Networks for Actor and Action Video Segmentation with Language QueriesAAAI 2020
[ACGA] Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language QueryICCV 2019Code
[A2D] Actor and Action Video Segmentation from a SentenceCVPR 2018Homepage
[Refer-DAVIS] Video Object Segmentation with Language Referring ExpressionsACCV 2018

3. Referring Audio-Visual Segmentation (RAVS)

TitleSourceCode / Homepage
[MQA-RefAVS]Audit After Segmentation: Reference-Free Mask Quality Assessment for Language-Referred Audio-Visual SegmentationarXiv 2026Code
[SeaVIS]SeaVIS: Sound-Enhanced Association for Online Audio-Visual Instance SegmentationarXiv 2026
[ATLAS]Can You Hear, Localize, and Segment Continually? An Exemplar-Free Continual Learning Benchmark for Audio-Visual SegmentationarXiv 2026Code
[RA-SSU]RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source UnderstandingTMM 2026Homepage
[SDAVS]Selective Noise Suppression and Discriminative Mutual Interaction for Robust Audio-Visual SegmentationTMM 2026Code
[WSAVSS]Look, Listen and Segment: Towards Weakly Supervised Audio-visual Semantic SegmentationICASSP 2026
[SOUPLE]SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt ContextsCVPR 2026
[MAR3]MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual SegmentationarXiv 2026
[SSP]How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?arXiv 2026
[DDAVS]DDAVS: Disentangled Audio Semantics and Delayed Bidirectional Alignment for Audio-Visual SegmentationarXiv 2025Homepage
[AVAGFormer]Learning Visual Affordance from AudioarXiv 2025Homepage
Layover or Direct Flight: Rethinking Audio-Guided Image SegmentationarXiv 2025
[FAVS]Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual SegmentationarXiv 2025
[CCFormer]Complementary and Contrastive Learning for Audio-Visual SegmentationTMM 2025Code
[SimToken]SimToken: A Simple Baseline for Referring Audio-Visual SegmentationarXiv 2025Homepage
[AVS Survey]From Waveforms to Pixels: A Survey on Audio-Visual Segmentationarxiv 2025
[AURORA]AURORA: Augmented Understanding via Structured Reasoning and Reinforcement Learning for Reference Audio-Visual Segmentationarxiv 2025
[TGS-Agent]Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentationarxiv 2025Code
[ICF]Implicit Counterfactual Learning for Audio-Visual Segmentationarxiv 2025
[Mettle]Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptationarxiv 2025
[TAViS]TAViS: Text-bridged Audio-Visual Segmentation with Foundation Modelsarxiv 2025
[OpenAVS] OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Modelsarxiv 2025
[Omni-R1] Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaborationarxiv 2025Code
[RAVS] Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignmentarxiv 2025
[AVSBench-Robust] Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?arxiv 2025
[AV2T-SAM] Audio Visual Segmentation Through Text Embeddingsarxiv 2025Code
[AVS-Mamba] AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentationarxiv 2025Code
[OmniAVS] Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual SegmentationICCV 2025Code
[SAM2-LOVE] SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual ScenesCVPR 2025Code
[DDESeg] Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio SemanticsCVPR 2025Code
[VCT] Revisiting Audio-Visual Segmentation with Vision-Centric TransformerCVPR 2025Code
[TSAM] TSAM: Temporal SAM Augmented with Multimodal Prompts for Referring Audio-Visual SegmentationCVPR 2025Homepage
[Dolphin] Aligned Better, Listen Better for Audio-Visual Large Language ModelsICLR 2025
[Co-Prop] Collaborative Hybrid Propagator for Temporal Misalignment in Audio-Visual SegmentationarXiv 2024
[3D AVS] 3D Audio-Visual SegmentationarXiv 2024Code
[AVIS] Audio-Visual Instance SegmentationarXiv 2024Code
[AVESFormer] AVESFormer: Efficient Transformer Design for Real-Time Audio-Visual SegmentationarXiv 2024Code
[SAVE] SAVE: Segment Audio-Visual Easy way using Segment Anything ModelarXiv 2024
[PMCANet] Progressive Confident Masking Attention Network for Audio-Visual SegmentationarXiv 2024Code
[MED-VT++] MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video TransformerarXiv 2024Code
[MoCA] Unsupervised Audio-Visual Segmentation with Modality AlignmentarXiv 2024
[AVSAC] Bootstrapping Audio-Visual Segmentation by Strengthening Audio CuesarXiv 2024
[VPO & CAVP] Unraveling Instance Associations: A Closer Look for Audio-Visual SegmentationCVPR 2024Code
[QDFormer] QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic DecompositionCVPR 2024Code
[LU-AVS] Benchmarking Audio Visual Segmentation for Long-Untrimmed VideosCVPR 2024Homepage
[COMBO] Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual SegmentationCVPR 2024Code
[TeSO] Can Textual Semantics Mitigate Sounding Object Segmentation Preference?ECCV 2024Code
[Stepping Stones] Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic SegmentationECCV 2024Code
[CPM] CPM: Class-conditional Prompting Machine for Audio-visual SegmentationECCV 2024
[Ref-AVS] Ref-AVS: Refer and Segment Objects in Audio-Visual ScenesECCV 2024Code
[TransAVS] TransAVS: End-to-End Audio-Visual Segmentation with TransformerICASSP 2024
[sun2024unveiling] Unveiling and Mitigating Bias in Audio Visual SegmentationACMMM 2024Homepage
[SelM] SelM: Selective Mechanism based Audio-Visual SegmentationACMMM 2024Code
[OV-AVSS] Open-Vocabulary Audio-Visual Semantic SegmentationACMMM 2024Code
[C3N] Cross-modal Cognitive Consensus guided Audio-Visual SegmentationTMM 2024Code
[BAVS] BAVS: Bootstrapping Audio-Visual Segmentation by Integrating Foundation KnowledgeTMM 2024
[PIF] Each Performs Its Functions: Task Decomposition and Feature Assignment for Audio-Visual SegmentationTMM 2024Code
[ST-BAVA] Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual SegmentationICIP 2024Code
[SBV] Segment beyond View: Handling Partially Missing Modality for Audio-Visual Semantic SegmentationAAAI 2024
[AVS-bigen] Improving Audio-Visual Segmentation with Bidirectional GenerationAAAI 2024Code
[AVSegFormer] AVSegFormer: Audio-Visual Segmentation with TransformerAAAI 2024Code
[GAVS] Prompting Segmentation with Sound Is Generalizable Audio-Visual Source LocalizerAAAI 2024Code
[UFE] Audio-Visual Segmentation via Unlabeled Frame ExploitationIJCAI 2024Code
[AVSBench-semantic] Audio-Visual Segmentation with SemanticsIJCV 2024Code
[CMSF] Leveraging Foundation models for Unsupervised Audio-Visual SegmentationarXiv 2023
[AuTR] Audio-aware Query-enhanced Transformer for Audio-Visual SegmentationarXiv 2023
[DiffusionAVS] Contrastive Conditional Latent Diffusion for Audio-visual SegmentationarXiv 2023Code
[AV-SAM] AV-SAM: Segment Anything Model Meets Audio-Visual Localization and SegmentationarXiv 2023
[LAVISH] Vision Transformers are Parameter-Efficient Audio-Visual LearnersCVPR 2023Code
[DeepAVFusion] Unveiling the Power of Audio-Visual Early Fusion Transformers with Dense Interactions through Masked ModelingCVPR 2023Code
[WS-AVS] Weakly-Supervised Audio-Visual SegmentationNeurIPS 2023
[AVS-Bench] Audio-Visual SegmentationECCV 2023Code
[ECMVAE] Multimodal Variational Auto-encoder based Audio-Visual SegmentationICCV 2023Code
[AVSC] Audio-Visual Segmentation by Exploring Cross-Modal Mutual SemanticsACMMM 2023
[CATR] CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video SegmentationACMMM 2023Homepage
[SAMA-AVS] Annotation-free Audio-Visual SegmentationWACV 2023Code
[AQFormer] Discovering Sounding Objects by Audio Queries for Audio Visual SegmentationIJCAI 2023

4. 3D-Referring Expression Segmentation (3D-RES)

TitleSourceCode / Homepage
[OpenVoxel]OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene UnderstandingarXiv 2026
[MVGGT]MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression SegmentationarXiv 2026Homepage
[NDTokenizer3D]Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language UnderstandingarXiv 2025
[PLM]Point Linguist Model: Segment Any Object via Bridged Large 3D-Language ModelarXiv 2025Code
[CaRF]CaRF: Enhancing Multi-View Consistency in Referring 3D Gaussian Splatting SegmentationarXiv 2025
[OV-BIS]OV-BIS: Open-Vocabulary Boundary Guide Zero-Shot 3D Instance SegmentationTMM 2025
[MORE3D] Multimodal 3D Reasoning Segmentation with Complex ScenesarXiv 2025
[3DResT] 3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression SegmentationarXiv 2025
[MLLM-For3D] MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning SegmentationarXiv 2025
[3D-LLaVA] 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint TransformerCVPR 2025Code
[MEN] Weakly-Supervised 3D Referring Expression SegmentationICLR 2025
[ReferSplat] ReferSplat: Referring Segmentation in 3D Gaussian SplattingICML 2025Code
[IPDN] IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression SegmentationAAAI 2025Code
[Reason3D] Reason3D: Searching and Reasoning 3D Segmentation via Large Language ModelarXiv 2025Code
[RG-SAN] RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression SegmentationarXiv 2024Code
[LESS] LESS: Label-Efficient and Single-Stage Referring 3D SegmentationarXiv 2024Code
[ConcreteNet] Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual GroundingarXiv 2024Homepage
[Reasoning3D] Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language ModelsarXiv 2024Homepage
[UniSeg3D] A Unified Framework for 3D Scene UnderstandingNeurIPS 2024Code
[MCLN] Multi-branch Collaborative Learning Network for 3D Visual GroundingECCV 2024Code
[SegPoint] SegPoint: Segment Any Point Cloud via Large Language ModelECCV 2024Homepage
[X-RefSeg3D] X-RefSeg3D: Enhancing Referring 3D Instance Segmentation via Structured Cross-Modal Graph Neural NetworksAAAI 2024Code
[RefMask3D] RefMask3D: Language-Guided Transformer for 3D Referring SegmentationACMMM 2024Code
[MDIN & 3D-GRES] 3D-GRES: Generalized 3D Referring Expression SegmentationACMMM 2024Code
[Uni3DL] Uni3DL: Unified Model for 3D and Language UnderstandingarXiv 2023Homepage
[3DRefTR] A Unified Framework for 3D Point Cloud Visual GroundingarXiv 2023Code
[3D-STMN] 3D-STMN: Dependency-Driven Superpoint-Text Matching Network for End-to-End 3D Referring Expression SegmentationAAAI 2023Code
[TGNN] Text-Guided Graph Neural Networks for Referring 3D Instance SegmentationAAAI 2021Code

5. Generalized Referring Expression x (GREx)

TitleSourceCode / Homepage
[LIHE]LIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression ComprehensionarXiv 2025Code
Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression ComprehensionarXiv 2025
[GREx]Generalized Referring Expression Segmentation, Comprehension, and GenerationIJCV 2026Homepage
[RSRefSeg]Generalized Referring Expression Segmentation on Aerial PhotosarXiv 2025Homepage
[InstanceVG]Improving Generalized Visual Grounding with Instance-aware Joint LearningTPAMI 2025Code
[HieA2G] Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression ComprehensionAAAI 2025
[nguyen2024instance] Instance-Aware Generalized Referring Expression SegmentationarXiv 2024Homepage
[SimVG] SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal FusionarXiv 2024Code
[CoHD] CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression SegmentationarXiv 2024Code
[HDC] HDC: Hierarchical Semantic Decoding with Counting Assistance for Generalized Referring Expression SegmentationarXiv 2024Code
[LaSagnA] LaSagnA: Language-based Segmentation Assistant for Complex QueriesarXiv 2024Code
[GSVA] GSVA: Generalized Segmentation via Multimodal Large Language ModelsCVPR 2024Code
[LQMFormer] LQMFormer: Language-aware Query Mask Transformer for Referring Image SegmentationCVPR 2024
[D-LISA] Multi-Object 3D Grounding with Dynamic Modules and Language-Informed Spatial AttentionNeurIPS 2024Code
[MDIN & 3D-GRES] 3D-GRES: Generalized 3D Referring Expression SegmentationACMMM 2024Code
[GNL3D] Advancing 3D Object Grounding Beyond a Single 3D SceneACMMM 2024
[MABP] Bring Adaptive Binding Prototypes to Generalized Referring Expression SegmentationTMM 2024Code
[R-RIS] Towards Robust Referring Image SegmentationTIP 2024Code
[GREC] GREC: Generalized Referring Expression ComprehensionarXiv 2023Code
[GRES& ReLA] GRES: Generalized Referring Expression SegmentationCVPR 2023Code
[Group-RES] Advancing Referring Expression Segmentation Beyond Single ImageICCV 2023Code
[DMMI] Beyond One-to-One: Rethinking the Referring Image SegmentationICCV 2023Code
[SgMg] Spectrum-guided Multi-granularity Referring Video Object SegmentationICCV 2023Code
[R2VOS] Robust Referring Video Object Segmentation with Cyclic Structural ConsensusICCV 2023Code
[Multi3DRefer] Multi3DRefer: Grounding Text Description to Multiple 3D ObjectsICCV 2023Homepage
[3DOGSFormer] Dense Object Grounding in 3D ScenesACMMM 2023

6. Application

TitleSourceCode / Homepage
[CAD-GD] Exploring Contextual Attribute Density in Referring Expression CountingCVPR 2025Code
[EDGS] Grasp What You Want: Embodied Dexterous Grasping System Driven by Your VoicearXiv 2024
[HiFi-CS] HiFi-CS: Towards Open Vocabulary Visual Grounding For Robotic Grasping Using Vision-Language ModelsarXiv 2024Code
[ETRG] A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot GraspingarXiv 2024Homepage
[Invigorate] INVIGORATE: Interactive Visual Grounding and Grasping in ClutterarXiv 2024
[Grasp-Anything++] Language-driven Grasp DetectionCVPR 2024Code
[Referring Image Editing] Referring Image Editing: Object-Level Image Editing via Referring ExpressionsCVPR 2024
[REC] Referring Expression CountingCVPR 2024Code
[CountGD] CountGD: Multi-Modal Open-World CountingNeurIPS 2024Homepage
[Reasoning Grasping] Reasoning Grasping via Multimodal Large Language ModelCoRL 2024Homepage
[CROG] Language-guided Robot Grasping: CLIP-based Referring Grasp Synthesis in ClutterarXiv 2023Code
[LAN-grasp] LAN-grasp: Using Large Language Models for Semantic Object GraspingarXiv 2023
[FLarG] Language Guided Robotic Grasping with Fine-Grained InstructionsIROS 2023
[VL-Grasp] VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor ScenesIROS 2023Code
[GraspCLIP] Task-Oriented Grasp Prediction with Visual-Language InputsIROS 2023
[xu2023joint] A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in ClutterICRA 2023Code
[INGRESS] Interactive Visual Grounding of Referring Expressions for Human-Robot InteractionarXiv 2018Code
[rao2018learning] Learning Robotic Grasping Strategy Based on Natural-Language Object DescriptionsIROS 2018
[shridhar2017grounding] Grounding Spatio-Semantic Referring Expressions for Human-Robot InteractionarXiv 2017

❤️ Citation

We would be honored if this work could assist you, and greatly appreciate it if you could consider starring and citing it:

@article{ReferringSurvey,
  title={Multimodal Referring Segmentation: A Survey},
  author={Ding, Henghui and Tang, Song and He, Shuting and Liu, Chang and Wu, Zuxuan and Jiang, Yu-Gang},
  journal={International Journal of Computer Vision (IJCV)},
  year={2026}
}

⭐️ Star History

Star History Chart