| [ORES]Refer to Any Segmentation Mask Group With Vision-Language Prompts |  | Homepage |
| [Falcon Perception]Falcon Perception |  |  |
| [DR2Seg]DR2Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models |  | |
| [CroBIM-U]CroBIM-U: Uncertainty-Driven Referring Remote Sensing Image Segmentation |  | |
| [IBISAgent]IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal |  | |
| [EVOL-SAM3]Evolving, Not Training: Zero-Shot Reasoning Segmentation via Evolutionary Prompting |  |  |
| [SSA]Spatial-aware Symmetric Alignment for Text-guided Medical Image Segmentation |  | |
| [Think2Seg-RS]Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Remote Sensing |  |  |
| [OmniRIS]Omni-Referring Image Segmentation |  |  |
| [UniGeoSeg]UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes |  |  |
| [SaFiRe]SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation |  | Homepage |
| [UniPixel]UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning |  |  |
| [LRB-WREL]Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions |  | |
| [CSINet]Referring Remote Sensing Image Segmentation with Cross-view Semantics Interaction Network |  | |
| [DGL-RSIS]DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation |  | |
| [RIS-FUSION]RIS-FUSION: Rethinking Text-Driven Infrared and Visible Image Fusion from the Perspective of Referring Image Segmentation |  | |
| [SVP]Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation |  | |
| [UGround]UGround: Towards Unified Visual Grounding with Unrolled Transformers |  |  |
| [CoT Referring]CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning |  | |
| [MedSeg EarlyFusion]A Text-Image Fusion Method with Data Augmentation Capabilities for Referring Medical Image Segmentation |  |  |
| [PixelRefer]PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity |  |  |
| [RefAM]RefAM: Attention Magnets for Zero-Shot Referral Segmentation |  | Homepage |
| [CoPatch]CoPatch: Zero-Shot Referring Image Segmentation by Leveraging Untapped Spatial Knowledge in CLIP |  | |
| [Latent-VG]Latent Expression Generation for Referring Image Segmentation and Grounding |  | |
| [RIS-LAD]RIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation |  |  |
| [X-SAM]X-SAM: From Segment Anything to Any Segmentation |  |  |
| [Seg-R1]Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning |  | |
| [ELBO-T2IAlign]ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-level Text-Image Alignment in Diffusion Models |  | |
| [MBA]MBA: Multimodal Bidirectional Attack for Referring Expression Segmentation Models |  | |
| [FOCUS]FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation |  | |
| [NWPU-Refer]A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark |  |  |
| [SegVLM]Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model |  | |
| [RemoteSAM] RemoteSAM: Towards Segment Anything for Earth Observation |  |  |
| [LISAT] LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery |  | Homepage |
| [PRS-Med] PRS-Med: Position Reasoning Segmentation with Vision-Language Model in Medical Imaging |  | |
| [RVTBench] RVTBench: A Benchmark for Visual Reasoning Tasks |  |  |
| [SAM-R1] SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning |  | |
| [VisionReasoner] VisionReasoner: Unified Visual Perception and Reasoning via Reinforcement Learning |  |  |
| [ReasoningSeg Survey] Reasoning Segmentation for Images and Videos: A Survey |  | |
| [PixelThink] PixelThink: Towards Efficient Chain-of-Pixel Reasoning |  | |
| [SynRES] SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data |  |  |
| [RESAnything]RESAnything: Attribute Prompting for Arbitrary Referring Segmentation |  | Homepage |
| [Pixel-SAIL] Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding |  |  |
| [LVLM_CSP] LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning Segmentation |  | |
| [SegEarth-R1] SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model |  |  |
| [UniRES++] Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities |  |  |
| [MediSee] MediSee: Reasoning-based Pixel-level Perception in Medical Images |  |  |
| [PLVL] Progressive Language-guided Visual Learning for Multi-Task Visual Grounding |  |  |
| [UFO] UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface |  |  |
| [GroundingSuite] GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding |  |  |
| [UniVG] UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing |  | |
| [AURA] Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURA |  | |
| [Seg-Zero] Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement |  |  |
| [AeroReformer] AeroReformer: Aerial Referring Transformer for UAV-based Referring Image Segmentation |  |  |
| [PixFoundation] PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models? |  |  |
| [MIRAS] Pixel-Level Reasoning Segmentation via Multi-turn Conversations |  |  |
| [NegRefCOCOg & NegationCLIP] Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP |  | |
| [DETRIS] Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation |  |  |
| [MVP-LM]Advancing Visual Large Language Model for Multi-granular Versatile Perception |  |  |
| [DeRIS]DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy |  |  |
| [RGA3]Object-centric Video Question Answering with Visual Grounding and Referring |  | Homepage |
| [SegAgent] SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories |  |  |
| [HybridGL] Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation |  |  |
| [READ] Reasoning to Attend: Try to Understand How Token Works |  |  |
| [WeakMCN] WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation |  |  |
| [MIMO] MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal Output |  |  |
| [DViN] DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression Comprehension |  |  |
| [POPEN] POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation |  | Homepage |
| [ADDP] Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception |  |  |
| [MMR] MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation |  |  |
| [SegLLM] SegLLM: Multi-round Reasoning Segmentation |  | Homepage |
| [Text4Seg] Text4Seg: Reimagining Image Segmentation as Text Generation |  | Homepage |
| [Segment Anyword] Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation |  | Homepage |
| [IteRPrimE] IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis |  |  |
| [PRIMA] PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation |  | |
| [MaTTR]Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval |  | |
| [RSVP] RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought |  | |
| [RESMatch] RESMatch: Referring Expression Segmentation in a Semi-Supervised Manner |  | |
| [FIANet] Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation |  |  |
| [InstructSeg] InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models |  |  |
| [MaskRIS] MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation |  |  |
| [CroBIM] Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation |  |  |
| [PVP] How Well Can Vision Language Models See Image Details? |  | |
| [EAVL] EAVL: Explicitly Align Vision and Language for Referring Image Segmentation |  | |
| [Shared-RIS] A Simple Baseline with Single-encoder for Referring Image Segmentation |  |  |
| [EVF-SAM] EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model |  |  |
| [S2RM] Spatial Semantic Recurrent Mining for Referring Image Segmentation |  | |
| [LaSagnA] LaSagnA: Language-based Segmentation Assistant for Complex Queries |  |  |
| [LLaVASeg] Empowering Segmentation Ability to Multi-modal Large Language Models |  | |
| [BSAP] Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning |  | |
| [GELLA] Generalizable Entity Grounding via Assistance of Large Language Model |  | |
| [CPRN] Collaborative Position Reasoning Network for Referring Image Segmentation |  | |
| [Grounded SAM] Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks |  |  |
| [F-LMM] F-LMM: Grounding Frozen Large Multimodal Models |  |  |
| [SESAME] See Say and Segment: Teaching LMMs to Overcome False Premises |  |  |
| [LQMFormer] LQMFormer: Language-aware Query Mask Transformer for Referring Image Segmentation |  | |
| [PerceptionGPT] PerceptionGPT: Effectively Fusing Visual Perception into LLM |  |  |
| [PixelLM] PixelLM: Pixel Reasoning with Large Multimodal Model |  | Homepage |
| [Osprey] Osprey: Pixel Understanding with Visual Instruction Tuning |  |  |
| [GLaMM] GLaMM: Pixel Grounding Large Multimodal Model |  |  |
| [RMSIN] Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation |  |  |
| [CaR] CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor |  |  |
| [GeoChat] GeoChat: Grounded Large Vision-Language Model for Remote Sensing |  |  |
| [PPT] Curriculum Point Prompting for Weakly-Supervised Referring Image Segmentation |  | |
| [Prompt-RIS] Prompt-Driven Referring Image Segmentation with Instance Contrasting |  | |
| [AnyRef] Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception |  | |
| [MRES & UniRES & Refcocom] Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation |  |  |
| [GSVA] GSVA: Generalized Segmentation via Multimodal Large Language Models |  |  |
| [Barleria] Barleria: An Efficient Tuning Framework for Referring Image Segmentation |  |  |
| [OneRef] OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling |  |  |
| [PCNet] Boosting Weakly-Supervised Referring Image Segmentation via Progressive Comprehension |  | |
| [OMGLLaVA]OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding |  | |
| [VRSBench] VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding |  |  |
| [PSALM] PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model |  |  |
| [NeMo] Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation |  | Homepage |
| [CoReS] CoReS: Orchestrating the Dance of Reasoning and Segmentation |  |  |
| [SAM4MLLM] SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation |  |  |
| [Pseudo-RIS] Pseudo-RIS: Distinctive Pseudo-supervision Generation for Referring Image Segmentation |  |  |
| [GTMS] GTMS: A Gradient-driven Tree-guided Mask-free Referring Image Segmentation Method |  |  |
| [ReMamber] ReMamber: Referring Image Segmentation with Mamba Twister |  |  |
| [SafaRi] SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation |  | Homepage |
| [SegVG] SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding |  |  |
| [SemiRES] SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation |  | |
| [NExT-Chat] NExT-Chat: An LMM for Chat, Detection and Segmentation |  |  |
| [LAVT-RS] Language-Aware Vision Transformer for Referring Segmentation |  |  |
| [Segment-Select-Correct] Segment, Select, Correct: A Framework for Weakly-Supervised Referring Segmentation |  |  |
| [yan2024fuse] Fuse & Calibrate: A bi-directional Vision-Language Guided Framework for Referring Image Segmentation |  | |
| [VATEX] Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding |  | |
| [HARIS] HARIS: Human-Like Attention for Reference Image Segmentation |  | |
| [yan2024calibration] Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation |  | |
| [DIT-SAM] Deep Instruction Tuning for Segment Anything Model |  |  |
| [ASDA] Adaptive Selection based Referring Image Segmentation |  |  |
| [DANet] Rethinking the Implicit Optimization Paradigm with Dual Alignments for Referring Remote Sensing Image Segmentation |  | |
| [PTQ4RIS] PTQ4RIS: Post-Training Quantization for Referring Image Segmentation |  |  |
| [CLIPU2Net] Robot Manipulation in Salient Vision through Referring Image Segmentation and Geometric Constraints |  | |
| [ETRG] A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping |  | Homepage |
| [CLIPUNetr] CLIPUNetr: Assisting Human-robot Interface for Uncalibrated Visual Servoing Control with CLIP-driven Referring Expression Segmentation |  | |
| [OPT-RSVG] Language-Guided Progressive Attention for Visual Grounding in Remote Sensing Images |  |  |
| [LQVG] Language Query-Based Transformer With Multiscale Cross-Modal Alignment for Visual Grounding on Remote Sensing Images |  |  |
| [RRSIS] RRSIS: Referring Remote Sensing Image Segmentation |  |  |
| [FAN] Fully Aligned Network for Referring Image Segmentation |  | |
| [CrossVLT] Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image Segmentation |  |  |
| [SkyScapes] Referring Image Segmentation for Remote Sensing Data |  | |
| [PVD] Parallel Vertex Diffusion for Unified Visual Grounding |  | |
| [LLM-Seg] LLM-Seg: Bridging Image Segmentation and Large Language Model Reasoning |  |  |
| [RIS-CQ] Towards Complex-query Referring Image Segmentation: A Novel Benchmark |  | |
| [RISCLIP] Extending CLIP's Image-Text Alignment to Referring Image Segmentation |  | |
| [LISA++] LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model |  |  |
| [ViLaM] Enhancing Visual Grounding and Generalization: A Multi-Task Cycle Training Approach for Vision-Language Models |  | |
| [BTMAE] Synchronizing Vision and Language: Bidirectional Token-Masking AutoEncoder for Referring Image Segmentation |  | |
| [MARIS] MARIS: Referring Image Segmentation via Mutual-Aware Attention Features |  | |
| [Ref-Diff] Ref-Diff: Zero-shot Referring Image Segmentation with Generative Models |  |  |
| [MMNet] MMNet: Multi-Mask Network for Referring Image Segmentation |  | |
| [LGFormer] Linguistic Query-Guided Mask Generation for Referring Image Segmentation |  | |
| [rewatbowornwong2023zero] Zero-guidance Segmentation Using Zero Segment Labels |  |  |
| [VPD] Unleashing Text-to-Image Diffusion Models for Visual Perception |  |  |
| [ETRIS] Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image Segmentation |  |  |
| [SaG] Shatter and Gather: Learning Referring Image Segmentation with Text Supervision |  |  |
| [weakly-ris] Weakly Supervised Referring Image Segmentation with Intra-Chunk and Inter-Chunk Consistency |  | |
| [TRIS] Referring Image Segmentation Using Text Supervision |  |  |
| [Partial-RES] Learning To Segment Every Referring Object Point by Point |  |  |
| [MCRES] Meta Compositional Referring Expression Segmentation |  | |
| [CGFormer] Contrastive Grouping with Transformer for Referring Image Segmentation |  |  |
| [MagNet] Mask Grounding for Referring Image Segmentation |  | Homepage |
| [VG-LAW] Language Adaptive Weight Generation for Multi-task Visual Grounding |  |  |
| [Peekaboo] Peekaboo: Text to Image Diffusion Models are Zero-Shot Segmentors |  |  |
| [GLEE] General Object Foundation Model for Images and Videos at Scale |  |  |
| [UNINEXT] Universal Instance Perception as Object Discovery and Retrieval |  |  |
| [X-Decoder] Generalized Decoding for Pixel, Image, and Language |  |  |
| [PolyFormer] PolyFormer: Referring Image Segmentation as Sequential Polygon Generation |  |  |
| [LISA] LISA: Reasoning Segmentation via Large Language Model |  |  |
| [Global-Local CLIP] Zero-shot Referring Image Segmentation with Global-Local Context Features |  |  |
| [SEEM] Segment Everything Everywhere All at Once |  |  |
| [BKINet] Bilateral Knowledge Interaction Network for Referring Image Segmentation |  |  |
| [CM-MaskSD] CM-MaskSD: Cross-Modality Masked Self-Distillation for Referring Image Segmentation |  | |
| [CARIS] CARIS: Context-Aware Referring Image Segmentation |  |  |
| [CVMN] Unsupervised Domain Adaptation for Referring Semantic Segmentation |  |  |
| [DIOR & MGVLF] RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data |  |  |
| [WiCo] WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image Segmentation |  | |
| [SLViT] SLViT: Scale-Wise Language-Guided Vision Transformer for Referring Image Segmentation |  |  |
| [JMCELN] Referring Image Segmentation via Joint Mask Contextual Embedding Learning and Progressive Alignment Network |  | |
| [TAS] Text Augmented Spatial-aware Zero-shot Referring Image Segmentation |  | |
| [SADLR] Semantics-Aware Dynamic Localization and Refinement for Referring Image Segmentation |  | |
| [MDSM] Multimodal Diffusion Segmentation Model for Object Segmentation from Manipulation Instructions |  | |
| [M3Att] Multi-Modal Mutual Attention and Iterative Interaction for Referring Image Segmentation |  | |
| [GraspNet-RIS] Towards Generalizable Referring Image Segmentation via Target Prompt and Visual Coherence |  | |
| [Omni-RES] Towards Omni-supervised Referring Expression Segmentation |  |  |
| [PCAN] Position-Aware Contrastive Alignment for Referring Image Segmentation |  | |
| [TSEG] Weakly-supervised segmentation of referring expressions |  | |
| [SHNet] Comprehensive Multi-Modal Interactions for Referring Image Segmentation |  |  |
| [GroupViT] GroupViT: Semantic Segmentation Emerges from Text Supervision |  |  |
| [CRIS] CRIS: CLIP-Driven Referring Image Segmentation |  |  |
| [ReSTR] ReSTR: Convolution-free Referring Image Segmentation Using Transformers |  |  |
| [LAVT] LAVT: Language-Aware Vision Transformer for Referring Image Segmentation |  |  |
| [CLIPSeg] Image Segmentation Using Text and Image Prompts |  |  |
| [CoupAlign] CoupAlign: Coupling Word-Pixel with Sentence-Mask Alignments for Referring Image Segmentation |  | |
| [SeqTR] SeqTR: A Simple yet Universal Network for Visual Grounding |  |  |
| [kesen2022modulating] Modulating Bottom-Up and Top-Down Visual Processing via Language-Conditional Filters |  |  |
| [GeoVG] Visual Grounding in Remote Sensing Images |  | Homepage |
| [feng2022learning] Learning From Box Annotations for Referring Image Segmentation |  |  |
| [ISF] Instance-Specific Feature Propagation for Referring Segmentation |  | |
| [PKS] Fully and Weakly Supervised Referring Expression Segmentation with End-to-End Learning |  | |
| [MaIL] MaIL: A Unified Mask-Image-Language Trimodal Network for Referring Image Segmentation |  | |
| [BUSNet] Bottom-Up Shift and Reasoning for Referring Image Segmentation |  |  |
| [LTS] Locate then Segment: A Strong Pipeline for Referring Image Segmentation |  | |
| [EFN] Encoder Fusion Network with Co-Attention Embedding for Referring Image Segmentation |  |  |
| [Referring Transformer] Referring Transformer: A One-step Approach to Multi-task Visual Grounding |  |  |
| [GbS] Detector-Free Weakly Supervised Grounding by Separation |  |  |
| [MDETR] MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding |  |  |
| [VLT] Vision-Language Transformer and Query Generation for Referring Segmentation |  |  |
| [CMPC] Cross-Modal Progressive Comprehension for Referring Segmentation |  |  |
| [TV-Net] Two-stage Visual Cues Enhancement Network for Referring Image Segmentation |  |  |
| [CMPC] Referring Image Segmentation via Cross-Modal Progressive Comprehension |  |  |
| [BRINet] Bi-directional Relationship Inferring Network for Referring Image Segmentation |  |  |
| [MCN] Multi-Task Collaborative Network for Joint Referring Expression Comprehension and Segmentation |  |  |
| [PhraseCut] PhraseCut: Language-based Image Segmentation in the Wild |  |  |
| [LSCM] Linguistic Structure Guided Context Modeling for Referring Image Segmentation |  |  |
| [ConvLSTM] Dual Convolutional LSTM Network for Referring Image Segmentation |  | |
| [CGAN] Cascade Grouped Attention Network for Referring Expression Segmentation |  | |
| [CMSA] Cross-Modal Self-Attention Network for Referring Image Segmentation |  |  |
| [CLEVR-Ref+] CLEVR-Ref+: Diagnosing Visual Reasoning with Referring Expressions |  |  |
| [NMTree] Learning to Assemble Neural Module Tree Networks for Visual Grounding |  |  |
| [STEP] See-Through-Text Grouping for Referring Image Segmentation |  | |
| [Lang2Seg] Referring Expression Object Segmentation with Caption-Aware Consistency |  |  |
| [MAttNet] MAttNet: Modular Attention Network for Referring Expression Comprehension |  |  |
| [RRN] Referring Image Segmentation via Recurrent Refinement Networks/document/8578700/) |  |  |
| [DMN] Dynamic Multimodal Instance Segmentation Guided by Natural Language Queries |  |  |
| [KWA] Key-Word-Aware Network for Referring Expression Image Segmentation |  |  |
| [CMN] Modeling Relationships in Referential Expressions with Compositional Modular Networks |  |  |
| [RMI] Recurrent Multimodal Interaction for Referring Image Segmentation |  |  |
| [MMI & G-Ref] Generation and Comprehension of Unambiguous Object Descriptions |  |  |
| [LSTM-CNN] Segmentation from Natural Language Expressions |  |  |
| [RefCOCO&RefCOCO+&RefCOCO(g)] Modeling Context in Referring Expressions |  |  |
| [ReferItGame] ReferItGame: Referring to Objects in Photographs of Natural Scenes |  |  |