SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detection (CVPR 2026)

April 14, 2026 ยท View on GitHub

License: MIT Paper Venue: CVPR 2026 Dataset Dataset

๐Ÿšง Update:

  • [2026/04/13]: The OVCOD-D benchmark dataset is released on Hugging Face and ModelScope!
  • [2026/03/25]: Repository created. Our paper has been accepted by CVPR 2026! [ArXiv Paper]
  • [Coming Soon]: The training and inference code will be available soon. Stay tuned!

๐Ÿ“– Introduction

Open-vocabulary object detection (OVOD) aims to detect known and unknown objects in the open world. However, camouflaged objects pose significant challenges due to high visual similarity with the background. We propose SDDF, which leverages specificity-aware sub-descriptions and a dynamic focusing mechanism to enhance the detector's discrimination capability.

๐Ÿ“Š OVCOD-D Benchmark

Hugging Face

Benchmark Pipeline Figure: Construction pipeline of OVCOD-D dataset. We extend COD10K-D, NC4K-D, and cleaned CAMO-D with YOLO-style detection labels and an additional red imported fire ant nest subset, then reorganize them into 40 base and 47 novel classes. Qwen3-VL-Plus generates fine-grained image descriptions from which we derive a semantic prompt library for open-vocabulary camouflaged object detection.

โš™๏ธ Architecture

SDDF Framework Figure: Overall architecture of the proposed specificity-driven open-vocabulary camouflaged object detector.

๐Ÿ† Main Results

1. Comparison with Open-Vocabulary Object Detectors

Evaluated on the union of base and novel classes on the OVCOD-D dataset.

MethodBackboneParamsPre-trainAPAP50AP75APmAPl
Grounding DINO-TSwin-T172MO365, GoldG34.843.937.7--
YOLOE-MYOLOv8-M94MO365, GoldG39.947.742.7--
YOLO-World-LYOLOv8-L110MO365, GoldG45.763.248.922.948.4
DOSOD-LYOLOv8-L108MO365, GoldG53.473.156.226.456.3
SDDF-L (Ours)YOLOv8-L109MO365, GoldG56.476.460.734.459.0

2. Comparison with SOTA COD Methods

Comparison with State-of-the-Art Camouflaged Object Detection methods.

MethodBackboneAPAP50AP75
SINet-V2ResNet-5040.269.339.4
FSPNetSwin-T47.976.249.4
CamoFormerSwin-T55.680.259.0
HDPNetViT-B56.381.559.6
SDDF-L (Ours)YOLOv8-L56.476.460.7

๐Ÿ–ผ๏ธ Qualitative Results

Qualitative Results Figure: Visualization of detection bounding boxes and heatmap representations.

๐Ÿ“ง Contact

If you have any questions, please feel free to contact us or open an issue.

๐Ÿ“ Citation

If you find our work helpful for your research, please consider citing:

@article{liang2026sddf,
  title={SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detection},
  author={Liang, Jiaming and Zhan, Yifeng and Liu, Chunlin and Zheng, Weihua and Peng, Bingye env, Liang, Qiwei and Cai, Boyang and Mai, Xiaochun and Nie, Qiang},
  journal={arXiv preprint arXiv:2603.26109},
  year={2026}
}