CLUE: Contrastive language-guided learning for referring video object segmentation

Referring video object segmentation (R-VOS), the task of separating the object described by a natural language query from the video frames, has become increasingly critical with recent advances in multi-modal understanding. Existing approaches are mainly visual-dominant in both representation-learni...

Celý popis

Uložené v:

Podrobná bibliografia
Vydané v:	Pattern recognition letters Ročník 178; s. 115 - 121
Hlavní autori:	Gao, Qiqi, Zhong, Wanjun, Li, Jie, Zhao, Tiejun
Médium:	Journal Article
Jazyk:	English
Vydavateľské údaje:	Elsevier B.V 01.02.2024
Predmet:	Contrastive learning Deep learning Multi-modal Video object segmentation Video object segmentation Deep learning Multi-modal Contrastive learning
ISSN:	0167-8655, 1872-7344
On-line prístup:	Získať plný text
Tagy:	Pridať tag Žiadne tagy, Buďte prvý, kto otaguje tento záznam!

Buďte prvý, kto okomentuje tento záznam!