CLUE: Contrastive language-guided learning for referring video object segmentation
Referring video object segmentation (R-VOS), the task of separating the object described by a natural language query from the video frames, has become increasingly critical with recent advances in multi-modal understanding. Existing approaches are mainly visual-dominant in both representation-learni...
Saved in:
| Published in: | Pattern recognition letters Vol. 178; pp. 115 - 121 |
|---|---|
| Main Authors: | , , , |
| Format: | Journal Article |
| Language: | English |
| Published: |
Elsevier B.V
01.02.2024
|
| Subjects: | |
| ISSN: | 0167-8655, 1872-7344 |
| Online Access: | Get full text |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Be the first to leave a comment!