Self-Chained Dynamic Context Perception to Tracking by Natural Language Specification
{{output}}
Vision-language cross-modal learning has significantly improved Tracking by Natural Language specification (TNL). Most existing TNL methods follow a Siamese-like matching paradigm, where visual search-region features and language-query features are aligned wit... ...