Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

43 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

【CAC'2025 🔥】 Text-Video Retrieval With Global-Local Contrastive Consistency Learning

Build Status Build Status Build Status

This is the implementation of CAC 2025 paper Text-Video Retrieval With Global-Local Contrastive Consistency Learning.

📌 Citation

If you find our method useful in your work, please cite:

@INPROCEEDINGS{11487249,
  author={Jing, Xiaolun and Yang, Xinxing and Yang, Genke},
  booktitle={2025 China Automation Congress (CAC)}, 
  title={Text-Video Retrieval With Global-Local Contrastive Consistency Learning}, 
  year={2025},
  volume={},
  number={},
  pages={1621-1626},
  keywords={Earth Observing System;Artificial satellites;Motion pictures;Broadcasting;Text to video;Videos;Video equipment;Protocols;HTTP;Location awareness;Global-Local Interaction;Contrastive Consistency Learning;Text-Video Retrieval},
  doi={10.1109/CAC67268.2025.11487249}}

📕 Overview

Text-video retrieval aims to find the most semantically similar videos with given text queries. However, since videos contain more diverse content than texts, the main semantics expressed by each text-video pair is often partially relevant. The primary methods involve the utilization of language-video attention module to better align texts and videos. Though effective, this paradigm inevitably introduces prohibitive computational overhead, resulting in inefficient retrieval. In this paper, we propose a simple yet effective method called Global-Local Contrastive Consistency Learning (GLCCL) to achieve texts and videos semantics alignment. Specifically, we design a parameter-free Global-Local Interaction Module (GLIM) to generate semantic-related frame and video features in a text-guided manner. Furthermore, we devise an auxiliary Contrastive Score Consistency (CSC) loss to promote consistency learning among different scores on positive pairs and suppress consistency learning on negative pairs. Extensive experiments on three benchmark datasets demonstrate the superiority and effectiveness of our approach, including MSR-VTT, DiDeMo and VATEX.

📚 Method

image

🚀 Quick Start

Setup code environment

pip install -r requirements.txt

Datasets

We train our model on MSR-VTT, DiDeMo and VATEX datasets respectively. Please refer to this repo for data preparation.

Datasets Google Cloud Baidu Yun Peking University Yun
MSR-VTT Download Download Download
DiDeMo TODO Download Download
VATEX TODO TODO TODO

Download CLIP Model

wget -P ./modules https://openaipublic.azureedge.net/clip/models/40d365715913c9da98579312b702a82c18be219cc2a73407c4526f58eba950af/ViT-B-32.pt
wget -P ./modules https://openaipublic.azureedge.net/clip/models/5806e77cd80f8b59890b7e101eabd078d9fb84e6937f9e85e4ecb61988df416f/ViT-B-16.pt

Compress Video

python preprocess/compress_video.py --input_root [raw_video_path] --output_root [compressed_video_path]

Train on MSR-VTT

python -m torch.distributed.launch --nproc_per_node=4 --master_port='30400' \
main_glccl.py --do_train --num_thread_reader=8 \
--epochs=5 --batch_size=128 --batch_size_val 64 --n_display=50 \
--train_csv ${FILE_DATA_PATH}/MSRVTT_train.9k.csv \
--val_csv ../DataSet/MSRVTT/data/file/MSRVTT_JSFUSION_test.csv \
--data_path ../DataSet/MSRVTT/data/file/MSRVTT_data.json \
--features_path ../DataSet/MSRVTT/data/file/clip4clip_video_frame_input \
--output_dir ../Model/glccl_msrvtt_vit32 \
--log_dir ../Log/glccl_msrvtt_vit32 \
--visualize_dir ../Visualize/glccl_msrvtt_vit32 \
--lr 1e-4 --max_words 32 --max_frames 12 \
--datatype msrvtt --expand_msrvtt_sentences \
--feature_framerate 1 --coef_lr 1e-3 \
--freeze_layer_num 0  --slice_framepos 2 \ 
--text_guided_flag --aggregation_weights_type softmax \
--csc_loss_flag --csc_loss_weight 0.5 \
--loose_type --linear_patch 2d --sim_header seqTransf \
--pretrained_clip_name ViT-B/32

Train on DiDeMo

python -m torch.distributed.launch --nproc_per_node=8 --master_port='30401' \
main_glccl.py --do_train --num_thread_reader=8 \
--epochs=20 --batch_size=64 --batch_size_val 32 --n_display=50 \
--data_path ../DataSet/DiDeMo/data/compressed/split_file \
--features_path ../DataSet/DiDeMo/data/compressed/split_video \
--output_dir ../Model/glccl_didemo_vit32 \
--log_dir ../Log/glccl_didemo_vit32 \
--visualize_dir ../Visualize/glccl_didemo_vit32 \
--lr 1e-4 --max_words 64 --max_frames 64 \
--datatype didemo \
--feature_framerate 1 --coef_lr 1e-3 \
--freeze_layer_num 0  --slice_framepos 2 \
--text_guided_flag --aggregation_weights_type softmax \
--csc_loss_flag --csc_loss_weight 0.5 \
--loose_type --linear_patch 2d --sim_header seqTransf \
--pretrained_clip_name ViT-B/32

Train on VATEX

python -m torch.distributed.launch --nproc_per_node=4 --master_port='30402' \
main_glccl.py --do_train --num_thread_reader=8 \
--epochs=5 --batch_size=128 --batch_size_val 128 --n_display=50 \
--data_path ../DataSet/VATEX/data/compressed/split_file \
--features_path ../DataSet/VATEX/data/compressed/clip4clip_video_frame_input \
--output_dir ../Model/glccl_vatex_vit32 \
--log_dir ../Log/glccl_vatex_vit32 \
--visualize_dir ../Visualize/glccl_vatex_vit32 \
--lr 1e-4 --max_words 32 --max_frames 12 \
--datatype vatex \
--feature_framerate 1 --coef_lr 1e-3 \
--freeze_layer_num 0  --slice_framepos 2 \
--text_guided_flag --aggregation_weights_type softmax \
--csc_loss_flag --csc_loss_weight 0.5 \
--loose_type --linear_patch 2d --sim_header seqTransf \
--pretrained_clip_name ViT-B/32

🎗️ Acknowledgments

The implementation of GLCCL relies on resources from CLIP4Clip and CLIP. We thank the original authors for their open-sourcing.

About

[CAC'2025] Text-Video Retrieval With Global-Local Contrastive Consistency Learning

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors

Languages