Open-Scenario Cross-Modal Retrieval
- Publication Type:
- Thesis
- Issue Date:
- 2025
Open Access
Copyright Clearance Process
- Recently Added
- In Progress
- Open Access
This item is open access.
Cross-modal retrieval enables searching across heterogeneous modalities such as images and text, and plays a critical role in large-scale multimedia information systems. However, real-world applications increasingly operate under open scenarios, where data are sparsely labeled, distributions shift across domains, semantic categories evolve over time, and data may arrive in a streaming manner with partially available modalities. These conditions fundamentally challenge conventional cross-modal hashing methods that assume static datasets, complete modality pairing, and sufficient supervision. This thesis investigates open-scenario cross-modal retrieval and proposes a series of cross-modal hashing frameworks to address its core challenges. First, a generative augmentation hashing approach is developed to improve retrieval performance under few-shot conditions by synthesizing semantically consistent multi-modal data. Second, a cross-domain transfer hashing framework is introduced to mitigate feature and label distribution shifts using semantic transfer technologies. Third, a prompt-infused continual hashing approach is proposed to accommodate incremental category evolution while preserving retrieval consistency. Finally, an online partial-modal hashing framework is presented to enable efficient real-time retrieval from streaming and incomplete data. Extensive experiments demonstrate that the proposed methods significantly enhance adaptability, robustness, and efficiency in dynamic cross-modal retrieval environments.
Please use this identifier to cite or link to this item:
