Cross-modal hybrid feature fusion for image-sentence matching

Xing Xu; Yifan Wang; Yixuan He; Yang Yang; Alan Hanjalic; Heng Tao Shen

doi:10.1145/3458281

Cross-modal hybrid feature fusion for image-sentence matching

Xing Xu, Yifan Wang, Yixuan He, Yang Yang, Alan Hanjalic, Heng Tao Shen^*

^*Corresponding author for this work

Intelligent Systems

Research output: Contribution to journal › Article › Scientific › peer-review

6 Citations (Scopus)

Abstract

Image-sentence matching is a challenging task in the field of language and vision, which aims at measuring the similarities between images and sentence descriptions. Most existing methods independently map the global features of images and sentences into a common space to calculate the image-sentence similarity. However, the image-sentence similarity obtained by these methods may be coarse as (1) an intermediate common space is introduced to implicitly match the heterogeneous features of images and sentences in a global level, and (2) only the inter-modality relations of images and sentences are captured while the intra-modality relations are ignored. To overcome the limitations, we propose a novel Cross-Modal Hybrid Feature Fusion (CMHF) framework for directly learning the image-sentence similarity by fusing multimodal features with inter- and intra-modality relations incorporated. It can robustly capture the high-level interactions between visual regions in images and words in sentences, where flexible attention mechanisms are utilized to generate effective attention flows within and across the modalities of images and sentences. A structured objective with ranking loss constraint is formed in CMHF to learn the image-sentence similarity based on the fused fine-grained features of different modalities bypassing the usage of intermediate common space. Extensive experiments and comprehensive analysis performed on two widely used datasets - Microsoft COCO and Flickr30K - show the effectiveness of the hybrid feature fusion framework in CMHF, in which the state-of-the-art matching performance is achieved by our proposed CMHF method.

Original language	English
Article number	3458281
Journal	ACM Transactions on Multimedia Computing, Communications and Applications
Volume	17
Issue number	4
DOIs	https://doi.org/10.1145/3458281
Publication status	Published - 2021

Keywords

Attention mechanism
Cross-modal retrieval
Image-sentence matching
Multimodal feature fusion

Access to Document

10.1145/3458281

Cite this

@article{6eac60f57cfc446ea8b2187034deaf7d,

title = "Cross-modal hybrid feature fusion for image-sentence matching",

abstract = "Image-sentence matching is a challenging task in the field of language and vision, which aims at measuring the similarities between images and sentence descriptions. Most existing methods independently map the global features of images and sentences into a common space to calculate the image-sentence similarity. However, the image-sentence similarity obtained by these methods may be coarse as (1) an intermediate common space is introduced to implicitly match the heterogeneous features of images and sentences in a global level, and (2) only the inter-modality relations of images and sentences are captured while the intra-modality relations are ignored. To overcome the limitations, we propose a novel Cross-Modal Hybrid Feature Fusion (CMHF) framework for directly learning the image-sentence similarity by fusing multimodal features with inter- and intra-modality relations incorporated. It can robustly capture the high-level interactions between visual regions in images and words in sentences, where flexible attention mechanisms are utilized to generate effective attention flows within and across the modalities of images and sentences. A structured objective with ranking loss constraint is formed in CMHF to learn the image-sentence similarity based on the fused fine-grained features of different modalities bypassing the usage of intermediate common space. Extensive experiments and comprehensive analysis performed on two widely used datasets - Microsoft COCO and Flickr30K - show the effectiveness of the hybrid feature fusion framework in CMHF, in which the state-of-the-art matching performance is achieved by our proposed CMHF method.",

keywords = "Attention mechanism, Cross-modal retrieval, Image-sentence matching, Multimodal feature fusion",

author = "Xing Xu and Yifan Wang and Yixuan He and Yang Yang and Alan Hanjalic and Shen, {Heng Tao}",

year = "2021",

doi = "10.1145/3458281",

language = "English",

volume = "17",

journal = "ACM Transactions on Multimedia Computing, Communications and Applications",

issn = "1551-6857",

publisher = "Association for Computing Machinery (ACM)",

number = "4",

}

TY - JOUR

T1 - Cross-modal hybrid feature fusion for image-sentence matching

AU - Xu, Xing

AU - Wang, Yifan

AU - He, Yixuan

AU - Yang, Yang

AU - Hanjalic, Alan

AU - Shen, Heng Tao

PY - 2021

Y1 - 2021

N2 - Image-sentence matching is a challenging task in the field of language and vision, which aims at measuring the similarities between images and sentence descriptions. Most existing methods independently map the global features of images and sentences into a common space to calculate the image-sentence similarity. However, the image-sentence similarity obtained by these methods may be coarse as (1) an intermediate common space is introduced to implicitly match the heterogeneous features of images and sentences in a global level, and (2) only the inter-modality relations of images and sentences are captured while the intra-modality relations are ignored. To overcome the limitations, we propose a novel Cross-Modal Hybrid Feature Fusion (CMHF) framework for directly learning the image-sentence similarity by fusing multimodal features with inter- and intra-modality relations incorporated. It can robustly capture the high-level interactions between visual regions in images and words in sentences, where flexible attention mechanisms are utilized to generate effective attention flows within and across the modalities of images and sentences. A structured objective with ranking loss constraint is formed in CMHF to learn the image-sentence similarity based on the fused fine-grained features of different modalities bypassing the usage of intermediate common space. Extensive experiments and comprehensive analysis performed on two widely used datasets - Microsoft COCO and Flickr30K - show the effectiveness of the hybrid feature fusion framework in CMHF, in which the state-of-the-art matching performance is achieved by our proposed CMHF method.

AB - Image-sentence matching is a challenging task in the field of language and vision, which aims at measuring the similarities between images and sentence descriptions. Most existing methods independently map the global features of images and sentences into a common space to calculate the image-sentence similarity. However, the image-sentence similarity obtained by these methods may be coarse as (1) an intermediate common space is introduced to implicitly match the heterogeneous features of images and sentences in a global level, and (2) only the inter-modality relations of images and sentences are captured while the intra-modality relations are ignored. To overcome the limitations, we propose a novel Cross-Modal Hybrid Feature Fusion (CMHF) framework for directly learning the image-sentence similarity by fusing multimodal features with inter- and intra-modality relations incorporated. It can robustly capture the high-level interactions between visual regions in images and words in sentences, where flexible attention mechanisms are utilized to generate effective attention flows within and across the modalities of images and sentences. A structured objective with ranking loss constraint is formed in CMHF to learn the image-sentence similarity based on the fused fine-grained features of different modalities bypassing the usage of intermediate common space. Extensive experiments and comprehensive analysis performed on two widely used datasets - Microsoft COCO and Flickr30K - show the effectiveness of the hybrid feature fusion framework in CMHF, in which the state-of-the-art matching performance is achieved by our proposed CMHF method.

KW - Attention mechanism

KW - Cross-modal retrieval

KW - Image-sentence matching

KW - Multimodal feature fusion

UR - http://www.scopus.com/inward/record.url?scp=85123272823&partnerID=8YFLogxK

U2 - 10.1145/3458281

DO - 10.1145/3458281

M3 - Article

AN - SCOPUS:85123272823

SN - 1551-6857

VL - 17

JO - ACM Transactions on Multimedia Computing, Communications and Applications

JF - ACM Transactions on Multimedia Computing, Communications and Applications

IS - 4

M1 - 3458281

ER -

Cross-modal hybrid feature fusion for image-sentence matching

Abstract

Keywords

Access to Document

Other files and links

Fingerprint

Cite this