Outline

Ingegneria Sismica

Ingegneria Sismica

Multi-modal English Translation Production Model Combining Cross-modal Alignment and Attention Mechanism

Author(s): Shufang Wang1
1School of Foreign Languages, Zhengzhou Shengda University of Economics and Management, Zhengzhou 451191, Henan, China
Wang, Shufang. “Multi-modal English Translation Production Model Combining Cross-modal Alignment and Attention Mechanism.” Ingegneria Sismica Volume 43 Issue 3: 1-22, doi:10.65102/is20261051.

Abstract

Aiming at the problems of traditional English translation models, such as insufficient scene constraints, limited ambiguity resolution ability, and unstable image-text semantic coordination, this paper constructed a multimodal English translation production model combining cross-modal alignment and attention mechanism. Based on the collaborative input of text and image, the model forms an integrated technical link of “alignment-attention-generation” through multimodal input representation, shared semantic space mapping, bidirectional cross-modal alignment and attention-driven decoding generation. Experimental results show that the BLEU, METEOR and ROUGE-L of the proposed model on the test set reach 37.4, 32.5 and 41.3 respectively, which are 5.6, 4.4 and 5.9 percentage points higher than those of the basic Transformer model. The accuracy of image-text consistency, ambiguity resolution and entity alignment reaches 85.9%, 84.2% and 85.1%, respectively. The results show that cross-modal alignment can effectively reduce the representation deviation between text semantics and visual semantics, and the attention mechanism can enhance the dynamic screening ability of key contexts in the translation generation stage, thereby improving the accuracy, stability and application adaptability of multimodal English translation production.

Keywords
Multimodal English translation; Cross-modal alignment; Attention mechanism; Translation production model

Related Articles

Zhihao Jiang1,2, Limi Chen1,2, Jing Yang1
1Hainan Vocational University of Science and Technology, Haikou 571126, China
2Institute for Mathematical Research, Universiti Putra Malaysia, Serdang 43400, Malaysia
Limi Chen1,2, Zhihao Jiang1,2, Jing Yang1
1Hainan Vocational University of Science and Technology, Haikou 571126, China
2Institute for Mathematical Research, Universiti Putra Malaysia, Serdang 43400, Malaysia
Hui Yuan1, Minjie Chai2, Siqing Xu1, Jinsong Li1, Jinwan Zheng1
1Electric Power Research Institute, State Grid Shanxi Electric Power Co., Ltd., Taiyuan, 030001, Shanxi, China
2Jincheng Power Supply Branch, State Grid Shanxi Electric Power Co., Ltd., Jincheng, 048000, Shanxi, China
Yanhan Zhu1,2
1China Academy of Cultural Heritage, Chaoyang District, 100029, Beijing, China
2Beijing University of Civil Engineering and Architecture, Xicheng District, 100044, Beijing, China
Ken Wang1, Jinhan Shu2, Kan Yuan1
1School of Digital Media, Shenzhen Polytechnic University, Shenzhen 518055, Guangdong, China
2Postdoctoral Mobile Station of Journalism and communication, Fudan University, Shanghai 200433, Shanghai, China