Cross-Lingual Multimodal Event Extraction: A Unified Framework for Parameter-Efficient Fine-Tuning
Article Information
Abstract
With the rapid development of multimodal large language models (MLLMs), structured event extraction (EE) has emerged as a critical intelligent information processing task, with increasing demand across multilingual and multimodal application scenarios. However, significant challenges remain in zero-shot multimodal and cross-language scenarios, including inconsistent cross-language outputs and the high computational cost of full-parameter fine-tuning. This study takes VideoLLaMA2 (VL2) and its improved version VL2.1 as the core models, and builds a multimodal annotated dataset covering English, Chinese, Spanish, and Russian (including 5,728 EE samples). It systematically evaluates the performance differences of zero-shot learning, and parameter-efficient fine-tuning (QLoRA) techniques. The experimental results show that for EE, QLoRA fine-tuning yields substantial gains across both models: VL2.1 achieves the highest trigger accuracy at 65.48%, while VL2 achieves the highest argument accuracy at 60.54%, with each figure representing the best performance obtained across the two fine-tuned models respectively. The study confirms that fine-tuning significantly enhances model robustness.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
Ethical Approval and Consent to Participate
References
- Mohammed, A., & Kora, R. (2025). A Comprehensive Overview and Analysis of Large Language Models: Trends and Challenges. IEEE Access, 13, 95851-95875.
[CrossRef] [Google Scholar] - Cheng, Z., Leng, S., Zhang, H., Xin, Y., Li, X., Chen, G., ... & Bing, L. (2024). Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms. arXiv preprint arXiv:2406.07476.
[Google Scholar] - Huang, K. H., Hsu, I. H., Natarajan, P., Chang, K. W., & Peng, N. (2022, May). Multilingual generative language models for zero-shot cross-lingual event argument extraction. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 4633-4646).
[CrossRef] [Google Scholar] - Lin, X. V., Mihaylov, T., Artetxe, M., Wang, T., Chen, S., Simig, D., ... & Li, X. (2022, December). Few-shot learning with multilingual generative language models. In Proceedings of the 2022 conference on empirical methods in natural language processing (pp. 9019-9052).
[CrossRef] [Google Scholar] - Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., ... & Stoyanov, V. (2019). Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116.
[Google Scholar] - Wadden, D., Wennberg, U., Luan, Y., & Hajishirzi, H. (2019). Entity, relation, and event extraction with contextualized span representations. arXiv preprint arXiv:1909.03546.
[Google Scholar] - Wu, J., Gan, W., Chen, Z., Wan, S., & Yu, P. S. (2023, December). Multimodal large language models: A survey. In 2023 IEEE International Conference on Big Data (BigData) (pp. 2247-2256). IEEE.
[CrossRef] [Google Scholar] - Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., ... & Le, Q. V. (2021). Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
[Google Scholar] - Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., ... & Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
[Google Scholar] - Holtzman, A., Buys, J., Du, L., Forbes, M., & Choi, Y. (2019). The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751.
[Google Scholar] - Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). Qlora: Efficient finetuning of quantized llms. Advances in neural information processing systems, 36, 10088-10115.
[Google Scholar] - Li, T., Wang, Z., Chai, L., Yang, J., Bai, J., Yin, Y., ... & Li, Z. (2024). Mt4crossoie: Multi-stage tuning for cross-lingual open information extraction. Expert Systems with Applications, 255, 124760.
[CrossRef] [Google Scholar] - Parthasarathy, V. B., Zafar, A., Khan, A., & Shahid, A. (2024). The ultimate guide to fine-tuning llms from basics to breakthroughs: An exhaustive review of technologies, research, best practices, applied research challenges and opportunities. arXiv preprint arXiv:2408.13296.
[Google Scholar] - Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., & Yuan, L. (2023). Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122.
[Google Scholar] - Zhang, H., Li, X., & Bing, L. (2023). Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858.
[Google Scholar] - Yang, A., et al. (2024). Qwen2 technical report. arXiv. https://arxiv.org/abs/2407.10671
[Google Scholar] - Chirkova, N., & Nikoulina, V. (2024). Zero-shot cross-lingual transfer in instruction tuning of large language models. arXiv preprint arXiv:2402.14778.
[Google Scholar] - Zhang, Q., Chen, M., Bukharin, A., Karampatziakis, N., He, P., Cheng, Y., ... & Zhao, T. (2023). Adalora: Adaptive budget allocation for parameter-efficient fine-tuning. arXiv preprint arXiv:2303.10512.
[CrossRef] [Google Scholar] - Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., ... & Chen, W. (2022). Lora: Low-rank adaptation of large language models. ICLR, 1(2), 3.
[Google Scholar] - Xiang, W., & Wang, B. (2019). A survey of event extraction from text. IEEE Access, 7, 173111-173137.
[CrossRef] [Google Scholar] - Marchisio, K., Ko, W. Y., Bérard, A., Dehaze, T., & Ruder, S. (2024, November). Understanding and mitigating language confusion in LLMs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 6653-6677).
[CrossRef] [Google Scholar] - Song, Z., Bies, A., Strassel, S., Riese, T., Mott, J., Ellis, J., ... & Ma, X. (2015, June). From light to rich ERE: Annotation of entities, relations, and events. In Proceedings of the 3rd workshop on EVENTS: Definition, detection, coreference, and representation (pp. 89-98).
[CrossRef] [Google Scholar] - Siriborvornratanakul, T. (2025, May). From Human Annotators to AI: The Transition and the Role of Synthetic Data in AI Development. In International Conference on Human-Computer Interaction (pp. 379-390). Cham: Springer Nature Switzerland.
[CrossRef] [Google Scholar] - Kamesh, R. (2024). Think Beyond Size: Adaptive Prompting for More Effective Reasoning. arXiv preprint arXiv:2410.08130.
[Google Scholar] - Zhang, Z., Zhang, A., Li, M., Zhao, H., Karypis, G., & Smola, A. (2023). Multimodal chain-of-thought reasoning in language models. arXiv preprint arXiv:2302.00923.
[Google Scholar] - Zhang, X., Wang, Z., & Li, P. (2023, June). Multimodal Chinese Event Extraction on Text and Audio. In 2023 International Joint Conference on Neural Networks (IJCNN) (pp. 1-8). IEEE.
[CrossRef] [Google Scholar] - Butcher, B., Zilka, M., Hron, J., Cook, D., & Weller, A. (2024). Optimising human-machine collaboration for efficient high-precision information extraction from text documents. ACM Journal on Responsible Computing, 1(2), 1-27.
[CrossRef] [Google Scholar]
Cited By (4)
-
Le Zou, Mengyu Ma, Jun Li, Hao Chen, Shuang Peng. Medical Vision-Language Models: Existing Technologies, Clinical Applications and Future Directions.
Sensors, 2026 , 26 (13).
[CrossRef] -
Leiquan Wang, Guixiang Lou, Xin Li, Xueqing Yang, Chunlei Wu, Zhongwei Li. Integrating global and local features for multisource remote sensing image classification with 2D Selective Scan Fusion.
International Journal of Remote Sensing, 2026 , 47 (13).
[CrossRef] -
Guixin Li, Bingjin Zhou, Minjian Ni, Huiying Chen, Yiwei Liu, Yinghua Zhang, Man Zhang, Minjuan Wang. MMIU-Net: an encoder–decoder architecture based on multimodal feature fusion for wheat yield prediction under drought stress.
International Journal of Remote Sensing, 2026 , 47 (13).
[CrossRef] -
Geng Gao, Tao Gan, Jinlin Yang, Nini Rao. Multimodal information compression, completion, and adaptive fusion for cancer survival prediction.
Information Fusion, 2026 , 136 .
[CrossRef]
Cite This Article
TY - JOUR AU - Hong, Sheng AU - Wang, Xuanqi AU - Mei, Zeyu AU - Wickramaratne, Thisura Bojitha PY - 2025 DA - 2025/10/04 TI - Cross-Lingual Multimodal Event Extraction: A Unified Framework for Parameter-Efficient Fine-Tuning JO - ICCK Transactions on Intelligent Systematics T2 - ICCK Transactions on Intelligent Systematics JF - ICCK Transactions on Intelligent Systematics VL - 2 IS - 4 SP - 203 EP - 212 DO - 10.62762/TIS.2025.610574 UR - https://www.icck.org/article/abs/TIS.2025.610574 KW - event extraction KW - QLoRA KW - multimodal LLMs KW - multilingual NLP AB - With the rapid development of multimodal large language models (MLLMs), structured event extraction (EE) has emerged as a critical intelligent information processing task, with increasing demand across multilingual and multimodal application scenarios. However, significant challenges remain in zero-shot multimodal and cross-language scenarios, including inconsistent cross-language outputs and the high computational cost of full-parameter fine-tuning. This study takes VideoLLaMA2 (VL2) and its improved version VL2.1 as the core models, and builds a multimodal annotated dataset covering English, Chinese, Spanish, and Russian (including 5,728 EE samples). It systematically evaluates the performance differences of zero-shot learning, and parameter-efficient fine-tuning (QLoRA) techniques. The experimental results show that for EE, QLoRA fine-tuning yields substantial gains across both models: VL2.1 achieves the highest trigger accuracy at 65.48%, while VL2 achieves the highest argument accuracy at 60.54%, with each figure representing the best performance obtained across the two fine-tuned models respectively. The study confirms that fine-tuning significantly enhances model robustness. SN - 3068-5079 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Hong2025CrossLingu,
author = {Sheng Hong and Xuanqi Wang and Zeyu Mei and Thisura Bojitha Wickramaratne},
title = {Cross-Lingual Multimodal Event Extraction: A Unified Framework for Parameter-Efficient Fine-Tuning},
journal = {ICCK Transactions on Intelligent Systematics},
year = {2025},
volume = {2},
number = {4},
pages = {203-212},
doi = {10.62762/TIS.2025.610574},
url = {https://www.icck.org/article/abs/TIS.2025.610574},
abstract = {With the rapid development of multimodal large language models (MLLMs), structured event extraction (EE) has emerged as a critical intelligent information processing task, with increasing demand across multilingual and multimodal application scenarios. However, significant challenges remain in zero-shot multimodal and cross-language scenarios, including inconsistent cross-language outputs and the high computational cost of full-parameter fine-tuning. This study takes VideoLLaMA2 (VL2) and its improved version VL2.1 as the core models, and builds a multimodal annotated dataset covering English, Chinese, Spanish, and Russian (including 5,728 EE samples). It systematically evaluates the performance differences of zero-shot learning, and parameter-efficient fine-tuning (QLoRA) techniques. The experimental results show that for EE, QLoRA fine-tuning yields substantial gains across both models: VL2.1 achieves the highest trigger accuracy at 65.48\%, while VL2 achieves the highest argument accuracy at 60.54\%, with each figure representing the best performance obtained across the two fine-tuned models respectively. The study confirms that fine-tuning significantly enhances model robustness.},
keywords = {event extraction, QLoRA, multimodal LLMs, multilingual NLP},
issn = {3068-5079},
publisher = {Institute of Central Computation and Knowledge}
}
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Portico