Cross-Lingual Multimodal Event Extraction: A Unified Framework for Parameter-Efficient Fine-Tuning
Research Article  ·  Published: 04 October 2025
Issue cover
ICCK Transactions on Intelligent Systematics
Volume 2, Issue 4, 2025: 203-212
Research Article Free to Read

Cross-Lingual Multimodal Event Extraction: A Unified Framework for Parameter-Efficient Fine-Tuning

1 School of Cyber Science and Technology, Beihang University, Beijing 100191, China
2 Nanchang University, Nanchang 330031, China
3 School of Information Engineering, Nanchang University, Nanchang 330031, China
4 International Business School, Beijing Foreign Studies University, Beijing 100089, China
5 School of Cyber Science and Technology, Beihang University, Beijing 100191, China
* Corresponding Author: Sheng Hong, [email protected]
Volume 2, Issue 4

Article Information

Abstract

With the rapid development of multimodal large language models (MLLMs), structured event extraction (EE) has emerged as a critical intelligent information processing task, with increasing demand across multilingual and multimodal application scenarios. However, significant challenges remain in zero-shot multimodal and cross-language scenarios, including inconsistent cross-language outputs and the high computational cost of full-parameter fine-tuning. This study takes VideoLLaMA2 (VL2) and its improved version VL2.1 as the core models, and builds a multimodal annotated dataset covering English, Chinese, Spanish, and Russian (including 5,728 EE samples). It systematically evaluates the performance differences of zero-shot learning, and parameter-efficient fine-tuning (QLoRA) techniques. The experimental results show that for EE, QLoRA fine-tuning yields substantial gains across both models: VL2.1 achieves the highest trigger accuracy at 65.48%, while VL2 achieves the highest argument accuracy at 60.54%, with each figure representing the best performance obtained across the two fine-tuned models respectively. The study confirms that fine-tuning significantly enhances model robustness.

Graphical Abstract

Cross-Lingual Multimodal Event Extraction: A Unified Framework for Parameter-Efficient Fine-Tuning

Keywords

event extraction QLoRA multimodal LLMs multilingual NLP

Data Availability Statement

Data will be made available on request.

Funding

This work was supported by the National Key Research and Development Program under Grant 2022YFB3103602.

Conflicts of Interest

Sheng Hong served as an Associate Editor of the ICCK Transactions on Intelligent Systematics at the time of manuscript submission. To ensure the integrity of the peer-review process, Sheng Hong was not involved in the editorial handling, peer review, or decision-making process for this manuscript, which was handled independently by another editor. The remaining authors declare no conflicts of interest.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Mohammed, A., & Kora, R. (2025). A Comprehensive Overview and Analysis of Large Language Models: Trends and Challenges. IEEE Access, 13, 95851-95875.
    [CrossRef] [Google Scholar]
  2. Cheng, Z., Leng, S., Zhang, H., Xin, Y., Li, X., Chen, G., ... & Bing, L. (2024). Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms. arXiv preprint arXiv:2406.07476.
    [Google Scholar]
  3. Huang, K. H., Hsu, I. H., Natarajan, P., Chang, K. W., & Peng, N. (2022, May). Multilingual generative language models for zero-shot cross-lingual event argument extraction. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 4633-4646).
    [CrossRef] [Google Scholar]
  4. Lin, X. V., Mihaylov, T., Artetxe, M., Wang, T., Chen, S., Simig, D., ... & Li, X. (2022, December). Few-shot learning with multilingual generative language models. In Proceedings of the 2022 conference on empirical methods in natural language processing (pp. 9019-9052).
    [CrossRef] [Google Scholar]
  5. Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., ... & Stoyanov, V. (2019). Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116.
    [Google Scholar]
  6. Wadden, D., Wennberg, U., Luan, Y., & Hajishirzi, H. (2019). Entity, relation, and event extraction with contextualized span representations. arXiv preprint arXiv:1909.03546.
    [Google Scholar]
  7. Wu, J., Gan, W., Chen, Z., Wan, S., & Yu, P. S. (2023, December). Multimodal large language models: A survey. In 2023 IEEE International Conference on Big Data (BigData) (pp. 2247-2256). IEEE.
    [CrossRef] [Google Scholar]
  8. Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., ... & Le, Q. V. (2021). Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
    [Google Scholar]
  9. Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., ... & Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
    [Google Scholar]
  10. Holtzman, A., Buys, J., Du, L., Forbes, M., & Choi, Y. (2019). The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751.
    [Google Scholar]
  11. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). Qlora: Efficient finetuning of quantized llms. Advances in neural information processing systems, 36, 10088-10115.
    [Google Scholar]
  12. Li, T., Wang, Z., Chai, L., Yang, J., Bai, J., Yin, Y., ... & Li, Z. (2024). Mt4crossoie: Multi-stage tuning for cross-lingual open information extraction. Expert Systems with Applications, 255, 124760.
    [CrossRef] [Google Scholar]
  13. Parthasarathy, V. B., Zafar, A., Khan, A., & Shahid, A. (2024). The ultimate guide to fine-tuning llms from basics to breakthroughs: An exhaustive review of technologies, research, best practices, applied research challenges and opportunities. arXiv preprint arXiv:2408.13296.
    [Google Scholar]
  14. Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., & Yuan, L. (2023). Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122.
    [Google Scholar]
  15. Zhang, H., Li, X., & Bing, L. (2023). Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858.
    [Google Scholar]
  16. Yang, A., et al. (2024). Qwen2 technical report. arXiv. https://arxiv.org/abs/2407.10671
    [Google Scholar]
  17. Chirkova, N., & Nikoulina, V. (2024). Zero-shot cross-lingual transfer in instruction tuning of large language models. arXiv preprint arXiv:2402.14778.
    [Google Scholar]
  18. Zhang, Q., Chen, M., Bukharin, A., Karampatziakis, N., He, P., Cheng, Y., ... & Zhao, T. (2023). Adalora: Adaptive budget allocation for parameter-efficient fine-tuning. arXiv preprint arXiv:2303.10512.
    [CrossRef] [Google Scholar]
  19. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., ... & Chen, W. (2022). Lora: Low-rank adaptation of large language models. ICLR, 1(2), 3.
    [Google Scholar]
  20. Xiang, W., & Wang, B. (2019). A survey of event extraction from text. IEEE Access, 7, 173111-173137.
    [CrossRef] [Google Scholar]
  21. Marchisio, K., Ko, W. Y., Bérard, A., Dehaze, T., & Ruder, S. (2024, November). Understanding and mitigating language confusion in LLMs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 6653-6677).
    [CrossRef] [Google Scholar]
  22. Song, Z., Bies, A., Strassel, S., Riese, T., Mott, J., Ellis, J., ... & Ma, X. (2015, June). From light to rich ERE: Annotation of entities, relations, and events. In Proceedings of the 3rd workshop on EVENTS: Definition, detection, coreference, and representation (pp. 89-98).
    [CrossRef] [Google Scholar]
  23. Siriborvornratanakul, T. (2025, May). From Human Annotators to AI: The Transition and the Role of Synthetic Data in AI Development. In International Conference on Human-Computer Interaction (pp. 379-390). Cham: Springer Nature Switzerland.
    [CrossRef] [Google Scholar]
  24. Kamesh, R. (2024). Think Beyond Size: Adaptive Prompting for More Effective Reasoning. arXiv preprint arXiv:2410.08130.
    [Google Scholar]
  25. Zhang, Z., Zhang, A., Li, M., Zhao, H., Karypis, G., & Smola, A. (2023). Multimodal chain-of-thought reasoning in language models. arXiv preprint arXiv:2302.00923.
    [Google Scholar]
  26. Zhang, X., Wang, Z., & Li, P. (2023, June). Multimodal Chinese Event Extraction on Text and Audio. In 2023 International Joint Conference on Neural Networks (IJCNN) (pp. 1-8). IEEE.
    [CrossRef] [Google Scholar]
  27. Butcher, B., Zilka, M., Hron, J., Cook, D., & Weller, A. (2024). Optimising human-machine collaboration for efficient high-precision information extraction from text documents. ACM Journal on Responsible Computing, 1(2), 1-27.
    [CrossRef] [Google Scholar]

Cited By (4)

  1. Le Zou, Mengyu Ma, Jun Li, Hao Chen, Shuang Peng. Medical Vision-Language Models: Existing Technologies, Clinical Applications and Future Directions. Sensors, 2026 , 26 (13).
    [CrossRef]
  2. Leiquan Wang, Guixiang Lou, Xin Li, Xueqing Yang, Chunlei Wu, Zhongwei Li. Integrating global and local features for multisource remote sensing image classification with 2D Selective Scan Fusion. International Journal of Remote Sensing, 2026 , 47 (13).
    [CrossRef]
  3. Guixin Li, Bingjin Zhou, Minjian Ni, Huiying Chen, Yiwei Liu, Yinghua Zhang, Man Zhang, Minjuan Wang. MMIU-Net: an encoder–decoder architecture based on multimodal feature fusion for wheat yield prediction under drought stress. International Journal of Remote Sensing, 2026 , 47 (13).
    [CrossRef]
  4. Geng Gao, Tao Gan, Jinlin Yang, Nini Rao. Multimodal information compression, completion, and adaptive fusion for cancer survival prediction. Information Fusion, 2026 , 136 .
    [CrossRef]
* Citation data provided by Crossref Cited-by.

Cite This Article

APA Style
Hong, S., Wang, X., Mei, Z., & Wickramaratne, T. B. (2025). Cross-Lingual Multimodal Event Extraction: A Unified Framework for Parameter-Efficient Fine-Tuning. ICCK Transactions on Intelligent Systematics, 2(4), 203-212. https://doi.org/10.62762/TIS.2025.610574
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Hong, Sheng
AU  - Wang, Xuanqi
AU  - Mei, Zeyu
AU  - Wickramaratne, Thisura Bojitha
PY  - 2025
DA  - 2025/10/04
TI  - Cross-Lingual Multimodal Event Extraction: A Unified Framework for Parameter-Efficient Fine-Tuning
JO  - ICCK Transactions on Intelligent Systematics
T2  - ICCK Transactions on Intelligent Systematics
JF  - ICCK Transactions on Intelligent Systematics
VL  - 2
IS  - 4
SP  - 203
EP  - 212
DO  - 10.62762/TIS.2025.610574
UR  - https://www.icck.org/article/abs/TIS.2025.610574
KW  - event extraction
KW  - QLoRA
KW  - multimodal LLMs
KW  - multilingual NLP
AB  - With the rapid development of multimodal large language models (MLLMs), structured event extraction (EE) has emerged as a critical intelligent information processing task, with increasing demand across multilingual and multimodal application scenarios. However, significant challenges remain in zero-shot multimodal and cross-language scenarios, including inconsistent cross-language outputs and the high computational cost of full-parameter fine-tuning. This study takes VideoLLaMA2 (VL2) and its improved version VL2.1 as the core models, and builds a multimodal annotated dataset covering English, Chinese, Spanish, and Russian (including 5,728 EE samples). It systematically evaluates the performance differences of zero-shot learning, and parameter-efficient fine-tuning (QLoRA) techniques. The experimental results show that for EE, QLoRA fine-tuning yields substantial gains across both models: VL2.1 achieves the highest trigger accuracy at 65.48%, while VL2 achieves the highest argument accuracy at 60.54%, with each figure representing the best performance obtained across the two fine-tuned models respectively. The study confirms that fine-tuning significantly enhances model robustness.
SN  - 3068-5079
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Hong2025CrossLingu,
  author = {Sheng Hong and Xuanqi Wang and Zeyu Mei and Thisura Bojitha Wickramaratne},
  title = {Cross-Lingual Multimodal Event Extraction: A Unified Framework for Parameter-Efficient Fine-Tuning},
  journal = {ICCK Transactions on Intelligent Systematics},
  year = {2025},
  volume = {2},
  number = {4},
  pages = {203-212},
  doi = {10.62762/TIS.2025.610574},
  url = {https://www.icck.org/article/abs/TIS.2025.610574},
  abstract = {With the rapid development of multimodal large language models (MLLMs), structured event extraction (EE) has emerged as a critical intelligent information processing task, with increasing demand across multilingual and multimodal application scenarios. However, significant challenges remain in zero-shot multimodal and cross-language scenarios, including inconsistent cross-language outputs and the high computational cost of full-parameter fine-tuning. This study takes VideoLLaMA2 (VL2) and its improved version VL2.1 as the core models, and builds a multimodal annotated dataset covering English, Chinese, Spanish, and Russian (including 5,728 EE samples). It systematically evaluates the performance differences of zero-shot learning, and parameter-efficient fine-tuning (QLoRA) techniques. The experimental results show that for EE, QLoRA fine-tuning yields substantial gains across both models: VL2.1 achieves the highest trigger accuracy at 65.48\%, while VL2 achieves the highest argument accuracy at 60.54\%, with each figure representing the best performance obtained across the two fine-tuned models respectively. The study confirms that fine-tuning significantly enhances model robustness.},
  keywords = {event extraction, QLoRA, multimodal LLMs, multilingual NLP},
  issn = {3068-5079},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Views
2090
PDF Downloads
800

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

Institute of Central Computation and Knowledge (ICCK) or its licensor holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
ICCK Transactions on Intelligent Systematics
ICCK Transactions on Intelligent Systematics
ISSN: 3068-5079 (Online) | ISSN: 3069-003X (Print)
Portico
Preserved at
Portico