SMALL LANGUAGE MODELS FOR MULTILINGUAL DISINFORMATION RETRIEVAL
DOI:
https://doi.org/10.18372/2310-5461.71.21448Keywords:
artificial intelligence, RAG, disinformation, small language models, large language models, multilingual retrieval, narrative, semantic retrieval, reranking, information operationsAbstract
Disinformation campaigns reuse known narratives while changing the language, actors, and publication context. Such variation makes it difficult for analysts to connect new material with cases already documented by fact-checkers. To address this task, a software application was developed for multilingual narrative normalization and the retrieval of documented debunks from the EUvsDisinfo corpus. The retrieval corpus contains 7,538 fact-checking records published between 2015 and 2026. Each record links a disinformation claim to its debunk or a contextual explanation. The experiment covered materials concerning Russia’s war against Ukraine, Russia–NATO relations, Syria, and COVID-19. Across these data, the same semantic patterns recur in different languages and event contexts. The language model converts an input text into a concise narrative query. It retains the main claim and the source’s position while removing names, dates, and other contextual details. The system uses the resulting query for semantic retrieval and corpus record reranking. This module can operate as part of a disinformation-monitoring system: it connects new material with known narratives and passes the retrieved debunks to an evidence-verification component. The study compares the capabilities of the commercial large language model Gemini Flash 3.5 with the locally deployed Qwen3-4B and MamayLM-Gemma-3-4B models. The results confirm that small language models can perform narrative normalization in a multilingual retrieval system with a moderate reduction in quality compared with the commercial model. Local deployment provides control over data and enables fine-tuning to improve model quality and stability.
References
Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” Advances in Neural Information Processing Systems, vol. 33, 2020, https://doi.org/10.48550/arXiv.2005.11401.
P. Santra, M. Ghosh, D. Ganguly, P. Basuchowdhuri, and S. K. Naskar, “The curious case of contexts in retrieval-augmented generation with a combination of labeled and unlabeled data,” WIREs Data Mining and Knowledge Discovery, vol. 15, no. 2, 2025, https://doi.org/10.1002/widm.70021.
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal, “FEVER: A large-scale dataset for fact extraction and verification,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2018, https://doi.org/ 10.18653/v1/N18-1074.
S. Shaar, N. Babulkov, G. Da San Martino, and P. Nakov, “That is a known lie: Detecting previously fact-checked claims,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, https://doi.org/ 10.18653/v1/2020.acl-main.332.
M. Pikuliak, I. Srba, R. Moro, T. Hromadka,
T. Smolen, M. Melisek, I. Vykopal, J. Simko, and M. Bielikova, “Multilingual previously fact-checked claim retrieval,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, https://doi.org/ 10.18653/v1/2023.emnlp-main.1027.
H. Peng et al., “SemEval-2025 Task 7: Multilin-gual and Crosslingual Fact-Checked Claim Retrieval,” in Proceedings of the 19th International Workshop on Semantic Evaluation, 2025, https://aclanthology.org/2025.semeval-1.323/.
G. Da San Martino, A. Barron-Cedeno, H. Wachsmuth, R. Petrov, and P. Nakov, “SemEval-2020 Task 11: Detection of propaganda techniques in news articles,” in Proceedings of the Fourteenth Workshop on Semantic Evaluation, 2020, https://doi.org/10.18653/v1/2020.semeval-1.186.
J. Piskorski, N. Stefanovitch, G. Da San Martino, and P. Nakov, “SemEval-2023 Task 3: Detecting the category, the framing, and the persuasion techniques in online news in a multilingual setup,” in Proceedings of the 17th International Workshop on Semantic Evaluation, 2023, https://doi.org/ 10.18653/v1/2023.semeval-1.317.
K. L. Anglin, A. Bertrand, J. Gottlieb, and J. Elefante, “Scaling Up With Integrity: Valid and Efficient Narrative Policy Framework Analyses Using Large Language Models,” Policy Studies Journal, vol. 54, no. 2, 2026, https://doi.org/ 10.1111/psj.70045.
D. Korencic, B. Chulvi, X. B. Casals, A. Toselli, M. Taule, and P. Rosso, “What distinguishes conspiracy from critical narratives? A computa-tional analysis of oppositional discourse,” Expert Systems, vol. 41, no. 11, 2024, https://doi.org/10.1111/ exsy.13671.
S. Robertson and H. Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond,” Foundations and Trends in Information Retrieval, vol. 4, nos. 1–2, pp. 1–174, 2009, https://doi.org/ 10.1561/1500000019.
J. Chen, S. Xiao, P. Zhang, K. Luo, D. Lian, and Z. Liu, “M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation,” arXiv, 2024, https://doi.org/10.48550/arXiv.2402.03216.
V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W. Yih, “Dense passage retrieval for open-domain question answering,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 2020, https://doi.org/10.18653/v1/2020.emnlp-main.550.
O. Khattab and M. Zaharia, “ColBERT: Efficient and effective passage search via contextualized late interaction over BERT,” in Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, 2020, https://doi.org/10.1145/3397271.3401075.
G. Lin, “Using cross-encoders to measure the similarity of short texts in political science,” American Journal of Political Science, vol. 69,
no. 4, pp. 1600–1616, 2025, https://doi.org/10.1111/ ajps.12956.
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 2019, https://doi.org/10.18653/v1/ D19-1410.
European External Action Service, “EUvsDisinfo: EEAS East StratCom Task Force disinformation database,” 2024. https://euvsdisinfo.eu/ (accessed Jun. 22, 2026).
J. A. Leite, O. Razuvayevskaya, K. Bontcheva, and C. Scarton, “EUvsDisinfo: A Dataset for Multilin-gual Detection of Pro-Kremlin Disinformation in News Articles,” in Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, 2024, https://doi.org/ 10.1145/3627673.3679167.
Google DeepMind, “Gemini Flash 3.5 technical documentation,” 2025. https://deepmind.google/ technologies/gemini/flash/ (accessed Jun. 22, 2026).
INSAIT Institute, “MamayLM v1.0 release blog,” Hugging Face, 2025. https://huggingface.co/ spaces/INSAIT-Institute/mamaylm-v1-blog (accessed Jun. 22, 2026).
Qwen Team, “Qwen3 technical report,” arXiv, 2025, https://doi.org/10.48550/arXiv.2505.09388.
O. V. Melnychuk, “Methodology of a multi-agent system for detection of manipulations and disinformation” [in Ukrainian], Problems of Infor-matization and Management, 2025, https://doi.org/ 10.18372/2073-4751.84.20898
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Oleh Melnychuk

This work is licensed under a Creative Commons Attribution 4.0 International License.
The scientific journal adheres to the principles of Open Access and provides free, immediate, and permanent access to all published materials without financial, technical, or legal barriers for readers.
All articles are published in Open Access under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Copyright
Authors who publish their works in the journal:
-
retain the copyright to their publications;
-
grant the journal the right of first publication of the article;
-
agree to the distribution of their materials under the CC BY 4.0 license;
-
have the right to reuse, archive, and distribute their works (including in institutional and subject repositories), provided that proper reference is made to the original publication in the journal.



