BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages

Lucas, J.1, Murtagh-White, M.2, Uchendu, A.3, Al-Lawati, A.1, Yamashita, M.1, Macko, D., Srba, I., Moro, R., Lee, D.1

1 The Pennsylvania State University, 2 Trinity College Dublin, 3 MIT Lincoln Lab

Multilingual falsehoods threaten information integrity worldwide, yet detection benchmarks remain confined to English or a few high-resource languages, leaving low-resource linguistic communities without robust defense tools. We introduce BLUFF, a comprehensive benchmark for detecting false and synthetic content, spanning 79 languages with over 202K samples, combining human-written fact-checked content (122K+ samples across 57 languages) and LLM-generated content (79K+ samples across 71 languages). BLUFF uniquely covers both high-resource “big-head” (20) and low-resource “long-tail” (59) languages, addressing critical gaps in multilingual research on detecting false and synthetic content. Our dataset features four content types (human-written, LLM-generated, LLM-translated, and hybrid human-LLM text), bidirectional translation (English?eftrightarrowX), 39 textual modification techniques (36 manipulation tactics for fake news, 3 AI-editing strategies for real news), and varying edit intensities generated using 19 diverse LLMs. We present AXL-CoI (Adversarial Cross-Lingual Agentic Chain-of-Interactions), a novel multi-agentic framework for controlled fake/real news generation, paired with mPURIFY, a quality filtering pipeline ensuring dataset integrity. Experiments reveal state-of-the-art detectors suffer up to 25.3% F1 degradation on low-resource versus high-resource languages. BLUFF provides the research community with a multilingual benchmark, extensive linguistic-oriented benchmark evaluation, comprehensive documentation, and open-source tools to advance equitable falsehood detection. Dataset and code are available at: https://jsl5710.github.io/BLUFF/

Cite: Lucas, J., Murtagh-White, M., Uchendu, A., Al-Lawati, A., Yamashita, M., Macko, D., Srba, I., Moro, R., Lee, D. BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages. In Proceedings of the 2026 ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery. 2026

Authors

Dominik Macko
Researcher
More
Ivan Srba
Researcher
More
Róbert Móro
Researcher
More