Research Group

Jaroslav Kopčan
Research areas: Natural Language Generation, Multilingual Language Processing, Deep Learning, Explainability, Knowledge Distillation
Position: Research Engineer
Jaroslav Kopčan is a Research Engineer whose PhD and earlier work focused on Deep Neural Networks and Explainable AI. Since then, his research has centered on applying NLP to social science, developing explainable methods for combating disinformation, language modeling for low-resource languages, knowledge distillation, and machine translation evaluation. He is particularly interested in agentic harnesses, both in the problems they can meaningfully solve and in what it takes to design robust and reliable ones.
Selected activities
- NLP for social science research. Applying language models and text analysis methods to study social phenomena at scale, turning large textual corpora into structured evidence that social scientists can work with.
- Explainable methods for combating disinformation. Building detection and analysis tools that not only flag disinformation but also surface why a given piece of content is suspicious, so that human moderators and researchers can verify and act on the output.
- Language modeling for low-resource languages. Training, evaluating language models and methods for languages with limited data and tooling, addressing the gap left by systems optimized primarily for English and other high-resource languages.
- Knowledge distillation. Compressing large models into smaller, cheaper ones that retain most of the original capability for given tasks, with attention to where and why distillation succeeds or breaks down.
Selected Projects
Selected Publications
Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection
Vykopal, I., Karamolegkou, A.1, Kopcan, J., Peng, Q.1, Javurek, T., Gregor, M., Simko, M. 1 University of Copenhagen, Denmark, Copenhagen Multilingual Large Language Models (LLMs) offer powerful capabilities for cross-lingual…
When the Dictionary Strikes Back: A Case Study on Slovak Migration Location Term Extraction and NER via Rule-Based vs. LLM Methods
This study explores the task of automatically extracting migration-related locations (source and destination) from media articles, focusing on the challenges posed by Slovak, a low-resource and morphologically complex language. We…




