Same Question, Different Answer? Auditing AI Chatbots with Simulated User Personas

Kukla, M., Koop, R., Malínský, T., Milotinský, T.

Millions of people now turn to AI chatbots for advice on their careers, their health, and their mental wellbeing. The companies behind ChatGPT, Claude, and Gemini publish specifications describing how their models should behave: fairly, safely, and without steering users’ opinions. But do the models actually behave that way? And does the answer change depending on who is asking? 

These are the questions behind a research project carried out by the students from the DELTA Secondary School of Informatics and Economics in Pardubice, Czech Republic, during their Erasmus internship at KInIT. DELTA students visited us during September 2026 for the third time and, similarly as their classmates in previous years, they again showcased their great expertise in computer science. In the following blog post, they describe what they were working on as well as their personal experience of being with us at KInIT.

Why the interface matters

Most AI audits test models through the API, the programmatic access developers use. It is convenient and easy to reproduce, but it is not what ordinary people see. The chat window in your browser adds hidden instructions, memory of past conversations, and personalization on top of the underlying model. Previous research shows that the interface and the API can behave quite differently. We therefore test the chatbots the way real users meet them: through their user interfaces. We run the same tests in English, Czech, and Slovak, because smaller languages are rarely examined even though their speakers use the same products every day.

A methodology built for reproducibility

An audit is only useful if someone else can repeat it and get comparable results. We designed a step-by-step methodology that covers the whole process, from framing the research question to archiving the results. The core part of this methodology consists of definition of user personas – users to be simulated during the audit. These personas define the demographic characteristics (such as age, gender, nationality/language) as well as their interests, job positions, etc. The purpose of these personas is to indicate different types of users the chatbot is communicating with.

Besides that, the auditor can propose a set of testing questions that serve as a probe into the internal workings of chatbots. Responses to such questions are assessed automatically as well as by human annotators, and we measure how consistently they agree or disagree for individual personas.

From methodology to tool

Based on this methodology, we built a web application that makes such audits possible at a much larger scale. It controls real chatbot interfaces in the browser, runs the same prompts across multiple providers, accounts, and personas, and records every conversation with full details about when and how it was collected. This allows several researchers to work on the same audit in parallel and makes it far easier for others to run their own studies.

Putting it into practice: career advice

We used the tool we developed to examine career advice. Each team member tested a different dimension: whether recommendations changed with the user’s gender, age, ethnicity, or social background, while everything else in the conversation stayed the same.

What we found

Across 271 conversations and 6 personas, we found no significant differences. The chatbots gave essentially the same career advice regardless of the user’s gender, age, ethnicity, or social background. This is likely because bias in career advice is one of the most widely studied and publicly discussed problems in AI, and the companies behind these models have likely invested heavily in addressing it. A result like this is still valuable: it suggests that, at least in this area, the models live up to their promises. It also suggests that future audits may uncover more by looking at less-explored areas, where vendors have had less reason to focus their attention.

What our internship at KInIT gave us

The internship at KInIT was our first real experience of research work. We learned how to take an initial idea all the way to a finished project: designing a methodology, building a tool to carry it out, and collecting and analyzing the results. Above all, we learned to work independently. We made our own decisions, solved problems as they came up, and took responsibility for the direction of our work.

Outside of work, we really enjoyed Bratislava, especially the cultural events and the city’s architecture. We would recommend the internship to any students who are interested in exploratory assignments where they can shape the direction of the work themselves.

srba

“It was fantastic to see how quickly the DELTA students embraced the challenging topic of auditing AI chatbots and how quickly they turned their ideas into a working solution. I was especially impressed to see them take it all the way from understanding the problem to demonstrating their approach in a real audit study.”

Ivan Srba

Researcher