Language models such as ChatGPT are often perceived as neutral tools, but they can silently reinforce existing bias related to gender, ethnicity and social groups. AI researcher Oskar van der Wal from the University of Amsterdam investigated how this bias emerges inside AI systems and what methods could help mitigate it.
According to Van der Wal, many existing methods for measuring bias in AI are too abstract and disconnected from real-world situations. Traditional tests often focus on explicit stereotypes, while bias in practice is much more subtle and highly context-dependent. As a result, these effects can remain largely invisible to both users and developers.
In his research, Van der Wal used realistic medical scenarios to test language models. AI systems were presented with cases in which only the ethnicity of the patient was changed. The models then produced subtle but consistent differences in diagnoses, risk assessments and recommendations. These variations often remained undetected in standard benchmark tests.
He also examined what happens internally while language models are being trained. AI systems learn associations based on recurring patterns in training data. When terms such as “doctor” are more frequently linked to men and “nurse” to women, models absorb and reinforce those relationships over time. Bias therefore emerges not only from the data itself, but also from how models organise and store information.
Van der Wal argues that these forms of bias can be reduced through targeted interventions. In experiments, he showed that specific adjustments inside a model can decrease biased behaviour while largely preserving overall model quality. At the same time, he stresses that responsible AI development requires interventions at multiple levels, including training data, model design, deployment and practical use.
The research connects to broader discussions around responsible AI in science, healthcare and government. As AI systems become increasingly involved in decision-making processes, questions around transparency, accountability and societal impact continue to grow in importance.