Researchers at the University of Amsterdam’s Faculty of Science have created NeoBabel, an advanced AI image generation system designed to work directly with prompts in multiple languages. This new tool addresses the digital inequality that arises because many existing AI image models perform best only with English input, disadvantaging people who speak other languages.
AI image generation technology has developed rapidly in recent years, but most models are optimised for English text. When non‑English prompts are used, they are frequently translated into English before image generation, which can lead to loss of linguistic and cultural nuance. This places non‑English speakers at a disadvantage when creating AI‑generated visuals.
To tackle this issue, a research team from the UvA’s Informatics Institute partnered with AI company Cohere Labs to integrate image generation with powerful multilingual text models. The result is NeoBabel, an AI image generator that can understand prompts written in six languages (English, French, Dutch, Chinese, Hindi and Persian) and produce images directly from each language without prior translation.
Unlike many commercial models developed by large companies that are not fully transparent, the NeoBabel research group has made its code and data fully open source. This open approach allows other researchers and developers to build on their work and contribute to a more inclusive AI landscape.
The team also improved training data quality by using multilingual models to create richer, descriptive labels for training images. This multilingual labelling process not only supports direct image generation across all six languages but also enables the model to learn stronger connections between words and visual elements.
One unique application of NeoBabel is a collaborative creative canvas where users who speak different languages can work together on the same visual project. Each participant can generate and modify the image using text in their own language, and the system updates the image based on all contributions.
Looking ahead, the research team aims to expand NeoBabel’s capabilities to video generation and culturally specific images, though they note that this will require additional compute resources and data. They are seeking collaboration partners to help advance these goals and explore new creative possibilities.