Development of Language Models in All EU Languages
Ali Levlog/Pexels
Date of publication:
The University of Ljubljana is participating in the European project LLMs4EU (Large Language Models for the European Union), which aims to develop large language models for all official languages of the European Union, including languages with fewer digital resources, such as Slovenian.
The project brings together 69 partners from 20 European countries under the coordination of the Alliance for Language Technologies (ALT-EDIC). Its aim is to strengthen the European artificial intelligence ecosystem by developing open, reliable language models tailored to European needs.
Particular emphasis is placed on developing solutions for five areas: science, public services, tourism, telecommunications, and energy. The goal is to enable organisations that do not have the capacity to develop their own models—particularly small and medium-sized enterprises and public institutions—to benefit from generative artificial intelligence.
At the University of Ljubljana, the project involves the Data Technologies Laboratory at the Faculty of Computer and Information Science, led by Assoc. Prof. Dr Slavko Žitnik. The team is leading the development of a search engine for a catalogue of language technology tools and models, which will allow users to identify the most suitable solutions based on language, application domain, tasks, licences, and other characteristics.
Researchers from the University of Ljubljana are also contributing to the development of the catalogue, which will support the addition of new models and their documentation, as well as to the establishment of a European network of centres for evaluating language models. Within this network, models will be assessed in terms of performance, robustness, safety, cybersecurity, compliance with European legislation, and alignment with human values.
The project is also developing infrastructure for collecting and sharing language data and preparing guidelines for compliance with the Artificial Intelligence Act (AI Act), the General Data Protection Regulation (GDPR), and intellectual property rules. In doing so, it aims to ensure the long-term and trustworthy use of European language models.
Six Slovenian partners are involved in the project: the University of Ljubljana, the Jožef Stefan Institute, Arctur, Event Registry, Telekom Slovenije, and XLAB. Their participation confirms the strong integration of Slovenia’s research and technology community into the development of European language technologies.
For the Slovenian language, the project represents an important step towards the development of higher-quality language models and better support for artificial intelligence in Slovenian.
More information about the project is available on the project's website.