Consent for the use of cookies and other tools

Tools and cookies used on the website collect information about visitors in anonymized form. Your consents enable us to ensure the functioning of all website features, customize certain content specifically for you, and continuously improve the website by analyzing visits.

Types of Cookies

I consent to the website use of tools, including cookies, which ensure full functionality and an appropriate level of security. I understand that without this, the website cannot offer proper functioning, such as website navigation, customization of appearance, and access to specific parts of the website.

I consent to the website use of tools, including cookies, which collect anonymized data about website visitors. I understand that without this, website administrators cannot analyze site traffic and usage patterns to improve the user experience on the website.

Changes were successfully saved

High-quality data – the foundation of the model

One of the greatest challenges in developing a large language model for Slovene was the availability of high-quality training data. Unlike major world languages, Slovene has a relatively small number of speakers; therefore, a citizen science approach and collaboration with the wider community were essential.

As part of a data collection campaign on the povejmo.si platform, Slovene speakers contributed texts for training the model. Additional training data were provided by organisations including the National and University Library of Slovenia, media house Dnevnik, and the Slovenian Press Agency.

The donated data enabled the development of a model that better understands the Slovene linguistic landscape and the specific ways Slovene is used in digital environments.

GaMS – a milestone for Slovene artificial intelligence

The technical development of GaMS took place in three key stages: pre-training, fine-tuning, and evaluation. The researchers built the model on Google’s open-source Gemma 2 and Gemma 3 models and trained it on extensive collections of original, machine-generated, and translated Slovene texts.

The model was then fine-tuned to better follow instructions, ensure safer use, and account for the linguistic and cultural specificities of Slovene. The development process also included extensive evaluation, comparing GaMS with other open-source and proprietary language models.

A foundation for Slovenia’s technological sovereignty

GaMS represents an important step towards reducing Slovenia’s dependence on foreign corporations in the field of language technologies. With large language models, performance is only one part of the equation; equally important is maintaining control over how the technology is developed, deployed, and accessed.

The model is available to users as a chatbot on the povejmo.si platform and is also released as an open-source model on Hugging Face. This allows companies, public administration institutions, and researchers to adapt it to their own needs while retaining control over their data.

By making advanced language technology openly available, GaMS supports the safer use of artificial intelligence, strengthens the presence of the Slovene language in the digital sphere, and enhances the competitiveness of the Slovenian economy.

From research to practical solutions for industry

An important part of the project was transferring the technology into practical applications. The companies Better, Semantika, Špica, and XLAB adapted GaMS to develop advanced industrial and business solutions.

Semantika adapted GaMS for creating presentation materials for museums and for advanced interactive museum applications. All these solutions are based on GalisOnline, a documentation system for museums and archives containing rich, structured information about cultural heritage artefacts.

Špica developed a computationally efficient version of GaMS for integration into advanced industrial speech recognition applications. The model accurately recognises dialects, foreign languages, and speech in noisy environments. It enables reliable communication between workers and industrial systems even under demanding working conditions with high levels of background noise and multilingual personnel.

Better fine-tuned GaMS using medical texts and clinical instructions to create a healthcare application that automates routine medical workflows. The application supports electronic medication signing, automatic speech transcription for discharge summaries, and the digitalisation of paper forms through text field recognition and automated data entry.

XLAB developed an advanced application that improves data efficiency while significantly increasing the accuracy and robustness of source code and documentation generation. Computationally efficient large language models and dedicated processing pipelines are used to generate software-defined descriptions of IT infrastructure.

The transfer of AI technologies from research into industry strengthens the productivity and competitiveness of the Slovenian economy. At the same time, GaMS ensures that Slovene is represented in AI systems on an equal footing with much larger language communities.

The development of GaMS demonstrates that the Slovenian research community can create world-class artificial intelligence solutions and successfully transfer them to industry. With GaMS, Slovenian has taken an important step towards equal representation in the world of modern AI systems.

GaMS was developed as part of the research and innovation programme Adaptive Natural Language Processing with Large Language Models (PoVeJMo), which ran from 2023 to 2026. The project was coordinated by Dr Simon Krek, Head of the Centre for Language Resources and Technologies at the University of Ljubljana, which has been developing language resources and technologies to support Slovenian in the digital age for more than a decade.

Chat with GaMS: https://povejmo.si/

GaMS in Open Access: https://gams.povejmo.si/odprtidostop/

Logos

  • Financira Evropska unija (NextGenerationEU) (en)
  • ARIS (en)