Amid the global race for artificial intelligence, Ukraine is taking a strategic step towards digital sovereignty. Head of Ukraine's Ministry of Digital Transformation, Oksana Ferchuk, confirmed the creation of a national large language model (LLM) named "Syayvo". The project, which will become a key element of the country's digital infrastructure, is based on Google's cutting-edge Gemma 3 architecture and is being adapted to the unique linguistic and cultural features of the Ukrainian language.

Technological Foundation and Infrastructure

The development of "Syayvo" is a large-scale public-private partnership. The telecommunications giant "Kyivstar" is taking on the technical infrastructure, financing, and organizational support. The choice of the Gemma 3 base model from Google is due to its high efficiency and openness, allowing engineers to focus on fine-tuning for specific tasks of the Ukrainian market. The key task at the current stage is the adaptation of the tokenizer—the component that breaks text into semantic units—to ensure the fastest and most accurate processing of Ukrainian vocabulary and grammar.

Creation of a National Data Corpus

The main challenge for developers remains the formation of a high-quality national data corpus. To train the model, the team is collecting and verifying information from more than 90 state institutions, scientific institutions, universities, and media resources. Special attention is paid to historical accuracy: the State Archive of Ukraine (Ukrhosarkhiv) has already provided more than 10 terabytes of materials, including historical documents, state acts, and scientific works. A significant part of this data requires complex digitization from paper media, which is a laborious but necessary process for preserving historical memory in digital format.

Expert Control and Security

To guarantee the quality and safety of AI operations, a special expert committee of more than 70 specialists has been formed. They evaluate the model in four critical areas: language accuracy and depth of context understanding, security (resilience to hallucinations, disinformation, and bias), absence of discriminatory responses, and comparative analysis with global counterparts. In June 2025, a small-scale prototype of "Syayvo" successfully passed the pre-training and supervised fine-tuning stages, confirming the viability of the chosen architecture.

Open Source and Integration Plans

After testing is completed and the project is handed over to the state, the base version of "Syayvo" and the Ukrainian datasets created during development will be published in the open source domain. This will allow developers, businesses, and scientists to obtain a technical pipeline for training the model for their own products. In the public sector, "Syayvo" will become the basis for future services, such as "Diia AI", an educational assistant in the "Mriya" app, and tools for legislative analysis. The test version of the model is scheduled to be completed by the end of 2026, and it will be opened to a wide range of users in January 2027.