Multimodal models (WP9)
Work Package WP9 focuses on the research and development of advanced multimodal models, with an emphasis on multilingualism. The research covers the processing of text, speech and visual input. A key component of the work is a complete data pipeline encompassing the acquisition, cleaning, structuring and annotation of large datasets, as well as the creation of metadata for all supported modalities and other WPs.
The main objective of WP9 is to develop robust base models that will serve as a common technological foundation for the project’s other work packages and which will find direct practical application through close collaboration with industrial partners on the one hand and government bodies on the other.
What problem does the WP solve, and why is it important?
01
Generic global AI models do not provide sufficient support for the Czech language
02
Existing AI models fail in highly specialised (domain-specific) fields involving technical terminology
03
Standard models cannot be safely applied to private and sensitive data without the risk of a data breach
04
Existing systems are unable to seamlessly combine and simultaneously analyse text, speech and visual input
05
High-quality, cleaned and structured training data with appropriate metadata is in critically short supply
06
Both the public sector and industry are encountering barriers when deploying advanced multimodal AI solutions in real-world operations
Examples of use / areas of application and benefits
Spoken language and audio processing
- Advanced audio and dialogue processing
- Operator status analysis and detection of voice impersonation
Public administration
- Assisted drafting of clear legal texts
- Data engineering for automated data management
Computer vision
- High reliability in non-standard conditions
- Detailed classification according to language requirements
Cognitive research
- Monitoring human perception using eye-tracking
- Greater credibility and naturalness of AI
Business processes
- Automating routine administrative tasks using AI
- Operational savings and privacy protection
Key people

Participating institutions










Used technologies and procedures
- Speech Processing and Synthesis
- A set of models for speaker-oriented speech processing
- A robust visual-linguistic model (Type R, software, #83)
- A compact voice biometrics system
- A service for advanced searching of datasets in a DCAT-AP-compliant data catalogue
- A hierarchical visual-language model for fine-grained classes (Type R, software, #84)
Achieved and planned results
2027
A robust visual-linguistic model
R – software
A compact voice biometrics system
R – software
A dialogue system with emotion recognition
R – software
A hierarchical visual-language model for fine-grained classes
R – software
Pipeline for the automation of business documents
R – software
Assistant for generating semantic metadata for public administration data sources – METAR
R – software
2028
A set of models for speaker-oriented speech processing (SPEAKER)
R – software
Client-Oriented Writing – PONK v. 1.0+
Gfunk – working sample
A service for searching for data sources in data catalogues – FINDIT
R – software
High-quality and efficient TTS for real-world applications
R – software
2029
A production service for entity extraction and document classification
R – software
2031
An adaptive, self-learning, dialogue-based system with emotion recognition
R – software
An adaptive, self-learning, dialogue-based system with emotion recognition
R – software
A multimodal model utilising source identity
R – software
Client-Oriented Writing – PONK v. 2.0+
R – software
Module for advanced phonetic normalisation of text
R – software
Production service for advocacy and compliance
R – software
Expansion of the National Data Catalogue’s open-source software to include the integrated METAR and FINDIT tools
R – software