Multimodal models (WP9)

Work Package WP9 focuses on the research and development of advanced multimodal models, with an emphasis on multilingualism. The research covers the processing of text, speech and visual input. A key component of the work is a complete data pipeline encompassing the acquisition, cleaning, structuring and annotation of large datasets, as well as the creation of metadata for all supported modalities and other WPs.

The main objective of WP9 is to develop robust base models that will serve as a common technological foundation for the project’s other work packages and which will find direct practical application through close collaboration with industrial partners on the one hand and government bodies on the other.

What problem does the WP solve, and why is it important?

Generic global AI models do not provide sufficient support for the Czech language

Existing AI models fail in highly specialised (domain-specific) fields involving technical terminology

Standard models cannot be safely applied to private and sensitive data without the risk of a data breach

Existing systems are unable to seamlessly combine and simultaneously analyse text, speech and visual input

High-quality, cleaned and structured training data with appropriate metadata is in critically short supply

Both the public sector and industry are encountering barriers when deploying advanced multimodal AI solutions in real-world operations

Examples of use / areas of application and benefits

Spoken language and audio processing

  • Advanced audio and dialogue processing
  • Operator status analysis and detection of voice impersonation

Public administration

  • Assisted drafting of clear legal texts
  • Data engineering for automated data management

Computer vision

  • High reliability in non-standard conditions
  • Detailed classification according to language requirements

Cognitive research

  • Monitoring human perception using eye-tracking
  • Greater credibility and naturalness of AI

Business processes

  • Automating routine administrative tasks using AI
  • Operational savings and privacy protection

Key people

prof. RNDr. Jan Hajič, Dr.

Institute of Formal and Applied Linguistics (ÚFAL) | MFF UK

prof. RNDr. Jan Hajič, Dr.

| MFF UK

prof. Dr. Ing. Jan Černocký

| BUT

prof . Ing. Jiří Matas, Ph.D.

| FEL CTU

doc. Ing. Jindřich Matoušek, Ph.D.

| UWB

Ing. Jan Šedivý, CSc.

| CIIRC CTU

Participating institutions

Used technologies and procedures

Achieved and planned results

2027

A robust visual-linguistic model

A compact voice biometrics system

A dialogue system with emotion recognition

A hierarchical visual-language model for fine-grained classes

Pipeline for the automation of business documents

Assistant for generating semantic metadata for public administration data sources – METAR

2028

A set of models for speaker-oriented speech processing (SPEAKER)

Client-Oriented Writing – PONK v. 1.0+

A service for searching for data sources in data catalogues – FINDIT

High-quality and efficient TTS for real-world applications

2029

A production service for entity extraction and document classification

2031

An adaptive, self-learning, dialogue-based system with emotion recognition

An adaptive, self-learning, dialogue-based system with emotion recognition

A multimodal model utilising source identity

Client-Oriented Writing – PONK v. 2.0+

Module for advanced phonetic normalisation of text

Production service for advocacy and compliance

Expansion of the National Data Catalogue’s open-source software to include the integrated METAR and FINDIT tools