Computercode

DVPS: Foundation Models for Future AI

Status: Ongoing Health Remote sensing Water

DVPS, a Horizon Europe research project, develops an open-source toolkit to streamline the design, pre-training, fine-tuning, and modality expansion of Multimodal Foundational Models (MMFMs). This toolkit supports the reuse and composition of pre-trained models, reducing development time and cost while enhancing adaptability across modalities.

Advancing Trustworthy Multimodal Foundation Models in Europe 

The DVPS project (“Diversis Viis, Plurima Solvens”, Latin for "Through diverse paths, I solve many issues") aims to advance the science and engineering of multimodal foundation models, a pioneering domain within AI research. These models can generalise across a wide range of data modalities and application domains. 

This project addresses the growing need for scalable, trustworthy, and adaptable AI systems that reduce reliance on domain‑specific models while strengthening Europe’s technological autonomy. To this end, DVPS develops AutoDVPS, an open‑source platform that automates the design, pre‑training, fine‑tuning, and expansion of multimodal models. Complementing this, the project delivers DVPSBench, a comprehensive benchmarking suite designed to assess robustness, safety, factual accuracy, and multimodal performance. 

Multidisciplinary Consortium Spanning Various Industries 

The consortium brings together 19 European partners (including VITO, Translated, the University of Oxford, EPFL, ETH Zürich, FBK, and AUMC) spanning academia, industry, and healthcare. Their work encompasses three planned application domains: geo‑intelligence, cardiology, and language communication, alongside two additional ‘surprise’ domains introduced to evaluate the generalisability of the developed methods. 

In the health and cardiology domain, DVPS investigates federated, privacy‑preserving multimodal models and tools for generating structured medical data and anonymising clinical text. In geo‑intelligence, it focuses on sensor‑agnostic models that integrate 2D and 3D remote sensing, Geographic Information Systems (GIS) sources, and multimodal forecasting approaches. In language communication, the project advances personalised, multilingual multimodal models capable of reasoning, cross‑modal interaction, and user adaptation. 

Open-Source Assets to Advance AI Science 

In line with its commitment to open science, DVPS releases all tools, models, benchmarks, and documentation as open‑source assets. Collectively, these efforts establish a rigorous scientific and operational foundation for building, evaluating, and deploying the next generation of European multimodal AI systems.

VITO’s Role (in DVPS): Driving AI Innovation in Health and Geo-Intelligence 

VITO is involved in the DVPS project across two applications domains: cardiology and geo-intelligence

Cardiology 

VITO plays a multidimensional role in the DVPS project: our research centre contributes expertise that spans trustworthy generative AI, patient‑facing clinical communication, and the transformation of medical information into standardised data. 

A key ambition of VITO’s work is the development of hallucination‑free generative medical AI, ensuring that every model output, whether a clinical summary, a patient letter, or an interactive explanation, is grounded in verifiable medical evidence. By combining domain‑specific consistency checks, high‑quality multimodal inputs, and rigorous evaluation methods, VITO aims to set a new benchmark for safe, reliable, and clinically responsible generative AI. 

In parallel, VITO leads the project’s multilingual patient letter translation and conversational interaction use case, which serves as an applied setting in which these trustworthy generative capabilities are deployed. This work improves the clarity, accessibility, and accuracy of medical communication across linguistic boundaries, while also acting as a practical testing ground for ensuring that generative systems remain factual, context‑aware, and aligned with clinical workflows. Through this use case, VITO demonstrates how advanced multimodal models can meaningfully support healthcare professionals in delivering information to patients in a safe and culturally sensitive way. 

Complementing these efforts, VITO is heavily engaged in building tools to convert medical data into well‑structured data that comply with international standards such as FHIR, SNOMED CT, and OMOP. This capability is essential for enabling trustworthy generative AI, as structured and high‑integrity data form the foundation for safe downstream reasoning, auditability, and regulatory acceptance. These tooling efforts strengthen the broader DVPS ecosystem by contributing interoperable, transparent, and reusable components that support both clinical research and real‑world deployments. 

Together, these three pillars (hallucination‑free generative medical AI, multilingual clinical communication, and robust clinical data structuring) underscore VITO’s commitment to improving the safety, usability, and regulatory alignment of next‑generation medical AI systems. Learn more about VITO’s work on regulatory science in the health domain and the numerous research projects in which we contribute to it. 

Geo-Intelligence 

VITO also operates within the Geo-Intelligence application domain: in addition to its coordinating role, VITO contributes a flood disaster management use case to assess the applicability of DVPS technologies in the context of climate adaptation and resilience. This use case targets challenges of societal relevance, particularly the need for timely and reliable information during rapidly evolving flood events. By focusing on both tactical and strategic decision-making levels, the initiative addresses gaps in how heterogeneous and incomplete data sources are currently utilised in crisis contexts. 

A key innovation lies in the development of tools capable of handling partial and continuously changing remote sensing inputs, including satellite, drone, and crowdsourced data. VITO advances downstream processing techniques by exploiting the multimodality of available data. For example, by fusing satellite optical and synthetic aperture radar (SAR) imagery, or drone and vector data to improve the reliability of flood extent detection and damage assessment under varying conditions. This multi-source data fusion enhances robustness compared to single-modality approaches and supports more accurate situational awareness during critical response phases. 

In parallel, VITO explores the integration of large language models (LLMs) with Geo-Intelligence foundation models, enabling advanced reasoning. This approach facilitates more context-aware analysis and interpretation of geospatial data. Additionally, the work investigates embedding these capabilities into agent-based workflows that incorporate GIS and vector data, aiming to create more interactive and comprehensive decision-support systems for first responders and policymakers. These innovations contribute to more adaptive, data-driven flood and crisis management strategies.

Project partners

  • Translated (coordinator)
  • The Alan Turing Institute
  • University of Oxford
  • Eidgenössische Technische Hochschule Zürich (ETH Zürich)
  • EPFL
  • KIT
  • Fondazione Bruno Kessler (FBK)
  • Cyfronet HPC
  • Pi School
  • Heidelberg University Hospital
  • Imperial College London
  • Deepset
  • Meteorological and Environmental Earth Observation (MEEO)
  • Sistema GmbH
  • Universitat de Barcelona
  • Stichting Amsterdam
  • Data Valley Consulting
  • Lynkeus
  • VITO
  • Vall d'Hebron Hospital Institute
  • July 2025 - June 2029

Contact person
Contact person