This work presents a modular framework that integrates vision-language models and multimodal Retrieval-Augmented Generation (RAG) to enhance interaction in Serious Games for cultural heritage. The system leverages CLIP to extract visual embeddings from in-game scenes and FAISS to perform efficient similarity search against a pre-indexed image database. Retrieved metadata is used to generate context-aware responses via a locally hosted large language model (LLM), enabling a conversational agent to provide semantically aligned and culturally inclusive dialogue. The architecture consists of three main components: a preprocessing pipeline for image embedding and metadata serialization, a Unity-based client for user interaction, and a FastAPI-based backend for coordinating retrieval and generation. Tested in an archaeological virtual museum, the framework demonstrates improved user engagement, educational potential, and support for adaptive narrative experiences. By combining real-time visual understanding with generative language capabilities, this approach advances the design of interactive and accessible cultural heritage experiences driven by AI.

A Modular AI-Powered Framework for Semantic Retrieval and Dialogue in Serious Games for Cultural Heritage

Simone Pio Barbagallo
;
Roberto Rizza
;
Dario Allegra;Anna Maria Gueli;Filippo Stanco
2025-01-01

Abstract

This work presents a modular framework that integrates vision-language models and multimodal Retrieval-Augmented Generation (RAG) to enhance interaction in Serious Games for cultural heritage. The system leverages CLIP to extract visual embeddings from in-game scenes and FAISS to perform efficient similarity search against a pre-indexed image database. Retrieved metadata is used to generate context-aware responses via a locally hosted large language model (LLM), enabling a conversational agent to provide semantically aligned and culturally inclusive dialogue. The architecture consists of three main components: a preprocessing pipeline for image embedding and metadata serialization, a Unity-based client for user interaction, and a FastAPI-based backend for coordinating retrieval and generation. Tested in an archaeological virtual museum, the framework demonstrates improved user engagement, educational potential, and support for adaptive narrative experiences. By combining real-time visual understanding with generative language capabilities, this approach advances the design of interactive and accessible cultural heritage experiences driven by AI.
2025
9783032113801
File in questo prodotto:
File Dimensione Formato  
Barbagallo et al_A Modular AI-Powered Framework.pdf

solo gestori archivio

Tipologia: Versione Editoriale (PDF)
Licenza: NON PUBBLICO - Accesso privato/ristretto
Dimensione 3.8 MB
Formato Adobe PDF
3.8 MB Adobe PDF   Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.11769/701696
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? 0
social impact