Direct imitation of humans by robots offers a promising directionfor remote teleoperation and intuitive task instruction, where ahuman can perform a task naturally and the robot autonomouslyinterprets and executes it using its own embodiment. Existing methods often rely on close alignment between human and robot scenes.This prevents robots from inferring the intent of the task or executing demonstrated behaviors when the initial states mismatch.Hence, it poses difficulties for non-expert users, who may needdomain knowledge to adjust the setup.To address this challenge, we propose a neuro-symbolic framework that unifies visual observations, robot proprioceptive states,and symbolic abstractions within a shared latent space. Humandemonstrations are encoded into this representation as predicatestates. A symbolic planner can thus generate high-level plans thataccount for the different robot initial states. A flow matching module then synthesizes continuous joint trajectories consistent withthe symbolic plan.We validate our approach on multi-object manipulation tasks.Preliminary results show that the framework can infer humanintent and generate feasible symbolic plans and robot motions undermismatched initial states. These findings highlight the potential ofneuro-symbolic models for more natural human-robot instruction.and they can enhance the explainability and trustworthiness ofrobot actions.

Human-to-Robot Imitation with Symbolic Planning in a Unified Latent Space

Alessandro Di Nuovo
Ultimo
2026-01-01

Abstract

Direct imitation of humans by robots offers a promising directionfor remote teleoperation and intuitive task instruction, where ahuman can perform a task naturally and the robot autonomouslyinterprets and executes it using its own embodiment. Existing methods often rely on close alignment between human and robot scenes.This prevents robots from inferring the intent of the task or executing demonstrated behaviors when the initial states mismatch.Hence, it poses difficulties for non-expert users, who may needdomain knowledge to adjust the setup.To address this challenge, we propose a neuro-symbolic framework that unifies visual observations, robot proprioceptive states,and symbolic abstractions within a shared latent space. Humandemonstrations are encoded into this representation as predicatestates. A symbolic planner can thus generate high-level plans thataccount for the different robot initial states. A flow matching module then synthesizes continuous joint trajectories consistent withthe symbolic plan.We validate our approach on multi-object manipulation tasks.Preliminary results show that the framework can infer humanintent and generate feasible symbolic plans and robot motions undermismatched initial states. These findings highlight the potential ofneuro-symbolic models for more natural human-robot instruction.and they can enhance the explainability and trustworthiness ofrobot actions.
2026
979-8-4007-2321-6
Human-to-Robot Imitation; Neuro-Symbolic Planning; MultimodalLearning
File in questo prodotto:
File Dimensione Formato  
Di-Nuovo_Human-to-Robot-Imitation-with-Symbolic-Planning-in-a-Unified.pdf

accesso aperto

Tipologia: Versione Editoriale (PDF)
Licenza: Creative commons
Dimensione 549.88 kB
Formato Adobe PDF
549.88 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.11769/724309
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact