Direct imitation of humans by robots offers a promising directionfor remote teleoperation and intuitive task instruction, where ahuman can perform a task naturally and the robot autonomouslyinterprets and executes it using its own embodiment. Existing methods often rely on close alignment between human and robot scenes.This prevents robots from inferring the intent of the task or executing demonstrated behaviors when the initial states mismatch.Hence, it poses difficulties for non-expert users, who may needdomain knowledge to adjust the setup.To address this challenge, we propose a neuro-symbolic framework that unifies visual observations, robot proprioceptive states,and symbolic abstractions within a shared latent space. Humandemonstrations are encoded into this representation as predicatestates. A symbolic planner can thus generate high-level plans thataccount for the different robot initial states. A flow matching module then synthesizes continuous joint trajectories consistent withthe symbolic plan.We validate our approach on multi-object manipulation tasks.Preliminary results show that the framework can infer humanintent and generate feasible symbolic plans and robot motions undermismatched initial states. These findings highlight the potential ofneuro-symbolic models for more natural human-robot instruction.and they can enhance the explainability and trustworthiness ofrobot actions.
Human-to-Robot Imitation with Symbolic Planning in a Unified Latent Space
Alessandro Di NuovoUltimo
2026-01-01
Abstract
Direct imitation of humans by robots offers a promising directionfor remote teleoperation and intuitive task instruction, where ahuman can perform a task naturally and the robot autonomouslyinterprets and executes it using its own embodiment. Existing methods often rely on close alignment between human and robot scenes.This prevents robots from inferring the intent of the task or executing demonstrated behaviors when the initial states mismatch.Hence, it poses difficulties for non-expert users, who may needdomain knowledge to adjust the setup.To address this challenge, we propose a neuro-symbolic framework that unifies visual observations, robot proprioceptive states,and symbolic abstractions within a shared latent space. Humandemonstrations are encoded into this representation as predicatestates. A symbolic planner can thus generate high-level plans thataccount for the different robot initial states. A flow matching module then synthesizes continuous joint trajectories consistent withthe symbolic plan.We validate our approach on multi-object manipulation tasks.Preliminary results show that the framework can infer humanintent and generate feasible symbolic plans and robot motions undermismatched initial states. These findings highlight the potential ofneuro-symbolic models for more natural human-robot instruction.and they can enhance the explainability and trustworthiness ofrobot actions.| File | Dimensione | Formato | |
|---|---|---|---|
|
Di-Nuovo_Human-to-Robot-Imitation-with-Symbolic-Planning-in-a-Unified.pdf
accesso aperto
Tipologia:
Versione Editoriale (PDF)
Licenza:
Creative commons
Dimensione
549.88 kB
Formato
Adobe PDF
|
549.88 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


