Doctorat.gouv.fr
LORIA - Laboratoire Lorrain de Recherche en Informatique et ses Applications
Vandoeuvre lès Nancy cedex
samedi 31 octobre 2026
ANR Financement d'Agences de financement de la recherche
Most work on VLM/LLMs for robotics focused on generating sequences of actions and plans from high level goals, offline, only targeting autonomous robots isolated from humans. A critical limitation to deploy VLM/LLMs for robots collaborating with humans is their ability to be used online, in a human-in-the-loop scenario, to generate suitable motions and 'safe' robot policies. Here, we use VLM/LLMs to generate a robot's motions online in collaborative scenarios where safety is critical: active exoskeletons and mobile manipulators assisting humans in object manipulation. The human vocally commands the robot interactively, online, to control the generation of its motion at the low level: start, stop, direct, and change its low-level parametrization (e.g., compliant behavior, the velocity, the maximal torque assistance, etc.). Extension of paradigms and comparison with existing and fine-tuning of VLAs is also considered, as this is part of the ongoing research of the team. The first objective is to design the robot's controller with the natural language interaction feature in mind: the human's commands, corrections and Approximate Numerical Expressions must be translated into meaningful quantities, coherent with the physics of the problem. What do 'faster', 'a bit higher', 'little to the right', and 'more assistance' mean? The second objective is to design new multimodal models fusing VLM/LLMs and multimodal pipelines to predict the human's intent and minimize the need for corrections. Natural language instructions may be incomplete or unclear, but cameras and microphones (or other sensors) could provide sufficient contextual information to generate an appropriate motion. For example, 'take that' could be easily translated into 'grasp the bottle', if it is the only item in front of the robot. 'Move a bit to the right' needs clarifications, but also estimation of physical quantities that are context dependent. The third objective is to detect emergency commands, leveraging both LLMs and audio processing models for nonverbal communication, and generating suitable robot's reactive behaviors. Humans are often unable to speak clearly when they interact with a robot: sometimes, fear takes over and they do not speak at all, or they mumble, or scream, when they could just say a clear 'stop'. Detecting emergency commands is critical to be able to deploy the robots into the real world. For example, 'Watch out', 'Attention!' are difficult to translate into precise motions, and require one-shot evaluations because of the urgent nature of the command. The PhD student will carry out research in the aforementioned objectives, and will benefit from our collaboration with E. Zibetti (Paris 8, SHS), expert in Approximate Numerical Expressions for Psychology, and D. Sadigh (Stanford University), leading the research in LLMs for robot actions. Real-world demonstrations with real robots and real humans interacting with the robots are mandatory in this PhD. École doctorale : IAEM - INFORMATIQUE - AUTOMATIQUE - ELECTRONIQUE - ELECTROTECHNIQUE - MATHEMATIQUES Direction : Serena IVALDI Financement : ANR Financement d'Agences de financement de la recherche
Source : Doctorat.gouv.fr · Récupérée le 24 septembre 2026