AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR

Article summary

AgentHands is an LLM-powered XR prototype that augments conversational agents with synchronized, expressive hand gestures to provide spatially grounded guidance, bridging the mental mapping gap and enhancing user engagement in physical tasks.

Captured article text

As AI assistants evolve from simple text interfaces to multimodal companions, we are seeing a shift toward more proactive, situated assistance. Recent innovations like Project Astra and Gemini 3.1 Flash Live already allow users to discuss their physical surroundings in real time, often utilizing visual bounding box overlays to identify objects in a camera feed. While these overlays are highly effective for 2D screens, the transition to immersive platforms like Android XR presents a unique challenge: how do we move beyond flat UI to create a truly embodied, spatially aware dialogue?

To bridge this gap, we introduce AgentHands, published at CHI 2026, a research prototype that brings the power of co-speech gestures to the 3D world. In human communication, our hands do more than just point; they describe shapes, mimic actions, and emphasize points, all synchronized with our voice. By leveraging the spatial understanding capabilities of Extended Reality (XR), AgentHands replicates this natural synergy. Following up our prior research in Human I/O and Sensible Agent, AgentHands further equips AI agents with expressive, synchronized hand gestures that transform abstract verbal instructions into intuitive, physical demonstrations, making conversations about your surroundings more natural and engaging.

A taxonomy for embodied hand agents in XR

To start, we conducted a formative study with XR and human–computer interaction (HCI) experts at Google to determine what makes a virtual hand “legible” in a 3D environment. We distilled these insights into a multi-dimensional taxonomy that defines how an agent should use its hands to ground a conversation within a user’s physical space.

  • Handedness & gesture: Choosing between one or two hands and selecting from a library of forms, such as a “palm” for caution or a “cylindrical grip” to mimic holding a tool.
  • Spatiality: Leveraging the depth of XR to determine where hands should live. They can be mid-air for general conversation, object-anchored for identifying specific parts, or user-relative for social cues.
  • Temporal dynamics & visual effects (VFX): Gestures in XR are not just static poses; they include animated motions like “pouring” or “tracing”. XR’s visual layer can add effects, such as a red glow to signify a heat warning.

The taxonomy diagram shows six dimensions: Handedness, Gesture, Spatiality, Temporal Dynamics, Interactivity, and Visual Effects.

The AgentHands workflow

The core innovation of AgentHands is its ability to map the high-level reasoning of LLMs into precise, real-time physical motions that match the agent’s “voice” and the user’s XR environment.

Environment awareness

The system begins with a lightweight object registration module. Using eye gaze and scene reconstruction, users can quickly “tag” items — like an orchid or a laptop — creating a spatial registry with 3D bounding boxes that the agent can reference.

Hand gesture event library

The system constructs a library of hand gesture behaviors across three semantic categories: deictic for referencing, iconic for depicting actions or forms, and expression for conveying social cues and emotion.

Gesture-embedded reasoning

When a user asks a question, the backend LLM generates a response that includes inline GestureEvents. Each event is attached to specific trigger words and encodes the primitives for a hand behavior following the taxonomy dimensions.

Synchronized XR execution

A local parser on the XR headset coordinates text-to-speech playback with the animation engine. Using word-level timestamps, the agent’s hands perform co-speech gestures in sync with the spoken words, providing clear, expressive spatial references.

By integrating these modules, AgentHands creates a bridge between linguistic intent and physical action. The system transforms a standard LLM output into a multimodal performance where the agent’s generated responses are manifested through speech and spatially accurate movement, allowing complex instructions to be demonstrated where they occur in the user’s environment.

Application scenarios

The article demonstrates how embodied gestures, paired with XR spatial awareness, can enhance understanding of physical surroundings.

  • Interactive tutoring: In an orchid-care scenario, the agent moves its hands to the base of the plant and outlines the air roots while explaining their function.
  • Technical walkthroughs: For 3D printer operations, the agent demonstrates the exact “turn and click” sequence needed to navigate control knobs and select files.
  • Lifestyle companionship: The agent can act as a wellness coach, using interactive warning gestures and visual effects to caution a user about an unhealthy behavior.

User study

The authors conducted a within-subjects study (N = 12) comparing AgentHands with a speech-only baseline. Both conditions used the same researcher-scripted verbal content, so the difference was the presence of embodied hands and synchronized gestures. Participants completed two procedural tasks: orchid care, including identifying plant parts and performing multi-step care activities; and 3D printer operation, including identifying hardware components, powering on the device, inserting an SD card, and navigating the control panel.

Results

The study reports that the combination of XR and co-speech gestures was effective for spatially grounded interactions:

  • Spatial grounding: Participants found it significantly easier to locate objects and identify directions referred to by the agent (p < 0.05).
  • Understanding complex actions: Required activities were rated as easier to follow (p < 0.05).
  • Safety cues: Warnings were more effective when paired with gestural cues and visual effects, such as a “burn” effect for a hot 3D-printer nozzle.
  • Cognitive load: Participants found instructions easier to understand and remember; qualitative feedback described the hands as a “partner guiding me”.

Conclusion and future directions

AgentHands represents a step toward AI systems that do not only analyze the world but dynamically operate within it. By using co-speech gestures and XR spatial grounding, the system aims to reduce the cognitive load of complex tasks and make spatial computing more accessible and human-centric.

The authors are exploring personalization such as adapting gestures to a user’s dominant hand or learning specific spatial routines, with the goal of more seamless human–AI collaboration.

Source attribution and limits

This capture preserves the Google Research article’s description of the AgentHands research prototype and its reported within-subjects study. The article reports a small N = 12 user study and product-style application scenarios; these results should not be treated as a general usability benchmark or evidence of deployment safety.

Labels

  • Human-Computer Interaction and Visualization
  • Machine Intelligence