Notableagents robotics

DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation

Haoran Yuan, Zekai Wang, Boning Shao, Haoran Lu, Trevor Darrell, Ismini Lourentzou, Wei Zhan

Published
Sep 21, 2026 17:55 UTC

Problem

The paper addresses a significant gap in the modeling of contact dynamics in dexterous manipulation tasks. Existing models often overlook the integration of tactile feedback, which is crucial for effective manipulation in real-world scenarios. This work is particularly relevant as it explores the intersection of visual and tactile information, which has not been extensively covered in the literature. The research is presented as a preprint and has not yet undergone peer review.

Method

The authors propose the Visuo-tactile World-Action Model (DexTacWAM), which consists of several key components:

  • Independent Fingertip Encoding: Each fingertip is encoded separately to capture the unique tactile information from each contact point.
  • Tactile Compressor: A finger- and pose-aware tactile compressor aggregates features from the fingertip encodings, allowing for a more nuanced understanding of the manipulation context.
  • Tactile Latent Injection: The model injects tactile latent variables into a video diffusion world model, enhancing the model's ability to predict outcomes based on both visual and tactile inputs.

The model is trained on six contact-rich dexterous manipulation tasks using a 22-DoF bimanual platform. The training process involves four hours of tactile-encoder adaptation with a frozen pretrained vision Variational Autoencoder (VAE). Approximately 100 demonstrations per task were collected without midtraining on tactile data.

Results

The results demonstrate significant improvements over existing baselines:

  • The average score achieved by DexTacWAM is 70.6, compared to 38.0 for the strongest baseline.
  • When evaluating a four-task mean with tactile world modeling, DexTacWAM scores 74.7, while the model without tactile world modeling scores only 26.6.
  • The visual prediction quality of DexTacWAM is within 0.5 dB of vision-only counterparts, indicating strong performance in visual tasks.
  • The model retains a contact recall of 89.4%, showcasing its effectiveness in maintaining contact information during manipulation.
  • Training speed is improved by a factor of 2.26x, and inference speed is enhanced by 1.29x, indicating efficiency gains in both training and deployment phases.

Limitations

The authors do not report any limitations in their work. However, it is important to note that the absence of reported limitations does not imply that the model is without potential weaknesses or areas for improvement.

Why it matters

The implications of this work are significant for downstream applications in robotics and automation, particularly in tasks requiring fine motor skills and dexterity. By effectively integrating tactile feedback with visual information, DexTacWAM could lead to advancements in robotic manipulation, enhancing the ability of robots to perform complex tasks in unstructured environments. This research opens avenues for further exploration into multi-modal sensory integration in machine learning and robotics.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI