Notableagents robotics

X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets

Prithwish Dan, Chenyang Ma, Wei Zhan

Published
Sep 28, 2026 — 17:46 UTC

Problem

The paper addresses the severe exploration problem encountered in training generalist policies with task-agnostic rewards in reinforcement learning (RL). This issue is particularly pronounced when attempting to generalize across multiple objects and tasks, which is critical for developing robust RL agents. The work is presented as a preprint and has not undergone peer review.

Method

The authors propose a novel framework called X-Reset, which leverages human hand-object demonstrations to enhance exploration in RL. Key components of the method include:

  • Kinematic Retargeting: The framework kinematically retargets hand-object states to noisy robot states, allowing the agent to learn from human-like interactions.
  • State Filtering: It filters out unstable states in simulation, ensuring that only reliable states are used for training.
  • State Sampling: The remaining stable states are sampled as resets during RL training, facilitating better exploration of the state space.
  • Reward Structure: The framework employs general-purpose object-centric rewards that are designed to be independent of specific tasks, focusing instead on the object state and the goal.
  • Training Objects: The experiments involve 20 different objects across three embodiments, including a 22-DoF hand on two robotic arms and a parallel-jaw gripper.

Results

The results indicate that the X-Reset framework successfully trains generalist policies on the 20 objects, resolving exploration challenges effectively. The framework demonstrates the ability to generalize to unseen objects and can learn from imperfect hand-pose estimates. Additionally, it shows the capability of transferring learned behaviors zero-shot from simulation to real-world applications. However, the available text does not report quantitative results or specific baselines for comparison.

Limitations

The authors do not report any limitations in their work. However, the lack of quantitative results and baseline comparisons may hinder the assessment of the framework's performance relative to existing methods.

Why it matters

The implications of this work are significant for downstream applications in robotics and RL, particularly in scenarios requiring generalization across diverse tasks and objects. By addressing exploration challenges and enabling effective behavior transfer, X-Reset could facilitate the development of more capable and adaptable RL agents, paving the way for advancements in autonomous systems.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI