Notableagents robotics

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer, Emiliano Penaloza, Zhengyan Shi, Alessandro Sordoni, Marc-Alexandre Côté, Xingdi Yuan, Minseon Kim

Published
Sep 28, 2026 — 17:54 UTC

Problem

This work addresses the gap in understanding how explanation-only training influences agent behavior, specifically in the context of software engineering tasks. The authors explore the potential of fine-tuning models using explanations rather than traditional reinforcement learning (RL) methods. This is particularly relevant as it offers a novel approach to improving model performance without the complexities associated with RL, which can be resource-intensive and challenging to implement.

Method

The authors propose a training procedure called Retrospection-Only Fine-Tuning (ROFT). This method involves fine-tuning a pre-trained model, specifically Qwen3.5-4B, using next-token prediction loss applied to explanation tokens derived from software-engineering problems. The training data consists of a mix of successful and unsuccessful attempts at solving these problems. The model undergoes 20 updates during the ROFT process, while a baseline method, Generalized Retrospection with Policy Optimization (GRPO), is trained with 40 updates for comparison. This setup allows for a direct evaluation of the effectiveness of explanation-based training against a more traditional approach.

Results

The results demonstrate that the ROFT method achieves a SWE-bench Verified Solve Rate of 49.2%, outperforming the GRPO baseline, which achieved a 48.0% solve rate. Additionally, the SWE-bench Pro Solve Rate for ROFT is 26.8%, compared to GRPO's 25.3%. Notably, ROFT exhibits faster early progress in training time and sampled attempts, indicating improved training efficiency. Furthermore, ROFT successfully solved tasks that all 64 sampled attempts from the base model failed to address, showcasing its potential to enhance model capabilities significantly.

Limitations

The authors do not report any limitations in their study. However, it is important to note that the absence of reported limitations does not imply that the method is without potential drawbacks or areas for improvement. Future work may need to explore the scalability of ROFT across different domains and its performance in more complex scenarios.

Why it matters

The implications of this research are significant for the development of agentic models in AI. By demonstrating that explanation-only training can enhance model performance without relying on RL, this work opens new avenues for training methodologies that are less resource-intensive and potentially more interpretable. This could lead to broader applications in areas where understanding model decisions is crucial, such as software engineering, and may inspire further research into the role of explanations in machine learning.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI