Notableevaluation benchmarksDeepSeek

Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark

Makar Ulesov, Vladislav Smirnov, Omar Ibrahim, Arsenii Bobovnikov

Published
Sep 23, 2026 13:25 UTC

{'Problem': 'The paper addresses the gap in backtest auditing, framing it as a calibration problem where high flaw recall is insufficient if the model generates false positives by flagging clean strategies. This issue is critical for ensuring the reliability of financial models, particularly in the context of machine learning applications in finance. The work is presented as a preprint and has not undergone peer review.', 'Method': 'The authors propose a benchmark consisting of 96 paired items, each containing flawed backtests and corresponding clean controls. The evaluation employs a deterministic scoring system that assesses flaw recall, clean-control false positives, evidence localization, and fix relevance. The dataset comprises 1440 cached audits sourced from four text endpoints. The model utilized for this task is named DeepSeek auditor, which is specifically designed to enhance the accuracy of backtest audits.', 'Results': "The results demonstrate impressive performance metrics: Closed and clean-aware code recall achieved 100.0%, while clean code control over-flagging was recorded at 93.8%. The clean-aware specificity across all three metrics reached 87.5%. Notably, DeepSeek's code false positives were 20.8% (95% CI 11.7–34.3) before implementing clean-aware warnings, which were reduced to 0.0% (0.0–7.4) after the warnings were applied. The clean-control rate separation was quantified at 79 points. The available text does not report quantitative results for any baselines.", 'Limitations': 'The authors do not report any limitations in their study, and no obvious limitations are identified in the provided text.', 'Why it matters': 'This work has significant implications for the field of financial machine learning, particularly in enhancing the reliability of backtest evaluations. By providing a robust benchmark and demonstrating the effectiveness of the DeepSeek auditor, the authors pave the way for improved auditing practices, which can lead to more trustworthy financial models and strategies.'}

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI