Notabletraining methods

Telescopic Language Models

Zhilin Guo, Boqiao Zhang, Hakan Aktas, Kyle Fogarty, Nursena Koprucu Aslan, Wenzhao Li, Canberk Baykal, Albert Miao, Siyu Hong, Yixiao Liu, Adam Wu, Ashish Kumar Singh, Sakar Khattar, Chenliang Zhou, Weihao Xia, Cristina Nader Vasconcelos, Cengiz Oztireli

Published
Sep 28, 2026 — 17:59 UTC

Problem

This work addresses the challenge of developing a language model capable of operating efficiently across multiple compute budgets without necessitating distinct training or compression processes. The authors highlight the limitations of existing fixed-exit suites, which do not generalize well across different capacities, often leaving models at chance level performance in non-optimal settings. This paper is a preprint and has not undergone peer review.

Method

The authors propose a novel architecture termed the nested-capacity Transformer. This model employs stochastic prefix supervision with a full anchor, allowing it to leverage a single training run to adapt to various compute budgets. The training mechanism involves one randomly truncated prefix being trained against the full next-token target, alongside a full-capacity pass. This dual approach results in two forward-backward passes per training step. The model is trained on a dataset comprising 20 billion FineWeb-Edu tokens. Notably, the training compute is approximately 12% lower per run compared to traditional fixed-exit suites, enhancing efficiency.

Results

The proposed model demonstrates a 43-44% reduction in area under the quality-budget curve when compared to fixed-exit suites, indicating a significant improvement in perplexity across varying compute budgets. At full capacity, the nested-capacity Transformer matches the performance of fixed-exit suites, showcasing its ability to maintain quality while being adaptable to different operational constraints. The available text does not report quantitative results beyond these metrics.

Limitations

The authors acknowledge that while their nested-capacity model performs well at full capacity, it suffers from limitations inherent to fixed-exit suites, which leave the model at chance level performance in other configurations, with perplexity ranging from 10^2 to 10^5 in baseline comparisons. This indicates that the model's effectiveness may be highly dependent on the specific compute budget utilized.

Why it matters

The implications of this work are significant for the development of more flexible language models that can efficiently allocate resources based on available compute budgets. By eliminating the need for separate training runs for different capacities, this approach could streamline the deployment of language models in resource-constrained environments, potentially leading to broader applications in real-world scenarios where compute resources are variable.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI