OpenAI Launches Improved Prompt Caching for GPT-6, Reducing Costs by 36%
- Published
- Sep 22, 2026 — 21:00 UTC
OpenAI's recent launch of an improved prompt caching system for GPT-6 on September 22, 2026, has resulted in a 36% reduction in inference costs. This enhancement allows for a 90% discount on cached input tokens and a 50% reduction in the share of prompt tokens requiring fresh processing. The system has achieved an increase in cache hit rates from 85% to 90% for the OpenAI model and from 83% to 91% for session agents, significantly improving efficiency.
Mario Rodriguez, OpenAI's Chief Product Officer, stated that the prompt caching plays a critical role in enabling GitHub Copilot to deliver fast, efficient experiences at scale. Arian Hanifi, Chief Technology Officer, noted that the caching diagnostics and dashboard led to a 20% cost reduction for their applications. Additionally, Bin Fan, Agent Team Lead, emphasized the importance of reliable caching for long-running agents like Manus, which are fundamental to economic viability.
In a specific case, Eugene Mikhantyev, an AI Engineer, reported that cache hit rates on evaluations improved from 83% to 91% within a week, demonstrating the system's effectiveness. The new caching system also resulted in a two-thirds reduction in cache writes, further optimizing performance. This follows OpenAI's recent initiatives to enhance AI's business value, including the introduction of an AI advertising product and a guide on connecting AI usage to business outcomes on September 16, 2026.
By Callan Zhang · Sep 22, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: OpenAI Blog
