GPT-6 Astra Ultrafast Is 8x Faster. The Price Isn't Public.
OpenAI's GPT-6 Astra Ultrafast runs on NVIDIA Blackwell and claims up to 8x faster token generation, but the launch announcement names no price at all.

OpenAI has released GPT-6 Astra Ultrafast, a faster inference mode it says generates tokens up to 8x faster than Astra's Standard mode. The rate card is the part that is missing.
What shipped
Ultrafast is not a new model. It is a faster inference mode of the existing Astra model, running on NVIDIA Blackwell GPUs, and it is available now through the OpenAI API and to eligible ChatGPT Work and Codex users.
The headline figure is up to 8x faster token generation than Astra Standard. OpenAI frames the gain around agentic workflows: coding agents running edit-test-debug cycles, and any application where a model writes code, calls a tool, checks the result, and decides what to do next.
Why agent loops feel it most
The reasoning OpenAI gives is structural rather than benchmark-based. Shortening each generation step in an agent loop shortens the whole cycle, because the delay repeats every time the agent takes an action.
That framing matters for how you read the number. A one-shot chat completion pays the generation delay once; a multi-step agent run pays it once per step. The 8x figure is stated as an upper bound on token generation, not as an end-to-end task-time guarantee — tool calls, network hops and your own retry logic do not get faster because the decoder did.
Who said what
OpenAI attributes the speedup to work done on NVIDIA's stack. Philippe Tillet, OpenAI's inference lead, said NVIDIA's tooling and documentation let OpenAI's models become "exceptionally good at programming Blackwell and Rubin GPUs," turning that into high-performance kernels that improve latency, throughput and cost on NVIDIA hardware.
Uday Ruddarraju, OpenAI's chief technology officer of compute, said the company used its own models to optimize the inference software running on NVIDIA GPUs, with NVIDIA's programmable platform enabling the acceleration behind Ultrafast. Both are company statements about company work; neither is an independent measurement.
The number OpenAI didn't publish
Developers can access Ultrafast through the API now, but OpenAI points to a separate Ultrafast guide for access, pricing and implementation details rather than stating a price in the announcement itself.
For context on what the baseline costs, a published rate-card comparison of frontier models lists GPT-6 Astra at $10 per 1M input tokens and $50 per 1M output, against $4/$20 for Claude Opus 5.5 and Gemini 4 Argon at regular pricing. Astra was already the most expensive of the three on sticker price before a speed premium entered the picture.
That is the open question for anyone budgeting around this: a mode that generates output tokens up to 8x faster is only a win if the per-token price has not moved to match.
How to evaluate it this week
Three things worth measuring before you switch an agent over. First, time a representative multi-step run end to end on both modes — the per-token speedup is the ceiling, not the result. Second, pull the actual Ultrafast rates from OpenAI's guide and recompute cost per completed task, not cost per million tokens. Third, check whether your agent's bottleneck is generation at all; if most of the wall clock is tool execution, a faster decoder moves very little.
More from DangMua