OpenAI releases GPT-6 Astra Ultrafast faster-response mode for Nvidia Blackwell
OpenAI has released GPT-6 Astra Ultrafast, a faster-response mode of its model that runs on Nvidia’s Blackwell GPU architecture, according to a release posted to Nvidia’s newsroom. The mode is aimed at coding agents and other latency-sensitive tools that call models repeatedly inside a workflow.
- Ultrafast is available now in the OpenAI API and to eligible ChatGPT Work and Codex users
- Nvidia says Ultrafast offers “up to 8x faster token generation than the Astra Standard mode”
- Watch for pricing and implementation detail in OpenAI’s separate Ultrafast guide, not published in this release
Nvidia detailed the launch in a release crediting “inference optimizations through OpenAI’s models that tap into the capabilities of the NVIDIA Blackwell architecture” for the speed gain. The release frames the improvement around agentic workflows, where a model “writes code, uses a tool, checks the result and decides what to do next,” and where repeated latency compounds across a session.
OpenAI says it used its own models to tune Nvidia’s chips
The release states that OpenAI is using its own models to refine the inference software that runs on Nvidia GPUs, describing this as an ongoing process rather than a one-time optimization. Nvidia’s programmability, it says, let OpenAI “test and implement improvements” to that software over time.
Philippe Tillet, OpenAI’s inference lead, is quoted saying: “NVIDIA’s deep investment in tooling and documentation has enabled us to make our models exceptionally good at programming Blackwell and Rubin GPUs.” Uday Ruddarraju, OpenAI’s chief technology officer of compute, added: “We used our internal models to optimize inference on NVIDIA GPUs, and NVIDIA’s programmability helped us deliver the acceleration behind Astra Ultrafast.”
What the release does not specify
Nvidia’s release does not disclose a price for Astra Ultrafast access, nor does it define what makes a ChatGPT Work or Codex user “eligible” for the mode. It also does not publish a benchmark methodology behind the 8x figure, leaving open whether that multiple holds across task types, model sizes or real customer workloads rather than a specific internal test.
The release points developers to “the Ultrafast guide for access, pricing and implementation details,” but that guide’s contents are not reproduced here. Nvidia also does not state how many developers or enterprise customers have access to Ultrafast today, or whether capacity constraints could limit the rollout.
Why the framing matters for compute buyers
Nvidia’s release leans on a broader argument it has made in related material published alongside this piece: that GPUs bought for training can be repurposed for inference and reinforcement learning as model needs shift, improving utilization. The release states that “a programmable NVIDIA platform allows developers and researchers to reuse infrastructure across training, inference and reinforcement learning as models evolve,” which it frames as a way to avoid overprovisioning hardware for a single workload.
That framing is relevant to anyone sizing Nvidia’s data center demand, since it ties a customer-facing product launch directly to Nvidia’s pitch that Blackwell deployments retain value across multiple AI workload types rather than becoming stranded capacity tied to one model generation.
The BlockWest read. Nvidia is using OpenAI’s own customer-facing speed claim to argue its chips hold value beyond initial training runs, which matters more to data center capex debates than to Ultrafast users themselves. The release offers no price, no independent benchmark and no user count, so the 8x figure functions as marketing language from both companies until OpenAI’s Ultrafast guide fills in the commercial terms.
The next concrete step is OpenAI’s publication of pricing and implementation details in its Ultrafast guide, referenced but not included in Nvidia’s release.
BlockWest is a news publication. Nothing here is investment advice. Read our disclaimer and editorial policy.
