AINews

OpenAI’s GPT-6 Astra Ultrafast launches with restricted Work and Codex access

NVIDIA says the Blackwell-powered service offers up to eight times faster token generation than Astra Standard. OpenAI says all API users have limited access, while Work and Codex access depends on plan and workspace.

Exterior of NVIDIA headquarters in Santa Clara, California, with a broad entrance and landscaped grounds
File photograph of NVIDIA’s headquarters in Santa Clara, California, taken on 4 August 2018. Coolcaesar, ‘NVIDIA Headquarters.jpg’ (perspective correction and crop by Jacek Halicki) (resized and converted to WebP). CC BY-SA 4.0.
LinkedInPostEmail
Save for later

OpenAI’s GPT-6 Astra Ultrafast became available through its API and to eligible ChatGPT Work and Codex users on 1 October, according to an NVIDIA announcement published that day. The service runs on NVIDIA Blackwell GPUs, and NVIDIA says it can generate tokens up to eight times faster than Astra Standard. For people building applications and coding agents, the claim points to less waiting for generated output, although the announcement does not establish the same improvement for a complete task.

OpenAI’s API guide says Ultrafast is available to all API users at low rate limits. Access in Work and Codex is narrower: OpenAI’s help page says that, at launch, it is available on Pro $500 and in eligible Enterprise and Edu workspaces, subject to workspace permissions. Plus, Pro $100, Pro $200 and Business users cannot access Ultrafast in Work and Codex at launch, including through credits.

What NVIDIA’s eightfold speed claim measures

NVIDIA describes the improvement as up to eight times faster token generation than Astra Standard. That comparison is with another speed mode for Astra, rather than with a different model. The phrase ‘up to’ describes a possible result, but the announcement does not say what result a typical user should expect. OpenAI’s API documentation identifies Ultrafast as a separate service tier and advises developers to use it when speed justifies the higher cost.

The NVIDIA announcement does not provide the test setup, workload, absolute Standard-mode token rate or range of observed results behind its figure. It also does not report an independently reproducible comparison. That leaves readers unable to calculate how much time an ordinary request would save from the published claim alone.

Token generation is one part of an application’s response time. OpenAI warns in its Ultrafast guide that network overhead can reduce the latency gains for applications that make many tool calls without a persistent connection. It recommends WebSockets for those applications. Consequently, NVIDIA’s token-generation figure should not be read as a promise that a complete answer or coding task will finish eight times faster.

Who can use Astra Ultrafast

For API developers, OpenAI says GPT-6 Astra Ultrafast is currently available to all users, subject to low rate limits. Its guide instructs developers to select the gpt-6-astra model and the ultrafast service tier for each request. Organizations working with an OpenAI account team can ask for higher rate limits, the guide says. These access terms describe the API; they do not extend the same entitlement to every ChatGPT subscription.

In Work and Codex, OpenAI says Pro $500 users can draw on included Work and Codex usage and credits for Ultrafast. Eligible Enterprise and Edu workspaces are charged workspace credits, and their access also depends on workspace permissions. OpenAI’s help page lists Plus, Pro $100, Pro $200 and Business among plans without Work and Codex Ultrafast access at launch. Eligible users select GPT-6 Astra and then Ultrafast in the model picker.

NVIDIA says faster generation could shorten coding agents’ edit, test and debugging cycles and reduce waiting between tool calls. Those are proposed uses of the service, rather than measured improvements in a published coding workflow. In an agent that repeatedly writes code, calls a tool and evaluates the result, time spent generating each response may add up across steps. OpenAI’s warning about connection overhead shows why application setup still matters.

How the companies say the speed was achieved

NVIDIA attributes Ultrafast’s performance to inference optimizations that use OpenAI models and capabilities of its Blackwell GPU architecture. In NVIDIA’s announcement, OpenAI inference lead Philippe Tillet said the companies’ work had helped OpenAI’s models produce high-performance kernels for NVIDIA hardware. Uday Ruddarraju, OpenAI’s chief technology officer of compute, said internal models had been used to optimize inference on NVIDIA GPUs.

The announcement also describes ongoing work to refine the software used to run models on NVIDIA hardware. It does not quantify how much of the advertised token-generation gain comes from a particular optimization, nor does it supply a separate result for any one application. The measurable public claim remains the companies’ up-to-eight-times comparison with Astra Standard.

Developers considering Ultrafast can consult OpenAI’s API guide for implementation, rate limits and its pricing link. For Work and Codex users, the relevant question is whether their plan and workspace currently permit the mode. Neither the announcement nor the cited product guidance establishes a general improvement in end-to-end task completion, and actual access may change as OpenAI updates its service terms.

Sources and context

AI-assisted article checked against the listed sources. NewsJaws did not conduct interviews or attend the reported events.

About NewsJaws Desk

AI-assisted reporting and explainers reviewed against the linked source documents. No claim of on-scene reporting or original interviews.