OpenAI launches GPT-6.1 Sol with lower cached-input pricing and new benchmark claims
OpenAI says the upgraded model is available in ChatGPT Work, Codex and its API. Its standard input and output prices match GPT-6 Sol, while cached input costs half as much.
OpenAI introduced GPT-6.1 Sol on September 29 for users of ChatGPT Work and Codex and for developers using its API. The company says the upgrade improves coding, computer use and professional work while charging $2 per million standard input tokens and $10 per million output tokens. Those prices are one-fifth of GPT-6 Astra’s listed standard rates; the new Sol model’s cached-input price is $0.10 per million tokens.
Developers can call the model as gpt-6.1-sol, according to OpenAI. It says the model is available to Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, but is not yet available in Chat. The Associated Press also reported the model announcement from OpenAI’s September 29 developer conference.
How GPT-6.1 Sol’s API prices compare
The September 22 GPT-6 Sol announcement listed $2 per million standard input tokens, $0.20 per million cached input tokens and $10 per million output tokens. On the new price list, standard input and output rates stay at those levels, while cached input falls to $0.10 per million tokens. OpenAI describes that cached rate as 95% below the new model’s standard input price and 50% below the earlier Sol rate.
The difference matters most to developers whose applications reuse context across requests: cached input has a separate, lower listed rate. The announcement gives prices per million tokens, rather than a fixed price for completing a task. OpenAI’s comparisons of cost per completed task therefore depend on the particular evaluation, its settings and how much work the models do within it.
What OpenAI’s coding and computer-use tests show
On DeepSWE v1.1, OpenAI says GPT-6.1 Sol matches GPT-6 Astra at roughly one-fifth of the cost and exceeds GPT-6 Sol’s best score by 6.4 percentage points at lower reasoning effort and cost. DeepSWE assesses agents on long software-engineering tasks in real codebases. Epoch AI describes its test set as 113 original tasks across 91 active open-source repositories and five programming languages. The agents run in sandboxed containers without internet access.
That coding result has a qualification. In a September 7 review, Epoch AI identified issues in at least 23 of DeepSWE’s 113 tasks, more than 20% of the set, and designated the benchmark flawed. The confirmed issues were false negatives, including cases in which verifier behavior could cause a submission to fail grading. The available evidence does not establish whether those problems affected OpenAI’s particular comparison or whether it used a corrected task set.
For computer use, OpenAI reports that GPT-6.1 Sol beat GPT-6 Sol by seven percentage points on the offline OSWorld 2.0 set at maximum reasoning effort while costing less than half as much per task. It says the new model finished within 2.1 points of Astra at roughly one-seventh of Astra’s cost per task. OpenAI specifies that its reported measure is partial reward on the offline set from the v2026.08.08 release, rather than a measure of everyday use.
Professional-work and factuality claims
OpenAI says GPT-6.1 Sol scored above Claude Opus 5.5 with fallbacks on GDP.pdf, its evaluation of questions about complex professional documents, at less than half the cost per task across the tested reasoning settings. It says the new model approached Astra’s score at roughly one-fifth of Astra’s cost per task. The documents include tables, charts, diagrams and fine print from fields including finance, healthcare and law.
On AutomationBench, OpenAI reports a score 2.2 percentage points above Opus 5.5 at medium reasoning effort and roughly one-third of its cost per task. It also reports a 4.8-point improvement over GPT-6 Sol at the same setting. The test covers workflows using 47 tools across sales, marketing, operations, support, finance and human resources. These are benchmark comparisons under stated settings, not independently established savings for customers’ own workflows.
The company reports a lower factual-error rate on a separate low-effort test: 7.7% of responses for GPT-6.1 Sol contained an error, compared with 11.4% for GPT-6 Sol. OpenAI says the test used de-identified ChatGPT conversations in which users had flagged an earlier model’s mistake. It cautions that these deliberately difficult prompts are not representative of typical use.
Availability and the limits of the comparisons
OpenAI says its GPT evaluations were conducted in its research environment or through its API. Results may differ from production ChatGPT because system prompts, available tools and effort settings can vary; the company says competitor-model results came from publicly available reports. The announcement does not provide independently reproduced results for GPT-6.1 Sol, so the performance and cost figures remain attributed company claims.
OpenAI also says it plans to offer GPT-6.1 Sol Ultrafast in Codex in the coming days, describing token generation as up to eight times faster than the standard option. That is an announced plan, with no subsequent availability established by the sources here. At launch, the confirmed access routes are ChatGPT Work, Codex and the API; OpenAI says the model is not yet in Chat.
Sources and context
- Introducing GPT-6.1 SolOpenAI
- OpenAI CEO announces new AI agent and avoids mention of security concerns at developer conferenceThe Associated Press
- DeepSWE v1.1Epoch AI
- DeepSWE v1.1 – Benchmark reviewEpoch AI
- Introducing GPT-6 Sol and LunaOpenAI
AI-assisted article checked against the listed sources. NewsJaws did not conduct interviews or attend the reported events.
About NewsJaws Desk
AI-assisted reporting and explainers reviewed against the linked source documents. No claim of on-scene reporting or original interviews.