Basis says GPT-6 Astra completed 50-tab tax workbook in half the time
OpenAI reports a faster result than GPT-5.6 Sol in a Basis test, but has not published the timing protocol or evidence of independently checked accuracy.
Basis completed a 50-tab tax workbook in half the time using GPT-6 Astra rather than GPT-5.6 Sol, according to an OpenAI announcement dated September 28. The reported result could matter to accountants using AI agents for lengthy spreadsheet tasks, but OpenAI has not published the elapsed times or a detailed account of how the comparison was conducted.
OpenAI describes Basis as a company building agents to automate manual accounting work. Its announcement says the comparison involved a complicated workbook with 50 tabs and a task to complete it accurately and reliably. Basis co-founder Mitch Troyanovsky put the speed claim plainly: “GPT-6 Astra is able to complete that workbook in half the time that GPT-5.6 Sol is able to.” The figure is a company-reported result from this test, rather than a general measure of tax-work performance.
What Basis measured in the 50-tab workbook test
The announcement gives the relative result, but no start-to-finish times for either model. It does not say how many times each model ran the task, describe the workbook's contents in enough detail to reproduce the exercise, or specify the tools and settings used for both runs. Those missing details limit what readers can infer from the twofold speed comparison.
OpenAI says Basis also recorded an approximate 20% improvement in its internal evaluation scores with Astra. It attributes that change to the model's interpretation of user intent, including when to ask questions, flag assumptions and follow instructions. The announcement does not provide the underlying scores, a scoring rubric or a breakdown of results, so the reported improvement cannot be translated into a measured reduction in tax errors.
Troyanovsky said Astra makes better decisions near the start of a task, reducing time spent correcting mistakes. OpenAI also says Basis adjusts the model's reasoning effort as a task progresses, using more computation for difficult steps and less for easier ones while retaining its cache. Troyanovsky links that approach to lower token use, cost and response time. Those are Basis's observations about its workflow; the announcement supplies no separate cost figures for the workbook comparison.
Why speed does not establish tax accuracy
A completed workbook and a faster completion time do not, by themselves, show how often an agent gives the right answer across varied tax cases. OpenAI says the task called for accuracy and reliability, but its page does not publish the workbook, an error rate or an independent assessment of the output. It also gives no evidence that the result extends to completed tax returns for customers. The distinction matters because a workflow can finish quickly while still requiring a professional to check its calculations and conclusions.
What accounting-agent benchmarks add to the picture
Mercor and Ramp's APEX-Accounting benchmark tested 160 tasks set in 10 simulated companies, using tasks and grading criteria developed by accounting professionals. Its designers ran models repeatedly because a correct answer on one attempt may not be a dependable outcome. Even the most consistent model completed only 2.6% of tasks correctly across all eight runs, Mercor reported; 58% of tasks were never fully solved by any model on any run.
Those figures describe the APEX-Accounting test, not Basis's tax workbook. Mercor says its benchmark covers month-end close and bookkeeping and expressly excludes tax preparation. Its results are useful context for why repeatability matters in accounting workflows, but differences in tasks, models and methods rule out a direct performance comparison with Astra's reported 50-tab exercise.
Rivet, a tax-technology company, makes a similar distinction in its description of TaxBench: it measures both first-attempt accuracy and whether a model gets the correct answer five times in a row. Rivet says a model can answer correctly once and fail on subsequent attempts. Its published description also says some questions draw on client workflows or internal systems and therefore cannot be released as a clean public dataset. TaxBench does not test the Basis workbook, but its design shows why a single reported completion time leaves reliability unanswered.
What remains unknown about the reported gain
The disclosed comparison establishes what OpenAI says Basis observed on one described workbook task: Astra finished in half the time of Sol. It does not establish how much time accountants would save across other workbooks, whether accuracy changed, or how often the same speed result would recur. Publishing elapsed times, repeated-run results, tool settings and an assessment of completed work would make those questions easier to evaluate. Until then, the twofold figure is best understood as a reported result from Basis's test.
Sources and context
- Basis completes a tax workbook 2x faster with GPT-6 AstraOpenAI
- APEX-Accounting: AI Productivity Benchmark for AccountingMercor, with Ramp
- TaxBench by RivetRivet
- Performance of LLMs on VITA test: potential for AI-assisted tax returns for low income taxpayersArtificial Intelligence and Law / Springer Nature
AI-assisted article checked against the listed sources. NewsJaws did not conduct interviews or attend the reported events.
About NewsJaws Desk
AI-assisted reporting and explainers reviewed against the linked source documents. No claim of on-scene reporting or original interviews.