LLMs & Generative AI

GPT-6 Astra arrives: benchmarks, pricing and what changes for AI work

GPT-6 Astra targets complex work across code, browsers and documents. We examine OpenAI's benchmark results, API costs, rollout and the limits behind the launch.

By 6 min read
Close-up photograph of chips and components on a laptop circuit board

GPT-6 Astra is OpenAI's new flagship for complex work involving reasoning, code and computer use. The practical question is whether it can finish a demanding assignment with less intervention: navigating an application, checking a change and delivering a usable result.

This article reflects information checked on September 5, 2026. Availability is still being phased in; a launch announcement does not mean every account already has access.

When is GPT-6 Astra available?

OpenAI's safety overview dated September 3 places the release in early September. The API documentation describes an initial rollout to enterprises in Trusted Access, followed by API access and ChatGPT Plus, Pro, Business and Enterprise in the coming days.

For readers deciding whether to subscribe or migrate an application, the useful check is the actual model selection on their account. Announced eligibility, visible access and a particular usage allowance are different things. These sources do not establish one universal message limit for all users.

The previous generation remains useful background: our GPT-5.6 launch coverage describes its original rollout. Prices in that older article are historical, rather than a current price list.

GPT-6 Astra benchmarks in one chart

The graphic below is our visualization of selected results in OpenAI's launch evaluation. These are vendor-reported results, not tests conducted by TreffikAI. Higher is better within each test; percentages from different tests are not interchangeable.

Selected OpenAI benchmarks: GPT-6 Astra and GPT-5.6 Sol compared on six tests; exact values appear in the table below.

EvaluationGPT-6 AstraGPT-5.6 Sol
ARC-AGI-399.9% [T7]7.8%
Terminal-Bench 4.057.9%37.3%
AutomationBench41.4%18.1%
OSWorld 2.072.6%65.7%
ScreenSpot-Pro, no tools92.7%76.9%
Agents' Last Exam59.3%53.6%

OpenAI reports best scores across effort settings, not a fixed equal-budget comparison. OSWorld uses v2026.08.08, the offline set and partial scoring. The ARC-AGI-3 result retains the source's T7 label. Research and API conditions can differ from ChatGPT.

What do those results actually tell us?

The largest gap in this selection is on ARC-AGI-3. It deserves attention, but a near-perfect benchmark score is not proof that a system can perform every human task. A benchmark defines a particular environment, scoring rule and set of problems.

For a team buying an assistant, the terminal and application results may be more directly useful. They give reasons to test Astra on workflows with several dependent steps. At the same time, scores well below 100% are a reminder to measure incomplete work and recovery from mistakes.

Our interpretation: treat the chart as a shortlist for your own evaluation. Run the same repository issue, spreadsheet task or browser workflow on both models, with the same tools and permissions. Record correctness, time, cost and the number of times a person has to intervene.

Context window and supported inputs

The model card lists a 1,050,000-token context window, up to 128,000 output tokens and an April 30, 2026 knowledge cutoff. Astra accepts text and images and produces text. Native audio and video are unsupported; generating an image through a separate tool is a different capability.

A large context window creates room for more source material, but it is not a guarantee that every detail will be used correctly. A useful document-analysis test includes a question whose answer is buried near the end of the input, plus a question the documents cannot answer.

The second test matters just as much: a polished answer is only useful if the model distinguishes evidence from missing information.

API pricing: calculate the whole task

Standard rates in the official model card, in USD per million tokens:

Token categoryPrice
Input$10
Cached input read$1
Cache write$12.50
Output$50

Above 272,000 input tokens, the full request uses twice the input/cache rates and 1.5 times the output rate. Tool charges can add to the bill.

For a simplified example, 100,000 uncached input tokens plus 10,000 billed output tokens cost $1.50, before tools and other applicable charges. This is arithmetic using the listed Standard rates, not a measured cost for an Astra workflow.

The economic question is cost per accepted result. A cheaper request can become expensive after repeated failures; an expensive model can also waste money on a task that a smaller one already handles reliably. Compare total bills alongside completion rates.

New features for longer agent workflows

The Astra developer guide describes asynchronous tool calls, instructions sent while a task is running, and reasoning-effort changes that preserve the cached prompt prefix. The model ID is gpt-6-astra; supported effort levels are low, medium, high, xhigh and max.

These features matter most when work takes time. For example, a user could correct an export requirement while an agent is still collecting inputs. An application must still manage pending tools and apply the updated instructions consistently.

The guide also notes two limits: there is no none reasoning setting, and Fast mode is unavailable with EU data residency. That second restriction concerns the API's data-residency configuration, rather than simply where a reader lives.

Safety is part of this release

In its safety overview, OpenAI classifies Astra at the Critical cybersecurity capability threshold in its Preparedness Framework. That is a capability assessment within OpenAI's framework, not a declaration that every deployed interaction is dangerous.

The company describes stronger protections and improved resistance to prompt injection, while also discussing residual risks. Better evaluation results should not be interpreted as a guarantee that a model will always follow the intended boundaries.

For practical adoption, test permissions alongside accuracy. A browser assistant that completes a form correctly still needs a clear rule about whether it may submit it. A coding assistant should be evaluated on the changes it makes and the evidence it provides that they work.

Should you switch to Astra?

Start with the work that currently consumes the most review time: difficult debugging, research involving many sources, or tasks that cross files and applications. Keep a small, repeatable set of examples so you can distinguish an actual improvement from one impressive demonstration.

For simpler classification or extraction, first establish whether the existing model is already good enough. The launch alone is not a reason to replace a working, economical setup.

Our ChatGPT vs Claude vs Gemini guide explains how to compare assistants by workflow. Astra adds a strong new candidate to that decision; the best choice still depends on the job, available access and the cost of a result you can trust.

Photo: Alexandre Debiève / Unsplash, used under the Unsplash License. Illustrative hardware photograph; it does not depict Astra's infrastructure. Benchmark graphic: TreffikAI, based on the OpenAI evaluation linked above.

Share: