LLMs & Generative AI

GPT-6.1 Sol arrives: benchmarks, pricing and the cost of getting work done

GPT-6.1 Sol approaches Astra at lower API token prices. We examine OpenAI benchmarks, independent Artificial Analysis and Vals AI results, availability and limitations.

By 10 min read
Illustration of a processor marked AI on a dark electronic circuit board

GPT-6.1 Sol is available. The useful question behind the version number is whether the new Sol can complete more demanding work within a reasonable time and budget. This article brings together the vendor's release data, independent measurements and practical examples for assessing a model change.

We checked the information and pricing on October 5, 2026. The results come from the linked evaluations; TreffikAI has not run its own GPT-6.1 Sol performance tests. The chart is our visualization of Artificial Analysis data.

GPT-6.1 Sol release date and availability

The OpenAI API changelog confirms the release on September 29, 2026. The model identifier is gpt-6.1-sol. This is an update to the Sol variant, rather than evidence that every member of the GPT-6 family has received a new version.

According to the current availability documentation, the rollout includes Plus, Pro, Business, Enterprise and Edu in Codex and ChatGPT Work. GPT-6.1 Sol is not available in ordinary Chat mode. Free and Go are excluded at launch. Enterprise and Edu administrators must enable it; availability also depends on the client and rollout.

Standard and Fast are listed as launch modes. The documentation still describes Ultrafast for GPT-6.1 Sol as coming later. An announcement of a faster option should not be read as confirmation that every account already has access to it.

Subscription access and API billing are separate questions. A model name in a release article does not establish how many tasks a particular account can perform before reaching its allowance.

OpenAI benchmarks: what does the launch evidence establish?

The OpenAI announcement positions Sol 6.1 closer to Astra. Selected reported changes relative to GPT-6 Sol:

BenchmarkChange reported by OpenAIComparison condition
DeepSWE v1.1+6.4 percentage pointsagainst the predecessor's best score, at lower effort for the new model
AutomationBench 1.0.6+4.8 percentage pointsboth models at Medium
OSWorld 2.0+7.0 percentage pointsboth at Max; offline, partial reward, v2026.08.08
Terminal-Bench Science 0.1more than double the scoreboth models at Max

These are vendor results from its research environment or API. Competitor results in the announcement come from public reports. Differences in effort, tools or budget mean that a comparison is not a controlled test of just one variable.

Our interpretation: progress on software work warrants a trial on your own repository. It does not establish equal performance across languages, test configurations and organizations. Similarly, a computer-use benchmark cannot directly tell you whether an agent will handle your particular accounting application.

We do not add these improvements together into a single percentage increase in intelligence. Each test uses different tasks and scoring. Our glossary entry on model evaluation explains the foundations of these comparisons.

Independent Artificial Analysis benchmarks: new Sol versus old Sol

The following figures come from Artificial Analysis's comparison, with both models at High, read on October 5. We retain the rounded values displayed by the source.

Independent GPT-6.1 Sol and GPT-6 Sol scores at High: Terminal-Bench 52 and 26 percent, AutomationBench-AA 64 and 60, Humanity's Last Exam 51 and 44, GDP.pdf 32 and 24, AA-LCR 82 and 84.

Test or metricGPT-6.1 Sol, HighGPT-6 Sol, High
Intelligence Index5042
Terminal-Bench 4.052%26%
AutomationBench-AA64%60%
Humanity's Last Exam51%44%
GDP.pdf32%24%
AA-LCR v1.182%84%

The Intelligence Index is an index score, not a percentage of correct answers, so it does not appear on the chart's percentage axis. A matching High setting also does not guarantee matching token consumption, time or cost.

Terminal-Bench shows the largest difference in this selection. The AA-LCR result is a useful reminder that a newer version need not win every test. For a project focused on long documents, an improved aggregate index is insufficient grounds for choosing a model.

AutomationBench-AA and the AutomationBench results in OpenAI's announcement belong to different evaluations. Subtracting a score in one publication from a score in another would not establish a model improvement.

What do Vals AI's evaluations show?

A second independent reference is the GPT-6.1 Sol model page on Vals AI. These are selected results read on October 5. The page lists Max as the default effort, while noting that individual benchmarks can use different parameters.

Vals AI benchmarkScoreSource-reported uncertainty
Terminal-Bench 4.055.05%±1.82
Terminal-Bench Science52.86%±6.01
Code Migration65.12%±4.36
Finance Agent v252.03%±0.31
Legal Research Bench38.46%±3.38

These figures describe different task sets. They do not mean that a user will receive exactly those proportions of correct responses in everyday work. We also avoid averaging the five rows: that calculation would lack a clearly defined interpretation.

The inclusion of code migration, finance and legal research is useful when designing your own trial. A team can choose examples from the work it actually performs rather than testing only puzzles. A domain benchmark result does not replace review by someone who understands the data and the acceptance criteria.

Coding, documents and applications: turning the release into a useful trial

For software work, a meaningful trial covers the full cycle: reproducing a bug, changing the code and running appropriate tests. A plausible suggestion in a chat window is insufficient. Check whether the model follows existing conventions, keeps changes within the required scope and can show that the fix works.

For document analysis, ask the model to compare several reports with references to specific pages and tables. Include a question the documents do not answer. This measures both the ability to find evidence and the response to missing evidence.

For application workflows, inspect the final state: whether the correct record was saved, whether a spreadsheet contains the right formula and whether an export is complete. A message saying the work is finished is not proof that it succeeded. Our AI agent entry provides more context for this kind of work.

A white humanoid robot with a tablet beside a wooden wall.

Illustrative photograph: Alex Knight / Unsplash, Unsplash License. This is not a demonstration of GPT-6.1 Sol controlling the robot.

Our suggested starting point is a small collection of repeatable tasks with preserved inputs and explicit success criteria. Record time, cost, corrections and human interventions for each result. That comparison can establish whether the change helps a particular process.

Context window, image input and model specifications

Key specifications from the official model page:

SpecificationGPT-6.1 Sol
API identifiergpt-6.1-sol
Context window1,050,000 tokens
Maximum output128,000 tokens
Training knowledge cutoffApril 30, 2026
Input / outputtext and images / text
Native audio / videounsupported
Fine-tuningunsupported

A large context window is capacity, not a guarantee of correctly interpreting everything inside it. Before supplying a full archive, define the question, the required output and which documents are actually relevant. This also makes it easier to check where an answer came from.

The knowledge cutoff is not the date of the newest information the model can retrieve with a tool. Current events require a current source. Image input also does not mean native photograph generation; producing images requires a separate tool.

API pricing: how much does Sol 6.1 cost compared with Astra?

Standard prices in USD per million tokens, from the official model comparison, for requests with up to 272,000 input tokens:

ModelInputCache hitCache writeOutput
GPT-6.1 Sol$2.00$0.10$2.50$10.00
GPT-6 Sol$2.00$0.20$2.50$10.00
GPT-6 Astra$10.00$1.00$12.50$50.00
GPT-6 Luna$0.10$0.01$0.125$0.50

Our calculation: 100,000 uncached input tokens and 10,000 output tokens, including reasoning, cost $0.30 on Sol 6.1. With identical token counts, Astra costs $1.50. This compares token bills; it is not a measured cost to solve a problem.

If all 100,000 input tokens qualify for a cache hit, the equivalent Sol 6.1 bill is $0.11, excluding the cost of creating the cache. This does not mean every long conversation automatically receives that rate. Context reuse needs to be considered in the application design.

The pricing documentation specifies additional conditions. Above 272,000 input tokens, multipliers apply to the full request: 2× input and cache rates, and 1.5× output. Fast is 2× Standard, Batch and Flex rates are 50% lower, and regional processing adds 10% where offered. Tool fees may also apply.

A production budget should include retries, unfinished tasks and review time. A cheaper token is less useful if an answer repeatedly needs repair; conversely, a higher effort may pay off if it reduces the number of attempts.

API integration: compatibility, reasoning and tools

The GPT-6 guide lists five reasoning.effort settings: low, medium, high, xhigh and max. The default is medium; Sol 6.1 does not support none or minimal.

Tool calling requires the Responses API. Chat Completions remains available for requests without tools. The changelog also lists Multi-agent support in beta.

Below is a minimal request body for POST /v1/responses, following the documentation. This is a configuration example, not a record of a test performed by TreffikAI:

{
  "model": "gpt-6.1-sol",
  "reasoning": { "effort": "medium" },
  "input": "Compare the supplied reports. Identify sources for numerical claims and missing information."
}

When replacing the previous Sol, check parameters, tool support and how you evaluate the result first. If the application used none, replacing the model identifier alone will not be enough. Our guide to writing prompts is a useful companion, particularly when defining the required outcome and task boundaries.

Factual errors and safety: what benchmarks cannot guarantee

The system card addendum covers hallucinations, prompt injection and respecting constraints. OpenAI treats the model's capabilities as Critical in cybersecurity and High in biology and chemistry. These are capability classifications within the vendor's framework, not a certificate of flawless behavior.

In OpenAI's difficult-prompt evaluation at Low, the share of responses with at least one factual error fell from 11.4% to 7.7%. The prompts came from conversations where users had previously reported an error and are not representative of ordinary usage. We do not turn that result into a universal hallucination rate.

A practical trial should include incomplete data, a failed tool and instructions embedded in somebody else's document. Check whether the model reports the problem and respects the task boundaries. Success on a straightforward example does not establish the robustness of an entire workflow.

For automation, retain an opportunity to review an output before sending or publishing it. The review should match the particular action: a working summary and a change to records in a live system call for different acceptance criteria.

Is GPT-6.1 Sol worth adopting?

Our assessment: GPT-6.1 Sol deserves comparison with its predecessor on tasks involving code, documents and applications. The most useful reason to try it is the combination of independently visible progress and unchanged Standard input and output prices for Sol.

We do not assume that it automatically replaces Astra. For the hardest tasks, compare both models on the same inputs and assess the result after verification. Our GPT-6 Astra release article explains the flagship model's positioning; its pricing figures belong to the date of that publication.

Before choosing a migration, answer four questions:

  1. Does the output satisfy the acceptance criterion you defined beforehand?
  2. What does a completed and accepted task cost, including retries?
  3. How long does the user wait for a usable result?
  4. How much checking does a person still need to perform?

If the change improves those measures on real work, there is a sound reason to adopt it. A version number and a single leaderboard position are weaker evidence than a verified result in your own process.

(Sources: OpenAI documentation and the linked OpenAI, Artificial Analysis and Vals AI evaluations. Cover: BoliviaInteligente / Unsplash; inline photo: Alex Knight / Unsplash; image licence. Chart: original TreffikAI visualization.)

Share: