GPT-6.1 Sol arrives: benchmarks, pricing and the cost of getting work done
GPT-6.1 Sol approaches Astra at lower API token prices. We examine OpenAI benchmarks, independent Artificial Analysis and Vals AI results, availability and limitations.

GPT-6.1 Sol is available. The useful question behind the version number is whether the new Sol can complete more demanding work within a reasonable time and budget. This article brings together the vendor's release data, independent measurements and practical examples for assessing a model change.
We checked the information and pricing on October 5, 2026. The results come from the linked evaluations; TreffikAI has not run its own GPT-6.1 Sol performance tests. The chart is our visualization of Artificial Analysis data.
GPT-6.1 Sol release date and availability
The OpenAI API changelog confirms the release on September 29, 2026. The model identifier is gpt-6.1-sol. This is an update to the Sol variant, rather than evidence that every member of the GPT-6 family has received a new version.
According to the current availability documentation, the rollout includes Plus, Pro, Business, Enterprise and Edu in Codex and ChatGPT Work. GPT-6.1 Sol is not available in ordinary Chat mode. Free and Go are excluded at launch. Enterprise and Edu administrators must enable it; availability also depends on the client and rollout.
Standard and Fast are listed as launch modes. The documentation still describes Ultrafast for GPT-6.1 Sol as coming later. An announcement of a faster option should not be read as confirmation that every account already has access to it.
Subscription access and API billing are separate questions. A model name in a release article does not establish how many tasks a particular account can perform before reaching its allowance.
OpenAI benchmarks: what does the launch evidence establish?
The OpenAI announcement positions Sol 6.1 closer to Astra. Selected reported changes relative to GPT-6 Sol:
| Benchmark | Change reported by OpenAI | Comparison condition |
|---|---|---|
| DeepSWE v1.1 | +6.4 percentage points | against the predecessor's best score, at lower effort for the new model |
| AutomationBench 1.0.6 | +4.8 percentage points | both models at Medium |
| OSWorld 2.0 | +7.0 percentage points | both at Max; offline, partial reward, v2026.08.08 |
| Terminal-Bench Science 0.1 | more than double the score | both models at Max |
These are vendor results from its research environment or API. Competitor results in the announcement come from public reports. Differences in effort, tools or budget mean that a comparison is not a controlled test of just one variable.
Our interpretation: progress on software work warrants a trial on your own repository. It does not establish equal performance across languages, test configurations and organizations. Similarly, a computer-use benchmark cannot directly tell you whether an agent will handle your particular accounting application.
We do not add these improvements together into a single percentage increase in intelligence. Each test uses different tasks and scoring. Our glossary entry on model evaluation explains the foundations of these comparisons.
Independent Artificial Analysis benchmarks: new Sol versus old Sol
The following figures come from Artificial Analysis's comparison, with both models at High, read on October 5. We retain the rounded values displayed by the source.

| Test or metric | GPT-6.1 Sol, High | GPT-6 Sol, High |
|---|---|---|
| Intelligence Index | 50 | 42 |
| Terminal-Bench 4.0 | 52% | 26% |
| AutomationBench-AA | 64% | 60% |
| Humanity's Last Exam | 51% | 44% |
| GDP.pdf | 32% | 24% |
| AA-LCR v1.1 | 82% | 84% |
The Intelligence Index is an index score, not a percentage of correct answers, so it does not appear on the chart's percentage axis. A matching High setting also does not guarantee matching token consumption, time or cost.
Terminal-Bench shows the largest difference in this selection. The AA-LCR result is a useful reminder that a newer version need not win every test. For a project focused on long documents, an improved aggregate index is insufficient grounds for choosing a model.
AutomationBench-AA and the AutomationBench results in OpenAI's announcement belong to different evaluations. Subtracting a score in one publication from a score in another would not establish a model improvement.
What do Vals AI's evaluations show?
A second independent reference is the GPT-6.1 Sol model page on Vals AI. These are selected results read on October 5. The page lists Max as the default effort, while noting that individual benchmarks can use different parameters.
| Vals AI benchmark | Score | Source-reported uncertainty |
|---|---|---|
| Terminal-Bench 4.0 | 55.05% | ±1.82 |
| Terminal-Bench Science | 52.86% | ±6.01 |
| Code Migration | 65.12% | ±4.36 |
| Finance Agent v2 | 52.03% | ±0.31 |
| Legal Research Bench | 38.46% | ±3.38 |
These figures describe different task sets. They do not mean that a user will receive exactly those proportions of correct responses in everyday work. We also avoid averaging the five rows: that calculation would lack a clearly defined interpretation.
The inclusion of code migration, finance and legal research is useful when designing your own trial. A team can choose examples from the work it actually performs rather than testing only puzzles. A domain benchmark result does not replace review by someone who understands the data and the acceptance criteria.
Coding, documents and applications: turning the release into a useful trial
For software work, a meaningful trial covers the full cycle: reproducing a bug, changing the code and running appropriate tests. A plausible suggestion in a chat window is insufficient. Check whether the model follows existing conventions, keeps changes within the required scope and can show that the fix works.
For document analysis, ask the model to compare several reports with references to specific pages and tables. Include a question the documents do not answer. This measures both the ability to find evidence and the response to missing evidence.
For application workflows, inspect the final state: whether the correct record was saved, whether a spreadsheet contains the right formula and whether an export is complete. A message saying the work is finished is not proof that it succeeded. Our AI agent entry provides more context for this kind of work.

Illustrative photograph: Alex Knight / Unsplash, Unsplash License. This is not a demonstration of GPT-6.1 Sol controlling the robot.
Our suggested starting point is a small collection of repeatable tasks with preserved inputs and explicit success criteria. Record time, cost, corrections and human interventions for each result. That comparison can establish whether the change helps a particular process.
Context window, image input and model specifications
Key specifications from the official model page:
| Specification | GPT-6.1 Sol |
|---|---|
| API identifier | gpt-6.1-sol |
| Context window | 1,050,000 tokens |
| Maximum output | 128,000 tokens |
| Training knowledge cutoff | April 30, 2026 |
| Input / output | text and images / text |
| Native audio / video | unsupported |
| Fine-tuning | unsupported |
A large context window is capacity, not a guarantee of correctly interpreting everything inside it. Before supplying a full archive, define the question, the required output and which documents are actually relevant. This also makes it easier to check where an answer came from.
The knowledge cutoff is not the date of the newest information the model can retrieve with a tool. Current events require a current source. Image input also does not mean native photograph generation; producing images requires a separate tool.
API pricing: how much does Sol 6.1 cost compared with Astra?
Standard prices in USD per million tokens, from the official model comparison, for requests with up to 272,000 input tokens:
| Model | Input | Cache hit | Cache write | Output |
|---|---|---|---|---|
| GPT-6.1 Sol | $2.00 | $0.10 | $2.50 | $10.00 |
| GPT-6 Sol | $2.00 | $0.20 | $2.50 | $10.00 |
| GPT-6 Astra | $10.00 | $1.00 | $12.50 | $50.00 |
| GPT-6 Luna | $0.10 | $0.01 | $0.125 | $0.50 |
Our calculation: 100,000 uncached input tokens and 10,000 output tokens, including reasoning, cost $0.30 on Sol 6.1. With identical token counts, Astra costs $1.50. This compares token bills; it is not a measured cost to solve a problem.
If all 100,000 input tokens qualify for a cache hit, the equivalent Sol 6.1 bill is $0.11, excluding the cost of creating the cache. This does not mean every long conversation automatically receives that rate. Context reuse needs to be considered in the application design.
The pricing documentation specifies additional conditions. Above 272,000 input tokens, multipliers apply to the full request: 2× input and cache rates, and 1.5× output. Fast is 2× Standard, Batch and Flex rates are 50% lower, and regional processing adds 10% where offered. Tool fees may also apply.
A production budget should include retries, unfinished tasks and review time. A cheaper token is less useful if an answer repeatedly needs repair; conversely, a higher effort may pay off if it reduces the number of attempts.
API integration: compatibility, reasoning and tools
The GPT-6 guide lists five reasoning.effort settings: low, medium, high, xhigh and max. The default is medium; Sol 6.1 does not support none or minimal.
Tool calling requires the Responses API. Chat Completions remains available for requests without tools. The changelog also lists Multi-agent support in beta.
Below is a minimal request body for POST /v1/responses, following the documentation. This is a configuration example, not a record of a test performed by TreffikAI:
{
"model": "gpt-6.1-sol",
"reasoning": { "effort": "medium" },
"input": "Compare the supplied reports. Identify sources for numerical claims and missing information."
}When replacing the previous Sol, check parameters, tool support and how you evaluate the result first. If the application used none, replacing the model identifier alone will not be enough. Our guide to writing prompts is a useful companion, particularly when defining the required outcome and task boundaries.
Factual errors and safety: what benchmarks cannot guarantee
The system card addendum covers hallucinations, prompt injection and respecting constraints. OpenAI treats the model's capabilities as Critical in cybersecurity and High in biology and chemistry. These are capability classifications within the vendor's framework, not a certificate of flawless behavior.
In OpenAI's difficult-prompt evaluation at Low, the share of responses with at least one factual error fell from 11.4% to 7.7%. The prompts came from conversations where users had previously reported an error and are not representative of ordinary usage. We do not turn that result into a universal hallucination rate.
A practical trial should include incomplete data, a failed tool and instructions embedded in somebody else's document. Check whether the model reports the problem and respects the task boundaries. Success on a straightforward example does not establish the robustness of an entire workflow.
For automation, retain an opportunity to review an output before sending or publishing it. The review should match the particular action: a working summary and a change to records in a live system call for different acceptance criteria.
Is GPT-6.1 Sol worth adopting?
Our assessment: GPT-6.1 Sol deserves comparison with its predecessor on tasks involving code, documents and applications. The most useful reason to try it is the combination of independently visible progress and unchanged Standard input and output prices for Sol.
We do not assume that it automatically replaces Astra. For the hardest tasks, compare both models on the same inputs and assess the result after verification. Our GPT-6 Astra release article explains the flagship model's positioning; its pricing figures belong to the date of that publication.
Before choosing a migration, answer four questions:
- Does the output satisfy the acceptance criterion you defined beforehand?
- What does a completed and accepted task cost, including retries?
- How long does the user wait for a usable result?
- How much checking does a person still need to perform?
If the change improves those measures on real work, there is a sound reason to adopt it. A version number and a single leaderboard position are weaker evidence than a verified result in your own process.
(Sources: OpenAI documentation and the linked OpenAI, Artificial Analysis and Vals AI evaluations. Cover: BoliviaInteligente / Unsplash; inline photo: Alex Knight / Unsplash; image licence. Chart: original TreffikAI visualization.)



