DocsQuick StartAI News
AI NewsGPT-6.1 Sol Is Here: One-Fifth the Price of Astra
New Model

GPT-6.1 Sol Is Here: One-Fifth the Price of Astra

2026-09-29T23:04:07.614Z
GPT-6.1 Sol Is Here: One-Fifth the Price of Astra

OpenAI has released GPT-6.1 Sol, focusing on agentic coding, computer use, and professional workflows. Delivering performance close to that of the flagship GPT-6 Astra, it costs just one-fifth as much for standard input and output.

GPT-6.1 Sol Is Here: One-Fifth the Price of Astra

OpenAI has released another GPT-6 model.

At its DevDay event on September 29, OpenAI officially launched GPT-6.1 Sol. This is an upgraded version released just one week after GPT-6 Sol, with a clear positioning: rather than pursuing the title of most capable model, it aims to bring the operating costs of agents and professional workflows down while offering capabilities close to the flagship GPT-6 Astra.

The most important numbers are the prices. GPT-6.1 Sol costs $2 per million tokens for standard input and $10 per million tokens for output; GPT-6 Astra costs $10 and $50, respectively. In other words, GPT-6.1 Sol's standard input and output prices are both one-fifth of Astra's.

This is not simply a "cheaper flagship." Based on OpenAI's published test results, GPT-6.1 Sol has moved significantly closer to Astra on tasks such as software engineering, computer use, complex document understanding, and automated workflows. For AI agents that need to run for long periods and repeatedly call tools, this change is more valuable than an improvement in one-off question-answering scores: every round of planning, retrieval, execution, and verification can directly affect whether a product can scale.

Comparison of the positioning and pricing of OpenAI GPT-6.1 Sol, GPT-6 Astra, and GPT-6 Luna

Sol's Upgrade Is About Getting Work Done, Not Chatting

GPT-6 Astra remains OpenAI's most capable frontier model overall, suited to tasks that demand extremely high result quality but involve relatively limited usage. GPT-6 Luna handles low-cost, high-throughput everyday work. GPT-6.1 Sol sits between the two: it aims to become the more common default model for developers building AI agents.

This positioning is already clear from the model's name. Sol is not intended to replace Astra. Instead, it moves some tasks that previously required a flagship model into a lower-cost range.

OpenAI highlighted four categories of use cases:

  • Agentic coding: Locating issues in real codebases, modifying code, running tests, and continuing to iterate based on the results.
  • Computer use: Understanding interfaces, clicking controls, filling out forms, and completing multistep tasks across applications.
  • Professional workflows: Processing complex PDFs, spreadsheets, charts, and business materials in fields such as finance, healthcare, and law.
  • Automated tasks: Executing business processes that require multiple rounds of tool calls and state evaluation, rather than simply generating a piece of text.

The common feature of these tasks is that the model does not finish after answering once. It must retain context, call tools, check results, and then decide what to do next. The model's per-call price is certainly important, but it matters even more whether the model can avoid detours, make fewer factual errors, and perform fewer unauthorized actions.

Scores Close to Astra, With a More Significant Cost Advantage

Based on the evaluation results disclosed by OpenAI, GPT-6.1 Sol is not attracting developers on price alone.

In the DeepSWE v1.1 software engineering evaluation, GPT-6.1 Sol reached a level comparable to GPT-6 Astra at higher reasoning intensity. At lower reasoning intensity and cost, it scored 6.4 percentage points higher than GPT-6 Sol. This result is especially important for coding agents because real-world software engineering tasks typically involve more than "writing a function": they require hours of ongoing investigation, modification, and verification in unfamiliar codebases.

In the OSWorld 2.0 computer-use evaluation, GPT-6.1 Sol scored 7 percentage points higher than GPT-6 Sol at maximum reasoning intensity, narrowing the gap with Astra to 2.1 percentage points, while costing approximately one-seventh as much per task. In other words, developers may not need to call the most expensive flagship model for every computer-use task.

In the professional document task GDP.pdf, GPT-6.1 Sol outperformed Opus 5.5 with fallback enabled at less than half the per-task cost; compared with Astra, it approached industry-leading performance at approximately one-fifth the cost.

In the AutomationBench automated workflow evaluation, GPT-6.1 Sol scored 2.2 percentage points higher than Opus 5.5 at medium reasoning intensity, at approximately one-third of the cost. Under the same settings, it also scored 4.8 percentage points higher than GPT-6 Sol.

The gap is even more pronounced on scientific terminal tasks. In Terminal-Bench Science 0.1, GPT-6.1 Sol scored more than twice as high as GPT-6 Sol, with an average per-task cost of approximately $5.47, significantly lower than Opus 5.5 at $23.21 and GPT-6 Astra at $23.80.

Of course, all of these are evaluation results published by the model provider. The test set, prompts, reasoning intensity, and cost calculation method can all affect the final conclusions. Developers should not equate benchmark scores directly with production performance. This is especially true in high-risk scenarios involving code changes, payments, or account operations, where regression testing on their own data and workflows is still necessary.

Fewer Factual Errors, but It Is Still Not "Self-Driving"

Another upgrade in GPT-6.1 Sol is reliability.

At low reasoning intensity, GPT-6.1 Sol's factual error rate fell from 11.4% for GPT-6 Sol to 7.7%, a decrease of approximately 32%. Across different reasoning settings, its error-rate gap with GPT-6 Astra remained within 1.9%.

OpenAI also stated that GPT-6.1 Sol is less likely to ignore failed search tools, better able to comply with restrictions explicitly set by users, and less likely to complete unauthorized actions. In safety testing, OpenAI did not observe it attempting to bypass automated safety reviewers, consistent with the performance of GPT-6 Astra and the earlier GPT-6 Sol.

For agents, these changes are more practical than simply making responses "more human-like." A model that clearly reports a problem when a tool fails instead of fabricating search results, stops when it lacks sufficient permissions instead of continuing to try, and requests confirmation when task boundaries are unclear can significantly reduce the maintenance cost of an automated system.

However, developers still cannot interpret these improvements as a safety guarantee. Any agent capable of operating browsers, code repositories, databases, or enterprise systems should be configured with permission isolation, human confirmation, operation auditing, and failure rollback. Improvements in model alignment cannot replace engineering safety boundaries.

How to Choose an API Pricing Mode: Cheaper Does Not Mean Every Mode Is Cheap

GPT-6.1 Sol offers four processing modes: Standard, Batch, Flex, and Fast. Developers can choose among them based on latency and budget.

| Mode | Suitable for | Short-context input | Short-context output | |---|---|---:|---:| | Standard | Everyday development and general applications | $2 / million tokens | $10 / million tokens | | Batch | Non-real-time, large-scale offline tasks | $1 / million tokens | $5 / million tokens | | Flex | Balancing cost and speed | $1 / million tokens | $5 / million tokens | | Fast | Real-time conversations and low-latency applications | $4 / million tokens | $20 / million tokens |

Once the context exceeds 272,000 tokens, Standard mode costs $4 for input and $15 for output; Batch and Flex cost $2 for input and $7.50 for output; Fast costs $8 for input and $30 for output.

Cached input is the part of Sol's pricing structure that deserves the most attention. For short contexts, cached input costs only $0.10 per million tokens, 95% less than the standard input price. For agents that repeatedly carry system prompts, project documentation, codebase indexes, or business rules, this means frequently reused context can be retained without being billed at the full input rate every time.

In practice, developers can adopt the following routing strategy:

  1. Use Standard for real-time interactions and tasks where users are waiting, avoiding the direct doubling of costs caused by Fast mode.
  2. Use Batch or Flex for overnight code scanning, batch document extraction, and offline data processing.
  3. Consider Fast only when latency directly affects conversion rates or user experience.
  4. Enable caching for fixed long prompts, project materials, and tool descriptions, and monitor the cache hit rate.
  5. Do not use Sol for every simple classification, summarization, and formatting task merely to pursue performance "close to Astra." These tasks can be assigned to a lower-cost model.

Calling GPT-6.1 Sol Through OpenAI Hub

GPT-6.1 Sol is now available through the API. Developers who want unified access to models such as GPT, Claude, and Gemini can use OpenAI Hub's OpenAI-compatible interface without separately modifying their entire business codebase for each model.

Here is a basic Python example:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_OPENAI_HUB_API_KEY",
    base_url="https://openai-hub.com/v1"
)

response = client.chat.completions.create(
    model="gpt-6.1-sol",
    messages=[
        {
            "role": "system",
            "content": "You are a rigorous software engineering assistant. Do not modify production environments without confirmation."
        },
        {
            "role": "user",
            "content": "Analyze this code for potential concurrency issues and provide recommendations for fixing them."
        }
    ]
)

print(response.choices[0].message.content)

For coding agents or browser automation systems, it is recommended to start with read-only access, record the reasoning, parameters, and results for every tool call, and then gradually grant write permissions. Do not skip permission design and human approval simply because model prices have fallen.

What OpenAI Really Wants to Sell Is a Cheaper Agent Runtime

The significance of GPT-6.1 Sol is not merely another model iteration. It represents a redivision of labor across OpenAI's GPT-6 product line.

Astra handles the capability ceiling, Luna handles large-scale low-cost tasks, and Sol targets professional workflows that require a certain level of reasoning but cannot absorb flagship pricing. In the past, many agent projects stalled at "the demo works well, but production costs are too high": the model needed multiple rounds of reasoning and tool calls, a single task could consume millions of tokens, and the final bill ended up far higher than expected. If Sol can maintain a success rate close to Astra's in real business scenarios, it could move a range of projects from the experimental stage into production.

But there is an easily overlooked prerequisite: total cost is not determined by token prices alone. Whether the model needs retries, whether tool calls are stable, whether context can be cached, and whether human intervention is required after a task fails will all change the final unit cost. A cheap model that frequently needs to rerun tasks may not be more cost-effective than an expensive model. Developers should therefore compare the "total cost of completing an acceptable task," rather than simply comparing prices per million tokens.

Based on the information currently available, GPT-6.1 Sol is best suited to three types of users: teams building coding or browser agents, enterprises that need to process professional documents in bulk, and API developers whose projects are already constrained by the cost of calling flagship models. For simple question answering, short-form text generation, and ordinary customer service, its capabilities may be underutilized, making a lower-cost model the more reasonable choice.

GPT-6.1 Sol has been available since September 29 to ChatGPT Work and Codex users on Plus, Pro, Business, Enterprise, and Edu plans. Developers can call it through the API using gpt-6.1-sol. Whether and when it will become available to ordinary ChatGPT users has not yet been announced.

The conclusion is straightforward: if Astra is "the most capable model," Sol is competing to be "the model most worth deploying." One-fifth the price does not automatically mean one-fifth the cost, but as agents begin moving from demos into production, this price difference is already large enough to make developers reconsider which tasks are worth assigning to large models.

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: