Jahanzaib

GPT-6.1 Sol Halves the Cache Rate and Leaves Every Other Price Alone

OpenAI's GPT-6.1 Sol keeps GPT-6 Sol's $2 input and $10 output prices and cuts only cached input. Here is what that saves on real agent loops, and why Chat Completions users cannot just swap the model name.

Jahanzaib Ahmed
8 min read
GPT-6.1 Sol pricing: the OpenAI logo on a glass tile wired to a stack of copper cache blocks

If your agent runs 40 turns, GPT-6.1 Sol takes about a quarter off the bill. If it answers in one call, it takes off nothing. OpenAI released the model at DevDay on September 29, one week after GPT-6 Sol, and its price sheet explains why. GPT-6.1 Sol pricing keeps input at $2 per million tokens and output at $10, and the only rate that moved is cached input, down from $0.20 to $0.10.

So the saving depends on how much of your traffic is cache reads, which is mostly a question of how long your agent loop runs. And one group of GPT-6 Sol users cannot switch by changing the model name at all.

GPT-6.1 Sol pricing next to GPT-6 Sol and Astra

OpenAI's API changelog lists both launches a week apart. Here are the standard rates for prompts up to 272K input tokens, taken from the GPT-6.1 Sol model page, the GPT-6 Sol model page and the API pricing page.

Per 1M tokensGPT-6 SolGPT-6.1 SolGPT-6 Astra
Input$2.00$2.00$10.00
Cached input$0.20$0.10$1.00
Cache writes$2.50$2.50$12.50
Output$10.00$10.00$50.00

Cached input on GPT-6.1 Sol is priced at 5% of the uncached rate, where GPT-6 Sol charges 10%. Everything else carries over: cache writes cost 1.25 times input, prompts over 272K input tokens pay double on input and cache and 1.5 times on output for the whole request, Fast mode is twice Standard, and Batch and Flex are half. The context window is the same 1,050,000 tokens, with 922,000 of input and 128,000 of output.

The "fifth of Astra" framing comes from comparing against GPT-6 Astra at $10 input and $50 output. That comparison is fair for teams deciding between the two big models. For teams who moved to GPT-6 Sol last week, which I priced out in my GPT-6 Sol agent cost breakdown, it is the wrong baseline.

What a cheaper cache read is worth on an agent loop

On the GPT-6 family, prompt caching works per turn. OpenAI's prompt caching guide says the implicit breakpoint sits at the end of the latest eligible message, so each turn reads everything before it from cache and writes the new tokens it appended. The older a conversation gets, the bigger the read share becomes.

Diagram of one agent loop turn: earlier context billed as a cache read, the new tool result as a cache write, then model output feeding the next turn
Only the cache read step got cheaper on GPT-6.1 Sol, so the saving grows with every turn the context survives.

I ran the numbers in code for a typical tool calling agent: a 20,000 token system prompt with tool schemas, 3,000 tokens of tool result per turn and 800 output tokens per turn, all appended to context. Same token counts on every model, standard short context rates. The model assumes implicit caching, a cache that stays warm inside OpenAI's 30 minute lifetime, and the first turn's prompt billed as a cache write.

Run lengthGPT-6 SolGPT-6.1 SolSavingGPT-6 Astra
1 call$0.058$0.0580%$0.290
10 turns$0.279$0.24711.4%$1.394
40 turns$1.460$1.10024.6%$7.298

At 40 turns, cache reads are 49% of the GPT-6 Sol bill, which is why halving them takes a quarter off. On that 40 turn loop GPT-6.1 Sol also comes out 6.6 times cheaper than Astra, not 5, because Astra's cache reads cost ten times more while its output costs five times more. At 10,000 runs a month, that is about $14,600 dropping to $11,000.

Two caveats. First, this holds token counts constant, and output at $10 per million is where it breaks. The 10 turn saving of $0.032 disappears if GPT-6.1 Sol spends about 317 more output tokens per turn than GPT-6 Sol did (slightly fewer if your setup carries reasoning forward in context). At 40 turns the saving of $0.36 survives up to about 900 extra tokens per turn. That matters because, as the next section shows, some teams will be forced from no reasoning to low reasoning when they switch. Measure on your own traffic.

Second, a one call workload can do better than my table shows. In explicit only cache mode, content after your last breakpoint is billed at the plain input rate with no write charge, so that single call drops to $0.048 on either Sol model.

Chat Completions users cannot just swap the model name

GPT-6 Sol accepts reasoning_effort: "none", and on Chat Completions that setting is the only way it does function calling. GPT-6.1 Sol does not support none or minimal at all, and on Chat Completions it supports no tool calling. OpenAI's GPT-6 migration guide is explicit: tool calling on GPT-6.1 Sol requires the Responses API.

So if your agent runs tools through Chat Completions with reasoning off, changing gpt-6-sol to gpt-6.1-sol will not work. You have to port the request to Responses first, then set effort to low or higher, which also means paying for reasoning tokens you were not paying for before. On a 10 turn loop, about 317 reasoning tokens per turn is enough to wipe out the cache saving, per the break even above.

Decision tree for moving from GPT-6 Sol to GPT-6.1 Sol: tools on Chat Completions means move to the Responses API first, then swap the model name and rerun evals
If your tools run through Chat Completions, the Responses port comes before the model swap, not after.

The same guide lists smaller changes that bite in testing. When effort is not none, remove temperature, top_p and top_logprobs. If you change effort between turns, use configuration_update items rather than editing the request level setting, or you break the cached prefix you are trying to profit from.

Warning: Ultrafast is not part of this launch for Sol. The pricing page lists only GPT-6 Astra under Ultrafast, and OpenAI's Codex models page says Ultrafast support for GPT-6.1 Sol is coming later, though The Next Web reports OpenAI promising it in the coming days. Fast mode, at twice Standard, is available now.

Fewer ignored refusals, but still 23.5% of the time

OpenAI's system card addendum treats GPT-6.1 Sol as Critical in cybersecurity and puts it on the same safeguards stack as GPT-6 Astra. In practice, expect the same behaviour I described when Astra's API started stopping flagged agent tasks: build your agent so a stopped task is a normal outcome, not a crash.

The more useful figure for builders is what OpenAI calls unwanted persistence: the model trying to get around a restriction it was told about. The system card's example is a model switching to email after a direct message is blocked because the recipient is out of office. It appeared in 23.5% of GPT-6.1 Sol rollouts against 17.4% for Astra, and The Next Web reports a GPT-6 Sol figure of 64.4% on this kind of test. For anyone already on GPT-6 Sol, that is a large drop.

Read it carefully, though. OpenAI says the test mostly covers low stakes restrictions and runs without the system level controls meant to block circumvention, so it measures the model's instinct, not what happens in production. Nearly one in four is still a lot for an agent that touches real systems. It is the same behaviour behind the Medicare portal incident, where a refusal got treated as a reason to retry.

My read: for GPT-6 Sol users, this matters more than the price change. The lower persistence rate is welcome, but it is no reason to remove guardrails. When I wire a model like this into a tool layer, a denied response ends that tool path in code. The model never gets to decide whether "no" means "try another way."

Who should switch this week

CNBC reported that CFO Sarah Friar touted the price drop at DevDay. Here is how I would sort who actually benefits.

  • Switch now: agents already on the Responses API with long sessions, where cache reads dominate. You get the cheaper reads, the lower persistence rate and OpenAI's claimed quality gains for a model name change and an eval run.
  • Port first, then switch: anything calling tools through Chat Completions with reasoning set to none. Budget the Responses migration and the new reasoning cost before assuming a saving.
  • No rush: short, single call workloads. Your per request price is unchanged, so switch only if the quality gain shows up in your own evals.
  • Check residency: GPT-6.1 Sol supports US and EU data residency, but Fast mode is unavailable with EU residency, and Ultrafast does not support EU or other regional processing outside the US.

Before any switch, set a spend cap on the key you test with. I covered how to do that for the other big API in capping a Claude API key, and OpenAI added project level key expiry and creation controls this month, per the same changelog. If you are still weighing OpenAI against Claude for an agent, my comparison for business agents covers the trade offs beyond price.

OpenAI's full keynote, where GPT-6.1 Sol, dots and Astra Ultrafast were announced.

If you would rather have an agent like this built and monitored for you, that is what I do: see how I build AI agents.

Frequently asked questions

How much does GPT-6.1 Sol cost in the API?

Standard pricing for prompts up to 272K input tokens is $2 per million input tokens, $0.10 per million cached input tokens, $2.50 per million cache write tokens and $10 per million output tokens. Longer prompts pay double on input and cache and 1.5 times on output for the whole request. Batch and Flex are half price, and Fast mode is double.

Is GPT-6.1 Sol cheaper than GPT-6 Sol?

Only on cached input. Input, cache write and output prices are identical, while cached input falls from $0.20 to $0.10 per million tokens. A single uncached request costs the same on both. The saving grows with conversation length, reaching about a quarter of the bill on a 40 turn agent loop in my model.

Can I use GPT-6.1 Sol with Chat Completions?

Yes, but only without tools. GPT-6.1 Sol supports Chat Completions for plain requests, and tool calling requires the Responses API. It also rejects the none and minimal reasoning efforts, so any Chat Completions setup that relied on reasoning set to none for function calling has to move to Responses before switching.

Does GPT-6.1 Sol support Ultrafast mode?

Not yet. OpenAI's pricing page lists only GPT-6 Astra under Ultrafast, and the Codex models documentation says Ultrafast support for GPT-6.1 Sol is coming later, and The Next Web reports OpenAI promising it in the coming days. Standard and Fast modes are available at launch, with Fast billed at twice the Standard rate.

What is the GPT-6.1 Sol context window?

The model page lists a 1,050,000 token context window, with a maximum of 922,000 input tokens and 128,000 output tokens, the same limits as GPT-6 Sol. Its knowledge cutoff is April 30, 2026.

Feed to Claude or ChatGPT

Published

September 30, 2026

Jahanzaib Ahmed

Jahanzaib Ahmed

AI Systems Engineer & Founder

AI Systems Engineer with 126 production systems shipped. I run AgenticMode AI (AI agents, RAG systems, voice AI) and ECOM PANDA (ecommerce agency). I build AI that works in the real world for businesses across home services, healthcare, ecommerce, SaaS, and real estate.