AI News

Gemini 4 Argon vs GPT-6.1 Sol: Same Price, Different Bet

Abhishek Bahukhandi

Abhishek Bahukhandi

•7 min read
Taqari cover art comparing Google Gemini 4 Argon and OpenAI GPT-6.1 Sol as frontier coding models
Taqari cover art comparing Google Gemini 4 Argon and OpenAI GPT-6.1 Sol as frontier coding models
Gemini 4 Argon vs GPT-6.1 Sol: two frontier coding models at the same $2/$10 price. We compare DeepSWE v1.1 scores, cache rates and who can use them today.

Gemini 4 Argon vs GPT-6.1 Sol is a comparison neither lab planned. OpenAI shipped GPT-6.1 Sol on 29 September 2026. Google announced Gemini 4 Argon the next day. Both are aimed at long-horizon agentic coding, both lead with the same third-party benchmark, and both arrived at exactly $2 per million input tokens and $10 per million output tokens.

The identical rate card is not the story. The story is what each lab decided to buy with that budget — and that one of these models is in production right now while the other is behind an application form.

What the two labs shipped, 24 hours apart

OpenAI's framing is cost-performance. Its GPT-6.1 Sol announcement describes an upgrade to GPT-6 Sol that nearly matches GPT-6 Astra's intelligence on agentic coding, computer use and professional work at one-fifth of Astra's standard token prices. It is available to Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, and to developers in the API as gpt-6.1-sol.

Google's framing is capability. Its Gemini 4 Argon announcement positions Argon as a frontier model for real-world coding, enterprise knowledge work and cyber defence, with a one-million-token limit for multi-step problems. Google claims a new state of the art on DeepSWE v1.1 and the lead on the Vals Index, which weights finance, coding, legal and tax work by each sector's contribution to US GDP.

The rate cards are the same three numbers

Per million tokens, from each lab's own published figures:

Model Input Output Cached input
Gemini 4 Argon$2$10$0.10 (95% off input)
GPT-6.1 Sol$2$10$0.10
GPT-6 Sol (previous)$2$10$0.20
GPT-6 Astra (flagship)$10$50not published here

Two labs with different hardware, different training runs and no reason to coordinate landed on $2, $10 and $0.10. That is what a commodity price looks like while it is forming. Google labels its number introductory, so treat it as a floor that may rise rather than a committed rate.

What is different underneath

  • Context: Argon advertises a one-million-token limit. GPT-6.1 Sol's model reference lists 1,050,000 tokens, with a 128,000-token maximum output.
  • Knowledge cutoff: 30 April 2026 for GPT-6.1 Sol. Google has not published an equivalent figure for Argon.
  • Speed: OpenAI says a GPT-6.1 Sol Ultrafast variant is coming with up to 8x faster token generation in Codex than the standard speed.
  • Availability: GPT-6.1 Sol is generally available. Argon is rolling out.

Gemini 4 Argon vs GPT-6.1 Sol on DeepSWE v1.1

DeepSWE v1.1 is a third-party benchmark from Datacurve. It scores coding agents on long-horizon engineering work inside real repositories, graded by purpose-written verifiers that check observable behaviour rather than diffing against one blessed solution. Both labs chose to lead with it, which tells you what they want to be judged on: planning and tool use across many steps, not single-shot recall.

Model DeepSWE v1.1 Price (in/out)
Gemini 4 Argon77.9%$2 / $10
GPT-6.1 Sol (high effort)75.2%$2 / $10
GPT-6 Astra74%$10 / $50
GPT-6 Sol (max effort)68.8%$2 / $10

The 2.7-point gap is the least interesting number here

Argon leads GPT-6.1 Sol 77.9% to 75.2%. Both figures are self-reported, each lab ran its own harness, and neither published a confidence interval alongside the headline. On a benchmark of long-horizon tasks, where a single tool call can cascade, 2.7 points is not a margin to rebuild a stack over.

Look at the third row instead. GPT-6.1 Sol scores 75.2% at $2/$10. GPT-6 Astra, OpenAI's own flagship, scores 74% at $10/$50. The cheap model beat the expensive one on the benchmark both were measured against, and it did so eight days after Astra's sibling tier shipped.

Frontier-level long-horizon coding went from $10/$50 to $2/$10 inside a fortnight, and two independent labs set the same price. If your cost model assumes the best coding model is the expensive one, it is out of date. Budget by cached-input rate and tokens per task, not by tier name.

Read the effort-level footnote before you quote the score

OpenAI's number comes with conditions that change what it means.

Reasoning effort is a pricing decision, not a toggle

The 75.2% is at high reasoning effort. GPT-6 Sol's 68.8% was its best score at maximum effort. So the new model beats the old one by 6.4 points while thinking less — OpenAI puts the saving at roughly 76% lower cost per task. Comparing two models at their defaults is not comparing them at equal compute, and the effort dial is where most of the real spend hides.

Cost per task, not cost per token

A model that reasons for 40,000 tokens at $2 per million can cost more per closed ticket than one that reasons for 8,000 tokens at $4. Identical headline rates tell you almost nothing about the bill. The only comparison that settles it is your own tasks, your own harness, measured per completed task.

The cache rate is where the real price cut landed

GPT-6 Sol already cost $2 and $10. The headline rate did not move at all. What moved is cached input: $0.20 per million on GPT-6 Sol, $0.10 on GPT-6.1 Sol. Google arrives at the same $0.10 from the other direction, discounting cached input by 95% off its $2 input rate.

For an agent loop that is the number that matters. A coding agent re-sends the same repository map, the same system prompt and the same conversation history on every turn; after the first call, most of the input is a cache hit. Halving the cache rate halves the dominant line item on a long session while the quoted price stays flat. We hit the same pattern comparing Anthropic's tiers — the Sonnet 5.5 and Opus 5.5 cache rates are identical, which is why the apparent 50% saving there is nearer 43% in practice. If you want the arithmetic worked through on real traffic, our breakdown of what Opus 5.5 costs to run walks through it.

Argon's second bet: patching vulnerabilities on its own

This is where the two models stop being substitutes. Google says Argon can autonomously find, validate and patch critical software vulnerabilities, and that for trusted defenders and its own internal teams it will be released without cyber guardrails so the full capability is available. That variant is gated through the Fairwind Program, a vetted set of cyber defenders, and can run standalone or alongside Google's CodeMender.

Read that as a product decision about who the customer is. OpenAI priced GPT-6.1 Sol for every developer with a Plus subscription. Google built a capability it does not consider safe to hand out, and is distributing it by invitation. Same $2 and $10, two different theories of what a frontier coding model is for.

Which one can you actually use today?

GPT-6.1 Sol — now

Live in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu, and in the API as gpt-6.1-sol. It has a published model reference with context window, cutoff and pricing. You can run your own evaluation this afternoon.

Gemini 4 Argon — soon, for some

Google describes Argon as rolling out, with the cyber-defence configuration going to Fairwind members first. An announced frontier model is not the same thing as a model you can put behind a product, and Argon's price is explicitly introductory. Until it is generally available, its 77.9% is a claim about what you will be able to buy, not a number you can act on.

That asymmetry should weight your decision more than 2.7 benchmark points. A model you can call today beats a better model you can read about.

What a $2 frontier coding model changes for developers

When near-frontier long-horizon coding costs the same as the mid tier did last month, the work that is scarce shifts. Closing a well-specified ticket inside a known repository is now cheap and increasingly reliable. What stays expensive is everything around it: deciding which ticket is worth closing, noticing that the agent solved the wrong problem, and reviewing a 400-line diff you did not write.

That reshapes interviews faster than it reshapes jobs. Reciting a known algorithm proves less every quarter. Reading an agent's output critically, catching a wrong abstraction before it ships, and defending a trade-off out loud are becoming the parts that are actually tested — which is why we build Taqari's free AI mock interviews around explaining your reasoning rather than reproducing a solution. If you are planning prep around this shift, our guide to preparing with AI interview tools is the practical version.

How we would choose

  1. Need it in production this week: GPT-6.1 Sol. It is the only one of the two you can call, and it beats OpenAI's own flagship on the shared benchmark at a fifth of the price.
  2. Running long agent sessions: compare cached-input rates and tokens per task, not headline prices. Both now quote $0.10 cached, so the deciding factor is which model finishes in fewer turns on your codebase.
  3. Doing security work at scale: Argon is the only one of the two pitching autonomous vulnerability patching, and the Fairwind application is the entry point.
  4. Picking on benchmarks alone: don't. Two self-reported scores 2.7 points apart, measured on different harnesses at different effort levels, is not evidence. Run twenty of your own tasks.

The useful takeaway is not which model won. It is that $2 and $10 is now the going rate for near-frontier agentic coding, set independently by two labs in two days, and that the flagship tier above it has to justify five times the price on something other than DeepSWE.

Frequently asked questions

Is Gemini 4 Argon better than GPT-6.1 Sol for coding?

+

On DeepSWE v1.1, Argon scores 77.9% against GPT-6.1 Sol's 75.2% at high reasoning effort. That is a 2.7-point lead on one benchmark, each lab running its own harness. Treat it as a lead, not a verdict, and note that only GPT-6.1 Sol is generally available.

How much do Gemini 4 Argon and GPT-6.1 Sol cost?

+

Both list $2 per million input tokens and $10 per million output tokens, with cached input at $0.10 per million. Google calls its figure an introductory price and discounts cached input by 95%. OpenAI publishes the same three numbers on its API pricing page.

Can I use Gemini 4 Argon today?

+

Not generally. Google describes Argon as rolling out, with the cybersecurity variant gated behind its Fairwind Program for trusted cyber defenders. GPT-6.1 Sol is already live in ChatGPT Work and Codex for paid tiers, and in the API as gpt-6.1-sol.

What is DeepSWE v1.1 and why do both labs cite it?

+

It is a third-party benchmark from Datacurve that scores coding agents on long-horizon engineering tasks inside real repositories, graded by programs that check behaviour rather than matching a diff. Both labs lead with it because it rewards planning and tool use, not recall.

Does GPT-6.1 Sol really beat GPT-6 Astra at a fifth of the price?

+

On DeepSWE v1.1, yes. OpenAI reports 75.2% for GPT-6.1 Sol at $2/$10 per million tokens against 74% for GPT-6 Astra at $10/$50. One benchmark is not the whole picture, but on long-horizon coding the cheaper model is no longer the compromise.

What does a cheap frontier coding model change for interview prep?

+

It moves the bar. When a $2-per-million model closes long-horizon tickets, reciting an algorithm proves less than reading an agent's diff, catching a wrong abstraction, and defending a trade-off aloud. Expect more debugging, review and design questions.

Sources

Did you find this helpful?

Share this guide with your circle.

#gemini 4 argon vs gpt-6.1 sol#gemini 4 argon#gpt-6.1 sol#gemini 4 argon pricing#deepswe v1.1#best ai model for coding#ai coding model comparison