Gemini 4 Argon Has Arrived, and Carbon Is Targeting Code Agents

Google has rolled out Gemini 4 Argon in phases and is internally testing new checkpoints under the codename Carbon. Argon focuses on million-token long contexts, complex coding projects, and agent tasks, while Carbon is expected to catch up with the Claude Opus series.
Google Is Betting Gemini 4’s Future on “Long Tasks”
According to the latest reports, Google is rolling out its next-generation frontier model, Gemini 4 Argon, in phases. At the same time, Google has begun testing a follow-up version codenamed Carbon internally and connected it to its internal code platform, Jetski.
This means Argon has not yet been broadly released, while the next-generation checkpoint has already entered internal iteration. For developers, what truly deserves attention is not an ordinary model upgrade, but the fact that Google is shifting Gemini’s competitive focus from “giving better answers in a single turn” to “whether it can continuously complete complex tasks over the course of several hours.”
Argon targets long-running software engineering, financial and legal research, cybersecurity defense, and multi-step automation workflows. According to Google’s disclosed test results, it supports an output limit of up to 1 million tokens, allowing it to process longer codebases, research materials, and execution traces within a single task.
The practical value of this capability is not simply a larger context window. In the past, developers who wanted an Agent to handle a large codebase usually had to design their own mechanisms for task partitioning, summary compression, state storage, and context assembly. Every time the model completed part of the work, the result had to be compressed and passed to the next round; otherwise, it would quickly hit the context limit.
A 1-million-token output limit means the model can continue analyzing, modifying, testing, and reviewing work within a longer single trajectory. The orchestration layer remains important, but some of the “memory transfer” work previously handled by developers is beginning to shift to the model itself.

Carbon: Not a New Name, but a Capability Calibration
At present, Carbon’s final product form has not been determined. Internal Google documents show that Argon, Barium, and Carbon have all been used for different test versions within the Gemini 4 family. However, Google’s internal codenames do not correspond one-to-one with public names. Carbon may ultimately become a follow-up checkpoint to Argon, or it may be released publicly as an independent model within the Gemini 4 family.
A checkpoint can be understood as a stage-specific snapshot based on the same model foundation. The model architecture may not undergo fundamental changes, but the training data, post-training methods, tool-calling strategies, and safety policies may all be adjusted. For users, updates of this kind can sometimes be more meaningful than changing the name of a major version: they may not change the model’s brand, but they can directly affect the stability of code generation, the execution success rate of Agents, and the error rate on long-running tasks.
A Google employee told the media that Carbon’s coding performance was “comparable to Opus 5.5.” However, this statement is still internal feedback. There are no public, reproducible third-party tests to support it, so it cannot be taken to mean that Carbon has comprehensively surpassed Anthropic’s models in the same tier.
This is also where Argon currently deserves the most cautious assessment. Google’s publicly presented results are impressive, but many of them come from internal cases and self-built benchmarks. Whether the model can work reliably in real development environments will depend on how it handles messy code, outdated dependencies, ambiguous requirements, failed tests, and permission restrictions.
Early internal feedback on Argon was also not entirely consistent. Some employees believed it was already very strong at complex programming and long-running tasks, while others pointed out that early versions still struggled with certain programming tasks, with an overall level roughly comparable to Anthropic’s earlier Claude Opus 5. Carbon’s significance may lie in closing this gap.
Coding Ability Will Determine Whether Argon Can Establish Itself
The internal cases disclosed by Google are almost all centered on whether the model can deliver results, rather than simply comparing its accuracy on individual questions.
In quantum computing research, Argon helped researchers optimize the space-time resources of quantum algorithms, meaning the combined cost of the number of qubits and quantum logic gates. Google said that in one real-world task, Argon improved publicly available baseline results by approximately 40% within a few minutes. The difficulty of such tasks lies not in generating code that merely looks plausible, but in understanding algorithmic constraints, proposing modifications, and then verifying through experiments whether those modifications actually work.
In a data center setting, the Argon Agent was used to analyze performance telemetry covering an entire server cluster. Google said the Agent autonomously identified and implemented a memory optimization方案, freeing more than 300 TiB of memory after deployment, with expected final savings potentially reaching 500 TiB to 1 PiB.
Cases like this clearly demonstrate the difference between Agent models and chat models. An ordinary model answers “how should this be optimized?” An Agent must continue through data retrieval, bottleneck identification, solution design, code modification, test verification, and deployment evaluation. If any step fails, the final result may be invalid.
Argon also participated in Google’s internal migration of C/C++ code to Rust, covering everything from core libraries with tens of thousands of lines of code to more than 800,000 lines of Fuchsia Zircon kernel code. For system-level code, this is not a matter of translating syntax line by line into Rust. It also requires addressing memory safety, concurrency semantics, performance regressions, FFI boundaries, and build-system compatibility.
Another case cited by Google involved the open-source video decoder libgav1. Through multiple rounds of performance analysis and experimentation, the Argon Agent rewrote approximately 32,000 lines of SIMD code, enabling the compiler to perform vectorization automatically. While maintaining completely identical video output, the migrated Rust version ran 2.7 times faster than the original Rust port and moved further toward the performance of the optimized C++ version.
The most important metric here is not “how much code the model wrote,” but whether it can establish a verifiable engineering loop:
- First understand the existing implementation and performance bottlenecks;
- Then propose multiple candidate changes;
- Use benchmark tests to confirm that the optimization is real;
- Validate the consistency of the output;
- Finally deploy only after automated testing and human review.
If Argon or Carbon can reliably complete this process across more real-world projects, Google will have genuinely entered the first tier of code Agents. Otherwise, 1 million tokens is merely a longer output window, not evidence of higher delivery capability.
Agents and Cybersecurity Are Google’s Two Strong Cards
In external benchmarks, Argon has been used for financial research, legal research, enterprise automation, long-video understanding, and cybersecurity tasks. Supplementary materials indicate that it achieved high rankings in the Vals Index, Vals Finance Agent v2, Harvey legal Agent tests, and Zapier’s AutomationBench, with an AutomationBench score of 51.3%.
However, these results still require a distinction between “leading model capability” and “leading product experience.” Enterprise users actually care about task success rates, recovery after failures, invocation costs, permission controls, log auditing, and accountability for results, rather than simply taking first place on a leaderboard.
Argon’s launch pricing has been disclosed as $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached input tokens. This pricing is highly significant for long-context Agents. Long-running tasks consume large amounts of input and output tokens. Without a caching mechanism, the cost of an individual task can easily spiral out of control. If the model can reuse stable codebases, system prompts, and historical context, the cache discount could substantially change the overall economics.
Cybersecurity is another major focus. Google says Argon can autonomously discover, validate, and fix software vulnerabilities, and has demonstrated strong capabilities in attack-surface identification, vulnerability discovery, and PoC generation in internal codebase tests and black-box penetration tests. In CWE-bench v1, Argon achieved a vulnerability-repair score of 68%, tied for the highest score.
However, the stronger a cybersecurity model becomes, the less security governance can be handled with a simple statement that “the model will refuse dangerous requests.” Through the Fairwind program, Google plans to first provide higher-permission versions to trusted cybersecurity defense experts and then expand access in phases. For ordinary developers and enterprise users, the capabilities, tool permissions, and network access ultimately available to them may differ substantially from those of the internal testing versions.
This exposes a new reality of frontier-model releases: security boundaries are no longer limited to the model’s internal refusal policies. They also include who can call the model, which tools it can invoke, whether it can access real systems, and whether its operations can be audited. For Agents, permission design is just as important as model capability.
What Developers Should Focus on Now
Argon’s release schedule has not yet fully stabilized, and Carbon has no officially confirmed public name or launch date. Developers should not immediately migrate production systems to the new model based solely on internal screenshots or employee evaluations. A more practical approach is to prepare an evaluation set focused on real-world tasks and pay particular attention to the following metrics:
- Long-task completion rate rather than single-turn answer quality;
- Test pass rates and regression rates after code modifications;
- Whether the Agent can identify the cause of a failure and recover;
- Whether tool calls comply with permission boundaries;
- Latency and actual cost under a 1-million-token context;
- Whether generated patches can pass human review and automated pipelines.
If you are already using an OpenAI-compatible API, you can first encapsulate model calls behind a unified Provider layer to avoid binding business logic to a single vendor. OpenAI Hub supports a unified OpenAI-format API, and once the model is officially available, you can switch models based on the model identifier actually provided by the platform. Example code:
from openai import OpenAI
client = OpenAI(
api_key="YOUR_OPENAI_HUB_API_KEY",
base_url="https://api.openai-hub.com/v1"
)
response = client.chat.completions.create(
model="gemini-4-argon",
messages=[
{
"role": "system",
"content": "You are a senior software engineer. Before modifying code, first analyze its dependencies, risks, and testing strategy."
},
{
"role": "user",
"content": "Please analyze the concurrency issues in this codebase, propose a remediation plan, and provide verifiable testing steps."
}
],
temperature=0.2
)
print(response.choices[0].message.content)
Note that the model identifier in the example is subject to OpenAI Hub’s actual availability list. Before integrating, confirm whether Argon has been released, as well as its context limit, pricing, rate limits, and tool-calling support. For code Agents, simply replacing the model name is usually not enough. You also need to reevaluate the tool protocol, sandbox permissions, timeout policies, and human approval checkpoints.
Assessment: Carbon Is More Worth Tracking Than Argon’s Release
Argon’s public release demonstrates that Google has shifted Gemini’s competitive focus toward complex workflows and engineering productivity. Its 1-million-token output limit, code migration cases, and cybersecurity capabilities are indeed closer to an Agent that can “participate in work” than traditional question-and-answer models.
However, whether Argon can change developers’ choices will ultimately depend on three questions: First, can the public version reproduce Google’s internal cases? Second, is its stability on long-running tasks sufficient for production environments? Third, do its pricing and permission restrictions make it affordable and accessible for small and medium-sized teams?
From this perspective, Carbon is actually the more important signal. It shows that Google does not view Argon as a flagship product whose story ends at launch, but is instead iterating intensively on coding and Agent execution capabilities. If Carbon can truly close Argon’s gap with the Claude Opus series on complex programming tasks, Gemini 4 may move beyond being “a model with very strong long-context capabilities” and become “a platform worth entrusting with an entire engineering task.”
As of October 10, 2026, most of Argon’s leading results still come from Google’s disclosed internal tests and phased access. Developers can follow its progress, but should not blindly chase rankings. For production systems, the real release day is not the day the model appears on the official website. It is the day the model can reliably complete your tasks and knows how to hand control back when it fails.
Sources
- ITHome: Google Gemini 4 “Argon” Model About to Be Released; Reports Say “Carbon” New Version Is Being Tested Internally — Coverage of Argon’s release progress, Carbon’s internal testing, and feedback from Google employees.
- ITHome — Reference source for technology and AI product developments.



