Google DeepMind has unveiled Gemini 4 Argon, a new AI model designed to handle complex, long-running workflows across software engineering, enterprise knowledge work, finance, legal research, and cybersecurity.
Unlike AI models built primarily around short question-and-answer interactions, Argon is designed to keep reasoning through complicated tasks for much longer. Google says the model can generate up to 1 million output tokens, giving it substantially more room to work through large, multi-step problems in a single run.
The company is already using Gemini 4 Argon internally for software engineering and research tasks, while a limited group of trusted cybersecurity defenders is getting early access.
Gemini 4 Argon can generate up to 1 million tokens
The biggest technical headline is Argon’s 1-million-token output limit, up from the 64,000-token limit of Google’s previous models.
That additional output capacity is intended for long-horizon tasks where an AI agent needs to inspect information, reason about a problem, write code or other content, test its work, and continue iterating.
Google says Gemini 4 Argon achieves a 77.9% score on DeepSWE v1.1, a benchmark for real-world, long-horizon software engineering tasks. The company also reports leading results on evaluations covering finance, legal research and drafting, and end-to-end business automation.
On Zapier’s AutomationBench, Google reports a 51.3% score, while Argon scored 91.7% on LVBench, an evaluation focused on long-video understanding.
These are Google’s reported results, and benchmark scores can vary depending on evaluation methodology, model configuration, and competing systems.
Google is already using Argon for complex engineering work
Google says its own engineers are using Gemini 4 Argon for specialized coding and research projects.
In one example, Argon helped optimize a quantum computing algorithm and beat a published baseline by 40% in a matter of minutes.
Google also used Argon agents to analyze profiling data across its data centers and identify memory optimizations. The company says the work ultimately freed more than 300 TiB of memory after deployment, with potential savings estimated at between 500 TiB and 1 PiB.
Argon is also being used for large-scale code migrations, including Google’s efforts to migrate C and C++ codebases to Rust.
One example involves Google’s libgav1 video decoder. Google says Argon replaced 32,000 lines of SIMD code in an existing Rust port through repeated profile-guided experiments and compiler analysis. The resulting memory-safe Rust implementation reportedly runs 2.7 times faster than the previous Rust port while producing identical video output.
That is the sort of workflow Google appears to have in mind for Argon: not simply generating code, but repeatedly analyzing, testing, optimizing, and refining a project.
Gemini 4 Argon targets cybersecurity, too
Cybersecurity is another major focus of the new model.
Google says Gemini 4 Argon can autonomously find, validate, and patch critical software vulnerabilities. The company is initially providing access to a limited group of trusted cybersecurity defenders through its Fairwind Program.
Google also says Wiz is using Argon through its Scan for Good initiative to identify and remediate high-risk exposures affecting critical public infrastructure.
According to Google, Argon uncovered a critical vulnerability involving sensitive personal information in healthcare software used by hospitals worldwide. The company says previous frontier models had missed the vulnerability.
On CWE-bench v1, an evaluation of vulnerability remediation, Google reports that Argon tied for first place with a score of 68%.
Because of the potential risks involved in giving a highly capable AI model more autonomous cybersecurity capabilities, Google isn’t releasing Argon broadly yet.
Gemini 4 Argon is rolling out gradually
Gemini 4 Argon is initially being made available to trusted cyber defenders and other testers while Google continues working on its safety systems.
The company says it is participating in the U.S. government’s voluntary pre-release model access process and using feedback from early testers to improve the model’s safeguards.
Google is also working on protections against cybersecurity and CBRN misuse, indirect prompt injection attacks, and situations where an AI agent could behave outside a user’s intended goals.
The company says Argon is its most resilient model yet against indirect prompt injection attacks. Google is also deploying systems that monitor the model’s reasoning and actions and can stop execution when necessary.
Google says broader availability will begin with paid API customers and Google AI Ultra subscribers, followed by wider access for developers, enterprises, and consumers.
Gemini 4 Argon pricing
Google says Gemini 4 Argon will launch with introductory API pricing of $2 per million input tokens and $10 per million output tokens.
Cached input tokens will receive a 95% discount from the input-token price.
After the introductory period, pricing will increase to $4 per million input tokens and $20 per million output tokens.
The pricing reflects the model’s intended use case. Argon isn’t being positioned as a lightweight model for casual chatbot conversations. Its target is complex work where an AI agent may spend substantial amounts of time researching, coding, analyzing information, and executing multiple steps.
Gemini 4 Argon is Google’s bet on longer-running AI agents
With Gemini 4 Argon, Google is pushing Gemini beyond the traditional prompt-and-response model toward AI systems that can work through much larger and more complicated tasks.
The 1-million-token output limit is certainly the headline specification, but the more interesting part may be what Google is doing with the model internally. From Rust migrations and algorithm optimization to data-center analysis and cybersecurity research, the company’s examples are centered on workflows that require sustained reasoning rather than a single generated answer.
The real test will come when developers and enterprises can put Gemini 4 Argon against their own workloads.
For now, Google is presenting Argon as a model built not just to answer questions, but to stick with difficult work for much longer
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
