A test checkpoint of Google's next flagship AI model surfaced last week on the public benchmarking site LMArena disguised under the name of an already-released model, giving outside developers an early, unofficial look at a system the company has not confirmed exists. Developer Pankaj Kumar was among the first to flag the checkpoint, which appeared on September 26 labeled "Gemini 3.8 Flash" but behaved unlike Google's actual Flash tier, according to reporting gathered by The Outpost and the AI-tracking site TestingCatalog.
Benchmarks outside the official roadmap
TestingCatalog reported the internal codename for the checkpoint as "Argon," attributing the detail to a source it calls Lentils; Google has not confirmed the name or the model's existence. Testers who ran it reported an input limit of 10 million tokens and an output limit of 256,000 tokens, up sharply from the 64,000-token output ceiling on Google's current Flash model, along with persistent memory across sessions. Leaked benchmark results circulating among developers put the checkpoint at 88% on DeepSWE v1.1, a coding-agent benchmark, 95.3% on Terminal-Bench 2.1, and 86.8% on OSWorld-2.0, a computer-use proficiency test — scores developers said outperformed OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 on coding, reasoning and agentic tasks.
Developers who tested the checkpoint's coding and design output described it as a notable jump from prior Gemini releases. One tester, posting under the handle Bee, called the design output from a 14-minute generation run "so good," adding that the inconsistent visual polish that dogged earlier Gemini models appeared resolved. Response times for the model's "high thinking" mode reportedly ranged from roughly 2.4 to 20 minutes depending on task complexity.
Google confirms a model, not a leak
Google has not acknowledged the Argon checkpoint specifically, but the company has publicly confirmed it is training a large successor model. Google DeepMind's Koray Kavukcuoglu said the company's intention is to release an early post-training checkpoint "as soon as possible — because we see the results and we are excited," and has previously indicated Gemini 4 should arrive well before year's end. Chief executive Sundar Pichai told investors on the company's second-quarter earnings call that Google was training Gemini 4 and "being very ambitious with it," and has signaled a push toward a near-monthly release cadence for future Gemini versions.
The leak lands amid unusually tight competitive pressure: Google has not shipped a new flagship Gemini since Gemini 3 in November, while OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 both launched in September. Pseudonym testing on arenas like LMArena — deploying an unreleased model under a decoy name to gather blind user feedback before a public launch — has become a standard part of how frontier labs validate a model ahead of release, making leaks like this one a predictable, if unofficial, preview of what is coming. Based on the current pace of testing, developers tracking the checkpoint expect a public release as early as October, though Google has set no confirmed date.