With major AI giants rolling out flagship models almost every one or two months, Google’s flagship AI model has been “missing” for three months. Gemini 3.5 Pro, originally set to debut in June, is still listed as “coming soon” on the official website, while Google has released four lightweight Flash models in succession over the past 106 days, with no word from the flagship Pro series. However, since early September, a string of leaked internal screenshots, anonymous platform tests, and reverse engineering of backend routing has gradually brought Google’s next-generation flagship model, Gemini 4 Pro, into focus.

Is Google Gemini 4 Pro About to Launch? Users Spot It Testing Incognito on Arena.ai
X user @lyra leaked: Google has officially deployed the first internal checkpoint of Gemini 4 Pro, expected to be publicly released in October. They also revealed a key decision: Google has quietly canceled Gemini 3.5 Pro and is putting all compute resources directly into Gemini 4.
A new Gemini 4 Pro checkpoint is being tested in https://t.co/6LtRlBPtTU https://t.co/oJ3rcgCjpt pic.twitter.com/Pjk6uCecqs
— lyra (@lyraxana) September 17, 2026
Tech researcher @Lentils80 leaked the first internal screenshot of Gemini 4 Pro on X. The model’s internal codename is “argonIn “High thinking effort” mode, it took 2.4 minutes to generate a complex visual scene. The screenshot shows a key specification: the output limit has been substantially increased from the previous 64K to 256K tokens。
Gemini 4 Pro checkpoints have finally started appearing internally a few days ago.
This is the first ever output from the model, internally codenamed “argon”. It took 2.4 minutes on High thinking effort.
It has a 256k token “output limit”, compared to 64k in previous Gemini… pic.twitter.com/x76ihjstTG
— Lentils (@Lentils80) September 14, 2026
September 17, one named gemini-3.8-flash A mysterious model appeared on LMSYS’s Arena.ai evaluation platform. Users quickly realized this was not an ordinary Flash update; the model took 8 to 10 minutesConducted deep reasoning and produced SVG vector graphics far above the normal standard. Generated an accurate vector rendering of a PlayStation 5 controller, a light-and-shadow design of a BMW M4, and a complex scene composition of a pelican riding a bicycle.
Gemini 4 Pro in https://t.co/ScMwTKhW8r it’s being tested under the name (gemini-3.8-flash)..
This output of Voxal Pagoda is so much better than the last checkpoint’s output a few days ago!
It worked for 8 minutes.. seems like google cooked this time! pic.twitter.com/NHVtpAH8hK
— Bee (@thtbee_) September 17, 2026
X user @LuminaBench shared the generated result of a PS5 SVG, saying, “This may be one of the craziest outputs I’ve ever gotten from an AI model.” @HarshithLucky3 ran a comparison test: when both were asked to generate a side-view SVG of a cat, Gemini 4 Pro’s output was clearly superior to GPT-6 Astra Max’s result in spatial accuracy and lighting and shadow effects.
🚨 Gemini 4 Pro is crazy (PS5 SVG example)
This might genuinely be one of the craziest outputs I’ve got from an AI model yet
For context this took around 10 minutes to complete too
We are in for a treat when this thing fully drops pic.twitter.com/6YvDYok8wg
— Lumina (@LuminaBench) September 17, 2026
Backend Reverse Engineering: Astonishing Architecture Specifications
While the community discusses visual output quality, another group of developers is trying to approach it from a lower level. X user @MaaSonder Claims to have successfully accessed Gemini 4 Pro’s backend route via a ghost route and obtained the following architecture specifications:
- 10 million token input context: It is 5 times the maximum context of previous Google models.
- 256K token output limitThe output limit of mainstream models is usually 4K-16K; 256K allows generating a complete software project or long technical document in a single response.
- Permanent memory across sessionsNative support for persistent user memory, with no external database integration required.
- Automatic sandbox isolationWhen malicious code is encountered, automatically route it to an isolated synthetic sandbox environment for execution.
- Native Robot Control ProtocolSupports motor control protocols, suggesting Google DeepMind is building a unified multimodal model that can run directly on physical automation hardware.

However, it must be emphasized that these specifications come from reverse engineering ghost route, and their reliability cannot be verified. It remains questionable whether a 10M context is technically feasible and whether it has entered actual testing.
Pichai personally confirmed: “We are training Gemini 4.”
Although Google never officially responded to the above leaks, CEO Sundar Pichai publicly confirmed during Alphabet’s Q2 2026 earnings call at the end of July that Gemini 4 is in training. Pichai said: “For the next generation of frontier models, you need a much larger foundation model. We’re training Gemini 4, and we’re very ambitious. I’m very excited about the progress inside Gemini 4, and I’m confident everyone will be satisfied when they see it released.”

He also acknowledged that Google lags behind in certain areas: “Coding and agentic coding are areas we need to improve, and the team is very, very focused on this.”
Sergey Brin Takes Personal Command: Code Strike Team
Google co-founder Sergey Brin has also personally stepped in to work on improving model capabilities. In April of this year, Brin assembled a “code strike team,” co-leading it with DeepMind CTO Koray Kavukcuoglu.
Brin told employees in an internal memo that improving coding capabilities is a step toward self-improving AI, and urged DeepMind to transform the model into the “primary developer” of code. This aligns with Google’s recent public stance on RSI (Recursive Self-Improvement). DeepMind researcher Shunyu Yao wrote when Gemini 3.8 Flash was released: “This is a small step for the model, a giant leap for RSI.”
In its release announcement for Gemini 3.8 Flash, Google also explicitly used the phrase “recursive evaluation and optimization” for the first time, stating that the model, through a long-running AI Agent loop, “recursively evaluates and optimizes the underlying model.”
Why was 3.5 Pro sacrificed?
The fate of Gemini 3.5 Pro is one of the most intriguing parts of the whole story. According to multiple reports, Google originally planned to release 3.5 Pro in June, but internal testing showed it failed to deliver a sufficient generational leap in multi-file code refactoring and autonomous Agent execution.
Google chose a bold strategy: instead of releasing a “just okay” intermediate product, it focused all its compute resources on pretraining Gemini 4.Wall Street Journal The report notes that because the Flash model is smaller, modifying it requires fewer computing resources, allowing multiple research teams to test different methods at the same time; each adjustment to the Pro model, by contrast, requires substantial resources.
Estimated Timeline and Observations
Combining information from all sources:
- SeptemberGemini 4 Flash-Lite and the updated NB2Lite (Nano Banana 2 Lite image model) are expected to be released, but news of Nano Banana 2.5 has already emerged today.
- OctoberThe industry widely expects Gemini 4 Pro to make its official debut.
- November-December9to5Google’s estimated timeline based on historical release cycles
Google is making a high-stakes bet: using rapid iteration of the Flash model to shore up commercial revenue and its developer ecosystem, while putting all its flagship-level resources behind Gemini 4 in the hope of reclaiming the lead in frontier models in one fell swoop. Leaked information suggests this card’s specifications are indeed impressive, but until Google officially unveils it, everything remains up in the air.
Source: KOCPC Chinese