If you spent any time on LLM evaluation platforms over the weekend, you likely noticed something unusual happening on Arena.ai. A mysterious model listed under the temporary handle gemini-3.8-flash began churning out responses that blew standard benchmarks out of the water. AI researchers and prompt engineers quickly realized they weren’t looking at a minor iterative […]
The post Gemini 4 Pro Stealth-Tested on Arena.ai: 10-Minute Inference, 10M Context, and Unbelievable Visual Outputs appeared first on NPowerUser.
If you spent any time on LLM evaluation platforms over the weekend, you likely noticed something unusual happening on Arena.ai.
A mysterious model listed under the temporary handle gemini-3.8-flash began churning out responses that blew standard benchmarks out of the water. AI researchers and prompt engineers quickly realized they weren’t looking at a minor iterative update. Instead, as ongoing community trackers and specialized coverage like the NokiaPowerUser Gemini 4 Pro coverage hub highlight, all signs point to an early, stealth checkpoint of Google’s upcoming flagship: Gemini 4 Pro.
From hyper-complex SVG vector art that took 10 minutes of deep reasoning to assemble, to rumors of ghost-routed backend specs featuring a 10-million token context window, the early trial runs suggest Google is preparing a massive leap forward in frontier AI models.
The Visual Proof: 10-Minute Reasoning and Complex Code Art🚨 Gemini 4 Pro is crazy (PS5 SVG example)
This might genuinely be one of the craziest outputs I’ve got from an AI model yet
For context this took around 10 minutes to complete too
We are in for a treat when this thing fully drops pic.twitter.com/6YvDYok8wg
— Lumina (@LuminaBench) September 17, 2026
What initially raised eyebrows was the sheer density and visual accuracy of the model’s visual code outputs. Standard language models usually struggle with precise spatial coordinates, often returning broken or simplified vector graphics when asked to generate raw code.
The new checkpoint on LMSYS Chatbot Arena handled spatial reasoning with astonishing accuracy:
The PS5 Vector Render: One user prompted the model for a vector render of a PlayStation 5 console. The model spent roughly 10 minutes thinking, compiling, and refining before delivering thousands of lines of flawless SVG code capturing every curve, shadow, and optical drive accent.
The BMW M4 & Voxel Pagoda: Other testers shared side-by-side comparisons showing hyper-detailed 3D voxel pagodas and a sleek BMW M4 design. Compared to outputs from older Gemini checkpoints, lighting effects, shading depth, and geometric alignment were vastly superior.
Gemini 4 Pro in Arena (under the name gemini-3.8-flash)
> clean SVG of a domestic cat in side view
left – Gemini 4 Pro
right – GPT 6 Astra max in codex pic.twitter.com/diS97x7I9j— Harshith (@HarshithLucky3) September 17, 2026
Complex Scene Composition: Another standout prompt—a pelican riding a bicycle under a twilight sky full of stars—demonstrated native spatial balance without clipping or misaligned vector shapes.
+--------------------------------------------------------------------+
| CHAKRA / ARENA.AI BENCHMARK |
+--------------------------------------------------------------------+
| Previous Checkpoint (Gemini 3.8 Flash): |
| - Fast response (~10-15s) |
| - Basic geometric shapes, approximate coordinate math |
| |
| New Stealth Checkpoint (Gemini 4 Pro Candidate): |
| - Extended inference reasoning (~8 to 10 minutes of execution) |
| - Precision coordinate placement, advanced lighting & voxel math |
+--------------------------------------------------------------------+
The long generation times—ranging from 8 to 10 minutes per response—indicate that Google is heavily deploying extended test-time compute, allowing the model to internally plan, execute, debug, and polish complex code before outputting the final result.
Leaked Backend Specs: 10 Million Context Ceiling & Native Agent Controlsuccessfully ghost routed into Gemini 4 Pro’s backend and here’s what I found
>10 million input ceiling
>256k output ceiling
>permanent cross session memory
>doesn’t need API’s to use the internet
>generates fake “trap” environments when feeding it malware
>compiles and runs… https://t.co/dP86qvDekf pic.twitter.com/0xzqSHdbOK— Qwinah (@MaaSonder) September 17, 2026
While visual outputs grabbed headlines, reverse engineers attempting to analyze the routing path behind the unreleased checkpoint uncovered even more intriguing technical details.
According to posts circulating on X from developers who managed to trace ghost-routed calls into the backend, the target architecture carries specs that leave current public models behind:
10 Million Input Context Window: The model appears capable of ingesting up to 10M tokens in a single prompt, doubling the previous industry milestone set by Google’s earlier context expansions.
256K Token Output Limit: Most frontier models cap maximum output generation between 4K and 16K tokens. A 256K output ceiling allows Gemini 4 Pro to generate entire software repositories, long-form novels, or full codebase refactors in a single pass.
Cross-Session Permanent Memory: The architecture reportedly supports persistent user memory natively, maintaining context and learned preferences across separate chat threads without requiring external database integrations.
Backend Execution & Sandbox “Trap” Environments: When fed malicious code or untrusted scripts, the system automatically routes the code into isolated, synthetic sandbox environments to execute safely before returning clean results to the user.
Built-in Motor Control for Robotics: Perhaps the most surprising discovery is native support for robotic motor control protocols, suggesting Google DeepMind is building unified multimodal models designed to run directly on physical automation hardware.
For users accustomed to Google Gemini, the shift in capability is immediately noticeable. Where current flagship models excel at rapid text generation and multimodal vision analysis, Gemini 4 Pro appears engineered for autonomous agent workflows and heavy reasoning tasks.
| Feature Category | Current Generation (Gemini 3.8 / 1.5) | Gemini 4 Pro (Stealth Checkpoint) |
| Max Context Window | 1M – 2M Tokens | 10M Tokens |
| Max Output Tokens | 8K – 16K Tokens | 256K Tokens |
| Code Reasoning | Instant syntax generation | Extended multi-minute planning & self-debugging |
| Memory | Session-bound / Basic context | Native permanent cross-session memory |
| Robotics Integration | Text/Vision API bridging | Direct motor control protocol translation |
Google has not officially confirmed the stealth testing on Arena.ai, but the company historically uses public evaluation platforms to stress-test release candidates weeks before a formal announcement.
Industry watchers speculate that if these blind tests continue at this rate, Google could officially reveal Gemini 4 Pro alongside new developer tooling during an upcoming fall hardware or AI launch event, likely targeted for October 2026.
If the early outputs are any indication, developers and creators are in for a significant upgrade in what generative AI can plan, code, and execute.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | Inside the Gemini Breakout: How Google’s AI Escaped Its Sandbox and Hacked 3 Real Companies | 0 | 9.41 | 19-09-2026 |
| 2 | Google Is Cooking: Gemini 3.5 Pro Mogs Claude Fable 5 in Arena Leak | 0 | 0 | 02-07-2026 |
| 3 | Google Upgrades Gemini Managed Agents: Unpacking the New Harness, Files API, and Credentials API | 0 | 7.11 | 18-09-2026 |
| 4 | Google Launches Gemini 3.8 Live Models That Can Reason While They Talk | 0 | 10.16 | 17-09-2026 |
| 5 | Google launches Gemini 3.6 Flash and 3.5 Flash-Lite, teases Gemini 4 | 0 | 27.92 | 21-07-2026 |
| 6 | Google unveils Gemini Omni 1.1 Flash that can create 4K AI videos of up to 40 seconds | 0 | 18.33 | 27-08-2026 |
| 7 | Gemini Spark rolling out to Google AI Pro users in the US | 0 | 7.75 | 23-07-2026 |
| 8 | Gemini AI Hacked Three Companies in a Testing Breakout, Google Says | 0 | 11.1 | 19-09-2026 |
| 9 | Google’s Gemini joins rogue AI cases after real-world hacks – WSJ | 0 | 12.78 | 19-09-2026 |