Info: Machine translation This post was machine-translated from the Chinese original. Wording may be rough in places — the Chinese version is authoritative.
Note This article was first published on the HyphenTech official WeChat account
Gemini enters the Agent era. What big move is Google pulling this time?
From “Answering Questions” to “Doing Things for You” · $180 Billion Bet on a Full-Stack Agent Ecosystem Managed Agents · Proactive Workflow · Gemini Spark 24/7 Backend Agent
May 19, 2026 · Google I/O · TechCrunch, Google official, CNET

In short: AI is no longer just a chatbot
In the early hours of May 19, Google I/O 2026 opened. Sundar Pichai said something on stage that I believe deserves to be recorded in AI history:
The most cutting-edge agents may have only reached 0.1% of the world’s population.
—— Sundar Pichai, Google CEO, I/O 2026 Keynote
The underlying message is: Agent technology penetration is still very low, and Google wants to push it to the masses.
Throughout the entire launch event, Google’s core narrative shifted in just one way—from ‘answering questions’ to 'doing things for you. It’s not just about adding several new model names; the platform’s strategy has shifted from ‘model capability competition’ to ‘Agent ecosystem war.’
Note 900+ million — Gemini App monthly active users (doubling within a year, now the world’s largest AI app)
The new model trio: Flash, Omni, and Spark
This time, Google released three new models/products, each with a clear focus:
Gemini 3.5 Flash — The king of speed and cost
This time, Google isn’t aiming for the “smartest” but is playing the **“cost per unit intelligence” card. Gemini 3.5 Flash inference speed is 289 tokens/second, 4 times faster than competing cutting-edge models, and internally running Antigravity even reaches 12 times.
| Indicators | Gemini 3.5 Flash | GPT-5.5 | Claude Opus 4.7 |
|---|---|---|---|
| Reasoning speed | 289 tokens/s ✅ | About 72 tokens/s | About 65 tokens/s |
| API output price | $9.00 per million tokens ✅ | $25-30 per million tokens | $25-30 per million tokens |
| Cost advantages | Benchmark ✅ | Three times more expensive | Three times more expensive |
| Terminal-Bench 2.1 | 76.2% | 83.4% | 69.2% |
※ Data source: Google I/O 2026 official demo, TechCrunch report | Speed comparison based on public API
Key Insight: While GPT-5.5 and Claude Opus 4.7 are still stronger in deep reasoning, Gemini 3.5 Flash covers 80% of daily Agent task demands at an extremely low price. Google’s strategy is clear: not beat individual champions, but win the entire ecosystem.
Note The Gemini 3.5 Pro has not been released at I/O; Pichai revealed that it is “already in use internally” and will debut in summer 2026.
Gemini Omni — a multimodal world model
Omni was the most visually striking release of this I/O. It is a native multimodal model supporting arbitrary modal input to arbitrary modal output:
-
Video Generation: Can generate video clips up to 10 seconds long
-
Video Editing: Modify video content using natural language commands (“One-sentence to Movie”)
-
1 Million Token Context Window
-
Physical simulation accuracy 77.1% (World Model physical consistency benchmark)
The real breakthrough lies in enabling models to understand the laws of the physical world through video generation.
—— Oriol Vinyals, Vice President of Google DeepMind, Co-Head of Gemini

Gemini Spark — your 24/7 backend AI butler
Spark is less popular than Antigravity, but may have the biggest impact on regular users. It’s a 24/7 backend AI Agent built into the Gemini App, running on a dedicated Google Cloud virtual machine—even if you turn off your computer, it’s still working.
| Function | Description | Status |
|---|---|---|
| Daily Brief | Automatically summarize Gmail+Calendar every morning to generate work briefs | Already online |
| Cross-application tasks | Perform multi-step tasks across Gmail, Drive, Calendar, and Docs | Already online |
| Customize Skills | Supports custom skills and integration with third-party services | Already online |
| Proactive warning | Able to proactively perform tasks and issue reminders | Already online |
| Booking Agent | Assists users in completing ticket booking, reservations, and other operations | Launched in summer |
※ Available only for Google AI Ultra subscribers ($99.99/month); currently only supports American English users
Note Spark still has hallucination issues—tests show it fabricates nonexistent links and tables. Leaked user terms also indicate that it “may share your information without permission.” The cost of trust is not low.

Core Concept: From ‘Passive Answer’ to ‘Active Execution’
The core conceptual shift at Google I/O 2026 is summed up in one word: Proactive Workflows.
In the traditional mode, you have to proactively ask the AI: “Help me summarize today’s email.” In the Agentic mode, the AI will continuously monitor→ proactively identify issues→ autonomously execute tasks→ report results. You don’t need to call it; it will keep an eye on you from the backend.

| Proactive workflow | Description |
|---|---|
| Daily Brief | Automatically summarize Gmail + Calendar every morning to generate work briefs |
| Universal Cart | Track product prices across Search/Gmail/YouTube and automatically notify you of discounts |
| Info Agents | A continuously running backend agent monitors changes in web pages and data sources |
| Gmail Live / Docs Live | Real-time AI collaboration features in Workspace |
※ These features will be rolled out gradually in summer 2026
InfoQ’s commentary on this hit the mark: ** “From answering questions to resident workflow systems—this could become a persistent workflow system.” ** Google’s ultimate vision is: Gemini is no longer just an app you open but an infrastructure that always runs in the background.
Developer nuclear weapons: Antigravity 2.0
If Spark is aimed at ordinary users, then Antigravity 2.0 is a nuclear weapon aimed at developers.
DeepMind CTO Koray Kavukcuoglu clearly stated: “Not an IDE, but an agent-first dev platform.” It’s not just an editor or CLI, but a four-piece set:
| Components | Positioning | By analogy |
|---|---|---|
| Desktop applications | Visualized Agent management interface for multitasking in parallel | Claude.ai + Task Scheduler |
| Antigravity CLI | Terminal version, Go, replacing Gemini CLI | Claude Code, but multiple Agents run in parallel |
| Python SDK | Customize the Agent workflow | LangChain, but Google’s own |
| VS Code plugin | Agent view + code completion within the IDE | Cursor, but in a shell Gemini |
※ Gemini 3.5 Flash runs at the base level, free to use throughout summer
Note 93 Agents collaborating simultaneously (a complete OS kernel built within 12 hours, costing only about $1,000)
This demo was the most explosive of the event: 93 agents built a complete OS kernel in 12 hours, costing only about $1,000. Previously, projects of this scale required a team to work for several months.
Note The internal speed of Antigravity running Gemini 3.5 Flash is 12 times faster than the public API. The public API is already “4x faster than similar frontier models”—what does 12x mean internally…
Managed Agents: The infrastructure of the Agent ecosystem
If Antigravity is a developer tool, then Managed Agents is the underlying infrastructure of the entire Agent ecosystem. It is provided through the Gemini API and serves as the foundation of Google’s Agent strategy.
| Characteristics | Google Managed Agents | OpenAI Agents API | Anthropic Tool Use |
|---|---|---|---|
| Operating environment | Isolating the Linux sandbox ✅ | Cloud function | Client execution |
| Arrangement layer | Antigravity 2.0 ✅ | and built it himself | and built it himself |
| Status persistence | ✅ Support | Limited | No support |
| Corporate governance | ✅ Built-in | Limited | Limited |
※ Source: Google I/O 2026 official documentation, EnterpriseDNA analysis
Each Agent runs in an independent sandboxed Linux container, supports cross-session state maintenance, multi-tool calls (code execution, API calls, file operations), and has a built-in enterprise-grade secure sandbox. Developers can deploy executable data analytics agents with a single call via the Gemini API, without needing to build their own infrastructure.

Three Kingdoms Kill: Google vs OpenAI vs Anthropic
The AI competition landscape in 2026 can be clearly seen in a single table showing the differentiated paths of the three companies:
| Dimension | OpenAI | Anthropic | |
|---|---|---|---|
| Strategic positioning | Ecosystem Core (Infrastructure + Platform) | Super App (ChatGPT-Centered) | Developer + enterprise focus |
| Agent mode | Full stack: Model + orchestration + sandbox + distribution | Model + API | Tool Use+API |
| Distribution channels | Search: 2.5 billion MAU + Android + Workspace | ChatGPT + API | Pure API distribution |
| Price advantage | $9 per million output ✅ | $25-30 per million output | $25-30 per million output |
| In-depth reasoning | 3.5 Pro (Summer Release) | GPT-5.5 leads the way | Opus 4.8 leads |
| Killer products | Spark + Antigravity | Codex + ChatGPT Ads | Claude Code + Dynamic Workflows |
※ Compiled based on the latest information as of June 2026

Everyone is a product manager’s analysis is incisive: Google’s moat has never been models, but search distribution and AI Overviews. **Google’s differentiating advantage lies in the trinity of channel, device, and enterprise entry point: Search AI Overview covers 2.5 billion MAU, Android devices + AR glasses (mass-produced with Samsung), and Workspace enterprise entry.
Google I/O 2026 Was Not Just a Model Launch. It was a platform shift. (This was not just a model release, but a platform-level shift. )
$180 billion: All in Agent ecosystem
Google’s capital expenditure budget for 2026 is $180–$190 billion, about six times that of 2022. This money covers the entire chain from chips (TPU) to models (Gemini) to platforms (Antigravity) to applications (Spark).
Note $180 billion — 2026 capital expenditure budget (about six times that of 2022, with a “significant increase” expected in 2027)
| Indicators | Data | Meaning |
|---|---|---|
| Gemini App MAU | 900 million+ | Doubling in one year, the world’s largest AI application |
| Search AI Overview MAU | 2.5 billion | Covering 30%+ of the global population |
| AI Mode users | 1 billion+ yuan | Break through within a year |
| Monthly processing of tokens | 3.2 trillion | A 300-fold increase |
| Monthly active developer | 8.5 million | Developer ecosystem scale |
| Cloud backlog of orders | $460 billion+ | Q1 nearly doubled |
※ Data source: Google I/O 2026 official data, Alphabet investor presentation
Google Cloud’s backlog nearly doubled to over $460 billion in Q1 2026, with monthly token processing growing 300-fold. These numbers illustrate: Enterprise demand for AI Agents is exploding.
Vinyals’ AGI roadmap: Agents will be able to build their own systems
Oriol Vinyals (Vice President of Google DeepMind, Co-Head of Gemini) systematically explained three main evolutionary threads of AGI in an in-depth interview in June:
| Main storyline | Core content |
|---|---|
| Multimodal learning | Mining knowledge from images, video, and audio |
| World Model | Simulating the operating laws of the physical world |
| Agent planning and memory | Taking initiative + continuously learning from experience |
※ Source: Sohu interview, Google DeepMind official blog
In the future, Agents will be able to build systems autonomously, dynamically generating toolchains and sub-agents. AGI requires the ability to continuously learn from real experience.
—— Oriol Vinyals, Vice President of Google DeepMind
This means Google is betting that Agents are not just about “performing predefined tasks,” but about autonomously designing the entire system’s execution framework. This is consistent with Antigravity 2.0’s “93 Agents Build OS Kernel” demo.
Risks and challenges: The Agent era is not so easy to enter
Of course, Google’s Agent vision is promising, but there are still many pitfalls in implementation:
-
Hallucination Issue: Gemini Spark has been tested to create links and tables, which raises trust costs
-
Price Threshold: The $99.99/month Ultra is not cheap, making ordinary users hesitant
-
Privacy Dispute: The leaked terms suggest Spark may share user information without permission
-
Deep Reasoning Gap: Still lags behind GPT-5.5 and Claude Opus 4.8 in tasks requiring precise reasoning
-
Debugging Complexity: When multiple agents run parallel and errors occur, it’s painful to identify which sub-agent to mislocate
Note Anthropic simultaneously issued a “brake pedal” warning—AI models may soon be able to improve themselves without supervision. Security governance in the Agent era is a challenge the entire industry must face.
My judgment: Google has a fight to win
Throughout the entire I/O, Google’s strategy was clear: **Not pursuing the smartest, but pursuing the cheapest and fastest; Not making super apps, but building the ecosystem core; From answering questions to doing things for you. **
Google holds three irreplaceable distribution channels—Search (2.5 billion MAU) + Android (3 billion devices) + Workspace (350 million paying users). This is a moat that no independent AI company can replicate in the short term.
-
If you’re a heavy Google Workspace user: The Spark + Antigravity combination is more useful than any external agent
-
If you are an Agent system developer: The Managed Agents API is worth serious study, and the multi-agent orchestration design is quite well-developed
-
If you mainly do deep coding tasks: For now, Claude Code and GPT-5.5 are still stronger. Wait until Gemini 3.5 Pro is released in the summer before re-evaluation
-
If you care about AI safety: Anthropic’s “brake pedal” warning deserves serious attention
Note In short
Google I/O 2026 marks the shift of AI competition from a “model capability race” to an “Agent ecosystem battle.” With an $180 billion bet, a full-stack Agent architecture, and a distribution advantage of 2.5 billion search MAU, Google has announced the arrival of a new era. Agent technology currently only reaches 0.1% of people—but soon, it will change the way everyone works.
Data sources for this article: Google I/O 2026 official keynote, TechCrunch, CNET, EnterpriseDNA, NetEase Technology hands-on test, Sohu interview with Oriol Vinyals, published May 19-20, 2026.
🧰 Tools I build
I maintain all of these tools myself. Preview builds are clearly labeled; the release pages are the source of truth for downloads, updates and known limits.
Info: HyphenBox Status: Official releases
A radar for free LLM APIs: availability is re-tested continuously, one local interface for all of them, and keys stay on your machine
Info: LocalBrain Status: Official releases
A multimodal MCP toolbox for local models: TTS, Whisper and video generation in one place
Info: ScreenLex Status: Official releases
Learn new words while you watch shows. Free, for Mac and Windows
Info: HyphenScreen Status: Official releases
Screen recording and smart editing in one: a DaVinci-style timeline, automatic redaction and a check of the finished video before export. Free
Quote: HyphenTech Make AI your superpower Local deployment · Free resources · Self-made software https://hyphentech.top
Late nights and burned API credits went in,a cup of tea comes back out — only if you feel like it.
Scan with WeChatPress and hold to save the image, then open it from your album in WeChat Scan

Comments
Loading comments…