Google's internal AI security agent has validated more than 500 cross-site scripting vulnerabilities across the company's first-party web applications, while finding only two such flaws in apps built on its hardened web frameworks — a gap the company is presenting as proof that secure-by-design architecture can withstand relentless automated attack.
The disclosure came September 24 in a blog post by information security engineer Michał Bentkowski of Google's Product Security team.
The agent first uses Gemini models — most scans run on Gemini 3.1 Pro and Gemini 3.5 Flash, though other models can be plugged in — to identify suspected flaws.
Only after a working exploit succeeds does the issue get flagged as a confirmed bug.
Repeated attempts give the agent multiple chances to land on a working exploit.
Google's internal AI security agent has validated more than 500 cross-site scripting vulnerabilities across the company's first-party web applications, while finding only two such flaws in apps built on its hardened web frameworks — a gap the company is presenting as proof that secure-by-design architecture can withstand relentless automated attack.
The disclosure came September 24 in a blog post by information security engineer Michał Bentkowski of Google's Product Security team. PageBreak began as a pilot in November 2025 and became a full project in January 2026, with a mission to autonomously scale vulnerability discovery while cutting down on manual toil.
Cross-site scripting, or XSS, occurs when an attacker injects a script into a page that another user loads. Depending on the application, the script can read data or take over a victim's logged-in session.
How PageBreak Separates Real Bugs from Hallucinations
What sets PageBreak apart from a conventional large language model scanner is a two-stage architecture. The agent first uses Gemini models — most scans run on Gemini 3.1 Pro and Gemini 3.5 Flash, though other models can be plugged in — to identify suspected flaws. Each candidate is then handed to a purpose-built validator that attempts to fire an actual exploit against a running instance of the application.
For an XSS finding, the validator injects a JavaScript payload, loads the page, and determines whether the script executes. Only after a working exploit succeeds does the issue get flagged as a confirmed bug. Google said this validation step has kept the false-positive rate near zero, and unverified candidates never reach product teams as confirmed vulnerabilities. Instead, they feed back into later scans or help engineers build the next validator.
The agent also tests for other attack classes beyond XSS. Its validators check whether injected inputs alter database queries, expose files through path traversal, or trigger code execution. A separate validator examines requests that may reach internal services.
Google runs the same seed over multiple iterations because models can abandon a productive route or pursue an attack path that fails under real application conditions. Repeated attempts give the agent multiple chances to land on a working exploit.