Treatmybrand


a Kainjoo SA Venture
Ch. du Vernay 14a
1196 Gland
+41.21.561.34.96
[email protected]

Support


Monday to Friday
8AM to 8PM
[email protected]
Back

Anthropic’s Browser Agent Vulnerability: 31.5% Hijacking Rate Before Safeguards Kick In

Anthropic’s latest browser-based AI model, Opus 4.8, showed a high prompt injection vulnerability, with attackers successfully hijacking it 31.5% of the time before safety measures activated. This figure, detailed in their extensive 244-page system card published on May 28, highlights the challenge of securing AI agents in diverse environments. Unlike OpenAI, Google, and Meta, Anthropic provided transparent, per-surface data across four operational surfaces, including browsers, coding, and tool use, revealing a stark contrast in exposure rates and protective effectiveness. For instance, in coding environments, the attack success dropped from 7.03% to 2.09% once safeguards engaged. Meanwhile, OpenAI reported a robustness score without direct attack rates, Google provided qualitative resistance claims without numeric data, and Meta focused on grading guardrails rather than direct model vulnerability. The article urges security teams to diligently assess AI agent deployment surfaces, request detailed attack success metrics from vendors, and implement rigorous red-teaming themselves prior to deployment to mitigate risk. It underscores that no unified industry standard exists for measuring prompt injection risk yet, making buyer vigilance crucial.

Venturebeat
Venturebeat