Claude Code Head Claims 720 Prompt Injection Attacks with 0 Success
Boris Cherny, head of Claude Code, stated that Anthropic has essentially resolved prompt injection attacks in practical use. A year ago, he never imagined reaching this level. Prompt injection attacks have been a significant issue for Agents, where malicious web pages might hide instructions that induce Agents to steal passwords, upload keys, or even perform actions not requested by users. Anthropic employs a three-layer defense mechanism: Claude itself resists malicious instructions, external content is checked before entering the context, and before executing any operation, Auto Mode conducts another review. Third-party testing utilized 72 previously unseen attacks, repeated 720 times, and neither Sonnet 5, Fable 5, nor Opus 5 succeeded after enabling Auto Mode. In contrast, the attack success rate for GPT-5.6 Sol using Codex Auto-review was 5.83%, while Full Access reached 19.03%. This also explains why Anthropic has set Auto Mode as the default; it serves as the final check against prompt injection attacks for Claude Code.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

OpenAI's Tibo Changes Profile Picture, Suggesting a Reset of Quotas

B.AI API Compatible with Codex Supports Calling GPT Series Models

Codex Weekly Users Reach 5 Million as OpenAI Expands Work Agent

SCAN 2026 Finals, 67 Participants in Solo Teams

web3: Z.ai Releases GLM-5.3 Programming Model Targeting Open Source Weight Market

OpenAI Tests Paid Reset for Codex, Pro Up to $80

Apple Limits Vulnerability Reports and Strengthens AI Verification

Tencent WorkBuddy Launches AI Collaborative Editing Feature Supporting Word, Excel, and PPT

Alibaba Qoder Releases Open Source AI Programming Tool Better Harness

Lingguang App Launches One-Click Deployment Feature Supporting Multiple AI Applications

Block Announces Open Source AI-Powered Chat "Buzz" to Compete with Slack and GitHub

Founder of Baixing.com: My Experience with Claude Code in Fourteen Points

Analysis: Anthropic and OpenAI have consecutively exposed security vulnerabilities, raising concerns about the safety of AI models

Tencent's Zhang Jun: WeChat Work CLI open source project launched on GitHub community

OpenAI resets Codex usage limits, in conjunction with the launch of the plugin system

OpenAI plans to merge ChatGPT, Codex, and the browser into a desktop super app

OpenAI releases GPT-5.4, positioning it as a cutting-edge professional work model

AI Workers Earn $400 Million Annually, Virtuals Aims to Make You a Shareholder

Investment Options for FAL: Only 8 ONs, Green Light for Sovereign Bonds, and Currently Excluded Provincial Bonds

Bitcoin Mining Uses 30% of Paraguay's Electric Energy: Crisis Warned for 2029

BCRA purchases exceeded $14 billion barrier in 2026

Irish drug dealer’s lost wallet moves $39.56M in Bitcoin

HyENA shuts down after processing $4B in trades

Avici attack drains over $1M from Solana users

Cosmos misjudged a critical bug for 4 months before hackers stole nearly $6 million across 6 chains

US Cities Eye AI Data Center Limits as Study Flags Water Risks

Tokenized gold is becoming productive collateral in crypto lending, Arch says

European Blockchain Convention 2026: Information and Exclusive Promotion

1200 Qubits Estimated, Pressure on Bitcoin and Ethereum Transition


