Welcome to this edition of the AI Security Newsletter. This week the plumbing that connects agents became the attack surface: the same MCP flaw turned up at Google, JPMorgan and two governments, and a Bouncy Castle bug exposed MLS group chats, a protocol increasingly used for agent-to-agent channels, to impersonation. On the containment side, NVIDIA’s OpenShell, the open-source Coi sandbox and Oracle’s zero-standing-access playbook show how teams are fencing agents in. Safety pressure also reached the boardroom: Sam Altman tied OpenAI’s IPO to “confident safety claims,” and days later a former OpenAI safety lead published a sharp critique of the company’s culture. Elsewhere, Anthropic opened tiered cyber access for defenders, and Google paused part of its open source bug bounty after a flood of invalid automated reports.
Risks & Security
The same MCP flaw was fixed at Google, JPMorgan and two governments
Researcher Syed Anas Mohiuddin found the same server-side request forgery bug in Model Context Protocol servers at Google, JPMorgan Chase, Weaviate, France’s DINUM and Indonesia’s Tangerang city government. Each server built outbound requests from agent-supplied URLs without checking where they resolved. Google’s flaw, CVE-2026-14540 (CVSS 8.0), is fixed. Five servers run by the US General Services Administration remain unpatched. Mohiuddin also describes “protocol pivoting,” where an orchestrator passes injected, task-shaped text to a subagent that trusts it.
References:
- MCP for agent-to-agent comms may be the riskiest protocol you’ve never heard of – Ars Technica
- Google, JPMorgan and two governments fixed the same MCP flaw – The Next Web
- Researcher Discloses Same MCP Flaw at Google, JPMorgan, Two Governments – Unite.AI
Bouncy Castle MLS flaw let attackers impersonate group members (CVE-2026-71885)
Bouncy Castle for Java before 1.86 never checked that an MLS member’s certificate key matched its signing key. Where external commits are admitted without a separate credential check, an unauthenticated attacker could join a group as the victim, evict them, decrypt later messages and post as them. Deployments using only basic credentials are unaffected. Version 1.86 fixes the flaw. Forkast rates it critical (CVSS 9.2) and notes MLS’s growing use for agent-to-agent channels.
References:
- NVD – CVE-2026-71885
- New Release: Bouncy Castle Java 1.86
- Bouncy Castle CVE-2026-71885 Harvested the Credential Binding That Authenticates Agent-to-Agent Channels (Forkast)
- Information on source package bouncycastle (Debian Security Tracker)
Google pauses its open source bug bounty after a flood of automated reports
Since October 1, Google has stopped accepting product vulnerability reports through its Open Source Software Vulnerability Reward Program, which covers projects such as Go and Angular. Google cited “a significant rise in automated submissions, the vast majority of which are not valid” and said it will rework the program, with an update in Q1 2027. Supply chain reports are unaffected, and the Patch Rewards Program still pays up to $15,000. Google had tightened the rules in March.
References:
- Google Pauses OSS Product Bug Bounty Rewards After Surge in Invalid Automated Reports
- Google Narrows Open Source Bug Bounty Amid Wave of Invalid Automated Reports
- AI slop submissions force Google to freeze its open-source bug bounty
- Google suspends part of its open source bug bounty
How Oracle lets AI agents in without handing over control
Oracle described the identity controls behind its internal AI rollout. For ChatGPT Enterprise, OCI IAM now supports MCP with OAuth, PKCE and user consent tied to the employee’s identity. Read consent cannot authorize writes or exceed the employee’s own rights. For Codex and operations agents, production access moved to just-in-time grants, protected deletions need human approval, and long-lived API keys gave way to temporary workload identity. Oracle calls the result zero standing access.
References:
Technology & Tools
NVIDIA OpenShell adds kernel-level guardrails and formal policy checks for agents
NVIDIA’s open-source OpenShell runtime sandboxes autonomous agents without rewriting them. Landlock and seccomp limit files and system calls, and an out-of-process supervisor inspects HTTP, GraphQL and MCP traffic, so it can allow a read while blocking a write on the same API. Agents never see real credentials; OpenShell injects them only for approved endpoints. A policy prover uses formal verification to flag risky changes, such as credentialed access to a new host, before approval.
References:
- Add runtime controls to AI agents with NVIDIA OpenShell
- NVIDIA/OpenShell on GitHub
- NVIDIA OpenShell | Open, Secure Runtime for AI Agents
- Overview of NVIDIA OpenShell
Coi keeps AI coding agents in a container, away from host credentials
Coi (Code on Incus) runs Claude Code, Codex and similar tools inside Incus system containers. Only the project is mounted, and SSH keys, tokens and .env files stay out unless explicitly forwarded. Kernel-level monitoring watches for reverse shells, exfiltration and DNS tunneling, pausing or killing the container depending on severity, and egress can be locked to an allowlist. One exposure remains: the AI tool’s own API token and the workspace are still reachable from inside.
References:
- GitHub – coipond/coi
- Releases · code-on-incus (v0.12.0)
- Claude on Incus – All the autonomy, securely – Closer to Code
Guardana: open-source verification for AI artifacts, endpoints and agent traces
Guardana is a new Apache-2.0 project that produces reproducible security evidence for AI release decisions. It checks model artifacts, probes live endpoints and MCP servers, and inspects recorded agent traces for unapproved side effects, working outside the request path. It is in beta, and it is neither a general SAST scanner nor a compliance certification. A sister project, guardana/control, is an MCP gateway with signed policies, approvals and an evidence trail.
References:
- scadastrangelove/awesome-ai-security-tools
- owasp-llm · GitHub Topics
- Zero0x00/Ai-Security-radar-
- audit-trail · GitHub Topics (guardana/control)
Cantina and Yeta Labs release Apex Flash-1, an open-weights security research model
Apex Flash-1 is an MIT-licensed reinforcement-learning post-train of the 321B-parameter GLM-5.3-Flash, trained on 150 tasks drawn from 50 real vulnerability cases. On Cantina’s 60-task held-out evaluation it solved 40 at about $2.38, against 43 at $74.68 for Claude Opus 5 High. Cantina pitches it as a cheap worker model directed by a larger agent. The BF16 weights need about 640 GB of GPU memory.
References:
- cantina-security/apex-flash-1 (Hugging Face model card)
- Cantina Research | Security Benchmarks and Field Reports
- Cantina releases open-weight security model trained on 50 vulnerability cases (RuntimeWire)
AWS open-sources Strands Decider 2B, a small model that picks instead of generating
Strands Decider 2B handles the small decisions inside agent workflows: which tool to call, where to route a request, whether an action passes a guardrail. Built on Qwen3.5-2B with a tiny pointer head instead of text generation, it returns one of the supplied options with a calibrated confidence in about 115 ms on an RTX 3090. Its reference hook checks whether tool arguments are grounded in what the user actually said.
References:
- Amazon unveils a free, fast, open source Jev killer: Strands Decider 2B – VentureBeat
- Amazon releases its own Jev clone as decision models flood the web | TechCrunch
- AWS Releases Decision Model for AI Agents That Routes Without Generating Any Text – Tech Times
- Running Strands Decider 2B on Amazon Bedrock AgentCore
Google releases EmbeddingGemma 2, an on-device multimodal embedding model
EmbeddingGemma 2 maps text, code, images, video and audio into one shared vector space under Apache 2.0. The 740M-parameter model pairs a 270M text base with optional vision and audio encoders, so developers load only what they need. Its 8K-token context window is four times the first version’s. Quantized on a Pixel 11 Pro, it needs about 191MB of RAM for text alone and 567MB for the full model.
References:
- EmbeddingGemma 2: an open, lightweight multimodal embedding model
- EmbeddingGemma 2: The Developer Guide
- EmbeddingGemma 2 model card | Google AI for Developers
Dyna-2.1 runs an hour-long laundry workflow on its own
Dyna Robotics’ Dyna-2.1 is a “physical agent” built to complete whole workflows rather than single tasks. It runs on Taku, a wheeled semi-humanoid robot, and is controlled by three layers: a vision-language orchestrator that picks the next step, a world-action model pre-trained on about a million hours of human video, and a 100 Hz whole-body controller. In an uncut demo, Taku ran a hotel laundry room for about an hour.
References:
Business & Products
Anthropic splits its Cyber Verification Program into three access tiers
Anthropic merged Project Glasswing into its Cyber Verification Program, now in three tiers with fewer blocking classifiers on Opus 5.5, Sonnet 5.5 and Mythos 5.1. Defense Access, open to individual researchers, covers SOC work and incident response. Red Team Access adds authorized penetration testing for organizations. Specialized Access covers critical systems such as power grids. On CyScenarioBench, Opus 5.5 was blocked on all 50 runs without program access but completed 34 unblocked under Red Team Access.
References:
- Expanding the Cyber Verification Program (Anthropic)
- Real-time cyber safeguards on Claude Opus and Sonnet (Claude Help Center)
- Anthropic folds Project Glasswing into an expanded three-tier cyber verification program (SiliconANGLE)
- Anthropic loosens Claude’s cyber restrictions for verified … (Help Net Security)
Claude for Government reaches general availability with FedRAMP High authorization
Anthropic made Claude for Government generally available to US federal and state agencies on September 30, ending a beta that began in July. It brings Claude’s coding and agent capabilities into a FedRAMP High authorized environment, with Claude Code CLI and Claude for Microsoft 365 in early access. Administrators get per-department spending controls, audit logs and SCIM user management. Pricing is usage-based with no seat fees, and the service is not a DoD IL5 environment.
References:
- Claude for Government is now generally available
- Claude for Government | Claude by Anthropic
- Anthropic Claude for Government Reaches GA: FedRAMP High Controls and Microsoft 365 Early Access
Altman: no OpenAI IPO until it can make “confident safety claims”
After his September 29 DevDay keynote, Sam Altman said OpenAI will not go public until it can “make confident safety claims” about its most capable models, calling an IPO during this shift “ill-advised.” He gave no timeline. The remarks follow incidents in which OpenAI agents hacked Hugging Face and government websites during testing, which OpenAI says took weeks or months to detect. The same day, a nonprofit sued OpenAI in California seeking stronger evaluation and monitoring.
References:
- OpenAI delays IPO over AI safety concerns (Financial Times)
- Sam Altman says OpenAI won’t go public until its models … (The Verge)
- Sam Altman: OpenAI won’t go public this year as IPO now … (Fortune)
- No OpenAI IPO Until the AI Stops Going Rogue, CEO Sam Altman Says (Gizmodo)
Cloudflare Birthday Week brings post-quantum certificates and agent payments
Cloudflare made 46 announcements during Birthday Week 2026. It plans to become a public certificate authority issuing free post-quantum Merkle Tree Certificates, and it aims to finish its own post-quantum migration by 2029 with help from an AI tool, CryptoLabe. For the agentic web, it launched Pay Per Use and a Monetization Gateway beta built on HTTP 402. It also released an agentic cf CLI and Kitesurf, a browser for agents.
References:
Opinions & Analysis
OpenAI safety lead David Robinson quits, saying the company’s culture is broken
In an October 3 Atlantic essay, David Robinson explained why he left OpenAI’s Safety Systems team, where he helped draft the Preparedness Framework and oversaw system cards for 12 frontier launches. He argues that “iterative deployment,” strengthening safeguards as problems appear, no longer works: “The time for trial and error is over.” He wants safeguards closer to those in aviation and nuclear power, and says stronger safety incentives must come from outside the company.
References:
- OpenAI safety employee resigns, claiming the company’s ‘culture is broken’ (TechCrunch)
- OpenAI safety employee quits, criticizes company’s approach to AI risks (Investing.com / Reuters)
- “The time for trial and error is over”: OpenAI safety … (CTech)
Teleport: Zero Trust is necessary but insufficient for AI agents
Teleport’s white paper “From Zero Trust to Agent Trust” argues that agents’ speed, scale and unpredictability can cause harm even within valid permissions. It extends the classic principles. “Enforce continuously” gives every agent an attested identity and a runtime that starts with zero privileges. “Bound collective autonomy” applies least privilege to swarms, so actions that are harmless alone but destructive together, such as changing security configurations, are escalated for review.
References:
- From Zero Trust to Agent Trust
- Agentic Zero Trust: Extending the Zero Trust Security Model (research paper)
- New tools and guidance: Announcing Zero Trust for AI
McKinsey: AI gains depend on redesigning the organization around human–agent teams
McKinsey argues that capturing value from AI requires redesigning roles, workflows and operating models, not just adding tools. In a survey of more than 700 executives, the companies furthest along organize cross-functional teams around end-to-end products and customer journeys. As agents take on work, organizational design becomes orchestration: deciding which humans and agents own an outcome. Flatter structures pay off only with clearer decision rights.
References:

Leave a comment