What a Top AI Jailbreaker Could Not Break Into, and Why It Matters for Your Business
A respected jailbreaker got five blind attempts at a personal AI agent and every attack was quarantined. The defenses that held, a human in the loop, a strong scanner, rate limits, are the same ones any business running AI on real data should copy.

Most businesses running AI agents are protecting against the wrong attack. The scenario that dominates every security conversation is data extraction: a bad actor reading private records, exposing credentials, or pulling sensitive information through a compromised system. That threat is real, and I am not dismissing it. But for the majority of small and mid-size businesses that are now giving AI agents access to their inbound messages and internal tools, the financial siege attack that costs real money without touching a single file is the more immediate and more overlooked danger. I, Madhuranjan Kumar, want to make this argument plainly, because every owner I speak with has thought about the data-theft scenario and almost none of them have set the one control that stops the other kind of damage.
A test published recently makes the case in concrete terms. One of the more respected AI jailbreakers active today, a person with a track record of cracking frontier AI models within minutes of their public release, was given a single email address. That address is the one a personal AI agent monitors. He had no knowledge of the model running underneath, no information about the architecture, no awareness of what hardening had been applied. He was given five structured attempts to break in. He estimated his own odds of cracking at least the early levels above 80 percent. Every attempt was quarantined. The defenses that held here are reproducible, and understanding why each layer earned its place is more useful than treating the outcome as proof that these systems are impenetrable.
The security test nobody expected a jailbreaker to lose
The test was deliberately unfair to the defender. A blind attack, with no advance knowledge of the target, is the hardest scenario any system can face because you cannot prepare for what you cannot see. Working blind, the first task is identifying which model is running, so you can tailor later attacks to its specific documented weaknesses.
The opening instrument was a token bomb: a payload disguised as something innocuous, a single element that secretly carries an enormous amount of hidden text. One such element in this test carried roughly three million characters, designed to flood the model and cause it to misbehave in ways that reveal its identity. Token bombs serve a dual purpose. They probe the model and, when coordinated into waves, they execute the financial attack that most owners have never considered.
After the probe came a sequence of escalating attempts. A jailbreak template with the obvious trigger phrases stripped out, designed to slide past pattern-matching filters that catch the well-known templates. A payload formatted to look like a legitimate internal system command, betting that the quarantine would treat an external attacker's instruction as a trusted next step from inside the architecture. A format-override attempt designed to gain just enough output control to confirm model identity. None of it reached the agent. Each attempt was caught and quarantined before any injection could occur in the live system.
The architecture that produced this result has four components. Each one earns its place independently, and removing any single layer degrades the protection below what a skilled attacker needs to succeed.

The siege attack costs you money without ever opening a file
Here is the threat that is almost never on a small business security checklist. A siege attack does not need to extract data or take any action inside your system. The attacker sends a coordinated wave of token-bomb payloads in rapid succession. Each one forces the AI agent to process an enormous volume of tokens. Do that across dozens of simultaneous inbound messages and the agent is running at high volume continuously until the API payment limit is reached. The attacker achieves disruption and potentially a genuinely painful bill, without ever reading a record, compromising a credential, or learning anything meaningful about your system.
This attack works because most AI API accounts have no hard spending cap applied by default. The major providers do not set a limit automatically. The account owner has to configure one deliberately, and most owners who have not thought specifically about siege attacks have not done it. Setting a per-day or per-source spending cap takes three to five minutes in the API account settings and permanently eliminates this entire attack category at zero ongoing cost. There is no engineering required. You type a number into a field in the billing dashboard.
Consider what a siege attack actually costs in practice. An agent processing inbound messages might handle a few hundred tokens per message at baseline. A token bomb containing three million characters forces millions of token-equivalents of processing through the system. If the attacker sends twenty of these in a rapid sequence, the cost of that single incident at standard API pricing can run to a figure that appears noticeably on a small business billing statement. Do this repeatedly across a day, targeting an account with no spending cap, and the damage is real before any human on the account side realizes what generated the charges.
The asymmetry between attacker effort and business impact is the reason this threat deserves more attention than the data-extraction scenario for most owners. A successful data extraction requires getting through a quarantine, past a capable scanner, executing an action without a human checkpoint intervening, and returning something valuable to the attacker. That is genuinely difficult against a well-configured system. A siege attack that generates unexpected charges requires sending large messages in volume. The skill and infrastructure required is orders of magnitude lower, and the business impact, a meaningful charge appearing before anyone checks the dashboard, can be operationally disruptive even when no data was ever at risk.
For businesses running paid acquisition through Meta ads or Google Ads, there is a compounding consideration. High-volume campaigns drive high volumes of external submissions into the AI pipeline. Every lead form, every chatbot contact, every inbound message from an ad-driven visitor is a potential injection point. The attack surface grows directly with marketing volume. A business that grows its inbound traffic without hardening the AI systems processing those submissions is building a vulnerability that expands in direct proportion to its marketing success.

A cheap model at the front line is how the whole stack collapses
When the winning architecture was revealed, the key detail was this: the front-line scanner was Claude Opus with extended reasoning enabled. A capable reasoning model sat at the first checkpoint, evaluating every inbound message before it reached the agent.
That decision is load-bearing in a way most owners underestimate. Tested directly during the experiment, Claude with reasoning turned on flagged embedded injection instructions, identified attempts to redirect it to external URLs, and refused to act on disguised commands. The attacker confirmed this when he tested the scanner directly as part of his probe. A lightweight, fast model chosen for the scanner position because it costs less falls for exactly the injection attempts that a reasoning model catches. The attacks that succeed against small front-line models are not obscure edge cases. They are the standard toolkit that any experienced attacker uses as a first step.
The cost calculation is worth making explicit. A reasoning model at the scanner position costs more per token than a lightweight alternative. For a business processing a few hundred inbound messages per day, that difference might amount to a few extra dollars per month. The cost of one successful injection into an agent handling patient records, financial data, or client credentials is orders of magnitude larger in direct and reputational terms. Optimizing cost at the scanner layer by choosing a lightweight model is the mistake that turns a layered defense into a single-point-of-failure system, because if the scanner passes a sophisticated injection attempt, every layer behind it is starting from a compromised position.
This applies directly to any business routing inbound contacts through an AI processing layer. If your SEO and organic search traffic and your paid acquisition flow into the same AI intake pipeline, the model quality at the first checkpoint determines the security posture of the entire inbound funnel. The scanner is not the place to reduce cost.
The human checkpoint is not a bottleneck, it is the architecture
The highest-ranked defense in this test was not a technical layer. It was a policy: no consequential action executes without a human seeing it first. This rule is the reason that even a hypothetical scenario where an injection gets through the quarantine and past the scanner still cannot produce a real outcome. An unauthorized action cannot run if a person must approve it before execution.
Owners often frame the human-in-the-loop requirement as a capability gap, something to remove as confidence in the AI grows. That framing misunderstands the architecture entirely. The human checkpoint is not a concession to limited AI capability. It is the structural guarantee that no action, however it arrived in the pipeline, can execute without a person in the approval chain. As AI agents grow more capable and take on more significant tasks with less friction, the importance of this checkpoint increases rather than decreases. A more capable agent that can take more significant actions is precisely the agent that most needs a firm checkpoint on anything consequential.
There is also a practical compliance benefit for any client-facing business with real professional liability. Every action the agent proposes is reviewed and approved by a named person before it executes. That creates an audit trail. The trail has value for regulatory compliance, for internal review when something goes wrong, and for demonstrating accountability to clients and regulators. A medical practice, a legal office, a financial advisory, any business with professional liability requirements, gets a compliance record as a natural side effect of the security architecture rather than as a separate administrative burden layered on top.
Four moves that price most attacks out of being worth attempting
Here is a concrete example that shows how the four layers interact in a real business context. A dental practice runs an AI agent that reads patient inquiry emails, sorts appointment requests by urgency, flags cases that need same-day attention, and drafts responses for staff review. The agent has access to the scheduling system and can see patient names and appointment histories. Without hardening, every inbound email is a direct injection point.
With the four layers in place, the flow is different at every step. Every inbound email passes through a quarantine step that checks its size, flags unusual structure, and isolates any message containing embedded instructions or exceeding a normal character count. A token bomb carrying millions of characters is detectable at the quarantine step long before it reaches any model. Messages that pass quarantine go to the Claude Opus scanner with reasoning enabled, which reads for injection attempts, disguised commands, and format-override requests. Messages that pass the scanner reach the agent, which drafts a response or scheduling proposal. That draft goes to a staff member for approval before anything is sent or any schedule is changed. The API account has a per-day spending cap set at a level that covers normal operation with meaningful headroom, so a coordinated wave of oversized messages cannot generate a surprise charge even if some pass the initial size check.
The numbers make the trade-off clear. The practice pays a modest premium for the Opus scanner over a lightweight alternative, perhaps an extra five to fifteen dollars per month at typical message volumes for a small clinic. It saves the equivalent of two to three staff hours per day in triage and response drafting. The security posture is dramatically stronger than an unhardened deployment, and the human approval step produces a compliance-ready audit trail at zero additional administrative cost.
For any business managing customer contacts through a CRM and website stack, the same four layers apply regardless of industry or message volume. Quarantine inbound content before the agent reads it. Put your strongest reasoning model at the scanner position. Define which agent actions require human approval before execution and make that policy explicit in the agent configuration. Set a hard spending cap per day and per source in the API account settings.
None of these four moves require engineering expertise. They require configuration decisions, each of which takes minutes rather than days. The spending cap is the fastest of the four and the one most commonly missing. Set it today. Then build the remaining layers in order, and the architecture that stopped five structured attacks from a skilled, motivated attacker is the one your business runs on.
The insight from this test is not that AI agents are invincible. It is that four deliberate decisions separate a system that stops professional-grade attacks from one that gets compromised on the first attempt. The siege attack, the one that costs real money before anyone checks the billing dashboard, is the threat most owners have not thought about. It is also the fastest threat to neutralize. Make that change first, then build the rest of the stack around it.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
