The hardest part of security operations isn’t finding alerts — it’s drowning in them. One alert tends to trigger others across an environment, and a human analyst has to figure out which are related, which are noise, and which deserve a page. Cloudflare’s answer, described in an October 7, 2026 post about its Managed Defense service, is a multi-agent harness built on its own developer platform — and the design choices are worth studying even if you never touch a SIEM.
Why the single-agent prototype failed
Cloudflare’s first prototype handed one general-purpose agent the whole investigation. The output was useful, but the agent hallucinated claims the evidence didn’t support. When telemetry, detector descriptions, policies, and threat intelligence all got flattened into one prompt, their distinct roles blurred together.
The team identified three recurring failure modes. Context became authority: a detection is a hypothesis, but the agent treated it as proof. Scope drifted: nothing stopped the agent from querying the wrong account or time range, because a prompt is not a boundary. And failure disappeared — a timed-out lookup was indistinguishable from “checked and not found.”
Recon before any model call
The fix is architectural, and it’s the most transferable idea in the whole design. The front half of the harness contains no AI at all. Deterministic code runs versioned API workflows to collect identity, detection history, traffic baseline, enforcement outcome, and network observations, storing each piece with its source, version, and timestamp.
This fixed snapshot does two jobs. It makes scope enforcement a property of application code rather than a suggestion to the model. And it makes evaluation reproducible: if agents fetched their own data, two runs could disagree because their inputs changed. With a fixed snapshot replayed across runs, differences between specialist agents come from interpretation, not retrieval.
Narrow specialists and a synthesis agent that can’t wander
Alerts that survive a lightweight triage pass — Clef, running on Workers AI, filters likely false positives deterministically — go to a coordinator running four specialists in parallel: traffic analysis, customer context, global telemetry, and threat intelligence. Each works only from the admitted evidence package and must cite it. Application code verifies every citation exists, belongs to the investigation, and supports the attached claim.
The synthesis agent combines their typed findings into an advisory, but it can’t fetch new evidence or pick a classification outside an approved vocabulary. That constraint is what makes catching unsupported claims feasible. It echoes a theme I covered before in Cloudflare’s observability platform updates: turn telemetry into something an agent can query within defined boundaries, rather than giving it raw access and hoping.
Global context is handled carefully too. The global telemetry specialist sees only privacy-preserving aggregates across Cloudflare’s network — never another customer’s individual records. So an IP scanning thousands of sites reads differently from one targeting a single customer, without cross-tenant leakage.
Building on your own primitives
The stack is a reminder of what running your own developer platform buys you. Workers admits evidence and validates results; Workflows coordinates stages and checkpoints completed work so a failure reuses validated findings instead of restarting; D1 and R2 hold investigation state and evidence artifacts; Durable Objects persist case-chat state.
Just as deliberate is how incomplete evidence is handled. The advisory distinguishes “not checked,” “checked with no match,” and “checked with evidence supporting absence.” When evidence is insufficient, the system makes no classification at all. Analysts keep the final decision and can inspect, revise, or expand any recommendation into a remediation rule.
The harness is in early beta for eligible Managed Defense customers. The broader takeaway for builders: the reliability of an agentic system came less from model quality and more from what Cloudflare refuses to let the model do — fetch its own evidence, cross tenant boundaries, or classify without sufficient proof. That’s a harness pattern worth stealing for any domain where a wrong answer costs real money.
