Governance · 5 min read · Updated 2026-10-08
Evidence before inference: what Cloudflare's new SOC agents get right
Cloudflare has rebuilt its Managed Defense alert triage around a rule most agentic AI projects skip - collect evidence deterministically first, let the model reason second. The architecture is a useful benchmark for any enterprise building agents that touch security decisions.
Cloudflare has published how its Managed Defense service now triages security alerts with a team of AI agents, and the detail worth reading twice is not that it uses agents. It is the order of operations it insists on (Cloudflare, 8 October 2026).
What happened
Managed Defense runs "a team of specialized AI agents built on Workers and global network telemetry to analyze security alerts," and the design separates deterministic evidence collection from model inference before any recommendation is generated (Cloudflare blog, "Building an evidence-grounded agentic security operations harness on Cloudflare", 8 October 2026).
Concretely: before a model is called, application code runs "a fixed set of reconnaissance workflows with versioned API calls" to gather customer identity, detection history, traffic baselines, enforcement outcomes and network observations. Four specialist agents then investigate in parallel - traffic behaviour, customer history, network-wide telemetry, threat intelligence - and a synthesis agent combines their findings into one advisory using a fixed vocabulary, without fetching new evidence of its own. Cloudflare says the earlier prototype failed in three specific ways it had to design against: "context became authority" (a detection was treated as proof rather than a hypothesis), "scope drifted" (agents queried the wrong account or time window), and "failure disappeared" (a timeout was indistinguishable from a clean result). The fix is that "application code checks that every citation exists, belongs to the investigation, and supports the attached claim," and the system records three distinct states for any fact - not checked, checked with no result, checked with evidence of absence - rather than filling the gap with an inference. Where evidence is insufficient, "it makes no classification or disposition recommendation" at all. The human Managed Defense Analyst still owns the decision; the system's job is to hand over the evidence behind the advisory, not the conclusion itself. The service is in early beta on application-security alerts.
Why this is not an isolated problem
Two independent data points explain why Cloudflare built it this way rather than shipping a single general-purpose agent.
First, the raw volume that security teams already drown in. Prophet's State of AI in Security Operations survey of roughly 300 CISOs and SOC leaders found the median team facing close to 960 alerts a day, with around 40% never investigated at all and roughly 90% of the ones that are investigated resolving as benign (Prophet, via NHI Management Group, 30 July 2026). An agent layered on top of that volume either cuts through the noise with something an analyst can trust, or it becomes one more unverified voice shouting into it.
Second, the reliability ceiling of the models doing the reasoning. Artificial Analysis's AA-Omniscience benchmark, evaluating 40 AI models in 2025, found that all but four were more likely to give a confident, incorrect answer than a correct one on difficult questions (Artificial Analysis, 2025; coverage: The Hacker News, May 2026). A model with that failure rate, asked to classify a security alert without a check on its citations, will occasionally classify confidently and wrongly. Cloudflare's architecture exists specifically to catch that case before it reaches an analyst.
Put the two together: enterprises are adding AI reasoning to a function that is already saturated with noise, using models that are documented to be confidently wrong on hard questions a meaningful fraction of the time. An agent that cannot show its evidence is not a shortcut through that problem. It is a second layer of it.
What it means for a regulated enterprise
If an AI system contributes to a security decision - flags an alert, recommends a WAF rule, scores a threat - a regulator, auditor or insurer can reasonably ask what evidence it had and what it did with the gaps. The EU AI Act's record-keeping expectations assume an organisation can reconstruct what a system did and why; a system that cannot distinguish "we checked and found nothing" from "we never checked" cannot answer that honestly. NIS2's management-level accountability for cybersecurity risk puts the same question on a named individual rather than a department. None of this is specific to Cloudflare's customers. It applies to any enterprise letting an agent touch a detection, a classification, or a remediation step, in security or anywhere else a wrong confident answer has consequences.
What actually addresses it
The mechanism, stripped of vendor framing, is simple to state and hard to retrofit: gather evidence with deterministic code first, version and timestamp every piece of it, run model inference only against that fixed evidence set, and verify every claim the model makes back against a citation before it is shown to anyone. Where the evidence is incomplete, the system has to be able to say so rather than fill the gap with a plausible guess. That last part is the one most agent builds skip, because "I don't know" is a harder output to engineer than a confident-sounding answer.
How AANCER answers
AANCER's append-only audit ledger exists for the same reason Cloudflare built citation checking: a recommendation an agent makes is only as useful as the evidence a human can pull up behind it afterwards. Every retrieval, every tool call and every model answer an agent produces is logged with its source, so a security or compliance reviewer can trace a conclusion back to what the agent actually saw, not what it is claimed to have seen.
Passports scope what each Certified Agent can reach before it runs, which narrows the "scope drifted" failure mode Cloudflare describes - an agent cannot query an account or a system its passport does not name, so a wrong-scope query is blocked structurally rather than caught after the fact. Approvals gate the steps that matter, so a classification or a remediation recommendation routes to a person before it becomes an action. None of this makes an underlying model more accurate on hard questions. It makes the gap between what a model asserts and what it can prove visible to the person who has to answer for the decision.
What to check on Monday
Pick one agent already running in production and ask it, or whoever owns it, three questions. What evidence did it use for its last five decisions, and can anyone retrieve that evidence today? Does it have a documented way to report "I don't know" rather than defaulting to its best guess? And if it made a wrong call last week, is there a log that would show you that - or would you only find out from the person downstream who acted on it?
Sources - Building an evidence-grounded agentic security operations harness on Cloudflare - Cloudflare, 8 October 2026 - Alert fatigue is a capacity problem, not a tuning problem - NHI Management Group, citing Prophet's State of AI in Security Operations survey, 30 July 2026 - How AI Hallucinations Are Creating Real Cybersecurity Risks - The Hacker News, May 2026, citing Artificial Analysis's AA-Omniscience benchmark, 2025
Related guides
Compliance
The EU AI Act Article 12 readiness guide
What record-keeping and human-oversight obligations actually require operationally from August 2026 — and the evidence an auditor will ask you to produce.
9 min read
Read the guide →Risk
The credentials nobody reviews
Your AI agents hold OAuth tokens, API keys and service accounts that went through no approval process. The agent was reviewed. The studio was reviewed. The identity behind them was not.
5 min read
Read the guide →Security
When the agents organised themselves: what the Hugging Face swarm means for accountability
Roughly 700 AI agents divided labour, traded favours and compromised production infrastructure across four regions. The uncomfortable part is not that it happened — it is that the account of what happened had to be reconstructed afterwards, by outside parties.
6 min read
Read the analysis →