
Deepfakes and voice clones change the verification ritual, not the underlying control. A callback to a known number still beats a model’s confidence score on a waveform. That is a human procedure. Automating the “yes, that sounded like the CFO” step is how you get a wire out the door.
Shadow AI inside the business is the same pattern inverted. Staff paste incident data, source code, and customer records into unapproved tools. Verizon’s later DBIR cycle has already flagged unapproved AI use as a non-malicious action path that still produces breaches. The control is policy plus monitoring plus a sanctioned alternative. It is not a banner that says “we use AI.”
What the models do well, and what they do not
That is not Luddism. It is how you treat any new principal with standing privileges. You already do this for service accounts. Treat the agent as a service account that talks. Log every tool call. Bind actions to a human-owned change window when the blast radius is large. Keep a kill switch that does not require the agent’s cooperation.
Defenders talk about AI as if it were a one-sided productivity grant. It is not. Generative models collapse the cost of personalized phishing, fluent BEC, and lures that no longer fail a grammar check. The 2025 DBIR still found 88% of basic web-application attacks involving stolen credentials. Identity remains the cheap door. AI makes the social engineering that produces those credentials cheaper and more convincing. It does not invent a new physics of access. It compresses the time between reconnaissance and a believable message.
GovTech’s reading of the same research cycle is that staff and budget cuts are raising perceived risk while AI adoption is rewriting which skills count. Shrinking the team because a copilot demo reduced mean time to first comment on a ticket is a category error. Comment latency is not containment quality.
The pitch is simple: models triage faster than people, so the SOC should shrink. That pitch collides with how incidents actually start and how they end. Verizon’s 2025 Data Breach Investigations Report still put a human element in 60% of breaches. Stolen credentials were the most common initial-access path, at 22%. Phishing accounted for 16%. Those are not model-scoring problems. They are decisions under incomplete information, inside a business that will not pause while a classifier finishes its next token.
Attackers bought the same tools
Eye Security’s operators put it without ornament: “Technology informs decisions; analysts take responsibility for them.” That sentence is the liability model. A copilot can be wrong and still be useful. An unsupervised containment action that is wrong is an availability incident with a security ticket attached.
Your IR retainers and tabletop exercises exist because the last mile is political. Who calls the board. Who talks to the insurer. Who accepts residual risk on a zero-day that has no patch and a compensating control that is ugly. Models do not sign the risk acceptance. People do. If you have ever watched a CISO choose to leave a vulnerable internet-facing system up because the alternative was a plant shutdown, you already know the limit of autonomous response.
The traditional SOC, as UnderDefense describes it, still holds institutional knowledge and relationship-driven response. That is not a compliment to process documentation. Most of that knowledge is not written down. It lives in people who remember the last time finance ran a one-off integration that looked like C2.
Judgment is the residue that remains after those tasks are automated. An agent constrained to certain data is access-controlled. An agent whose planning loop is monitored for behavioral drift is governed. Those are different security postures. Vendors selling “autonomous SOC” blur them on purpose. Their validation is operationally meaningful but not independent in the audit sense: it measures time-to-triage on the vendor’s own alert corpus, not whether the model should have pulled a kill switch on a payment processor during a peak window.
Automation bias is the new alert fatigue
Panther notes that trust collapses when AI systems generate too many wrong calls, and that teams then turn the tool down. That is the other half of the same problem. Over-trust and under-trust are not opposites in operations. They are sequential. First the dashboard is gospel. Then it is ignored. Neither state is a detection program.
You already see this in the inbox and in the help desk. The old tells (bad English, odd domains, mismatched logos) are decaying as discriminators. What remains is process: out-of-band confirmation, dual control on irreversible actions, and an analyst willing to slow a transaction that looks clean in the tooling.
The contrast is not “AI versus humans.” It is AI-on-AI noise versus a person who can refuse the statistically plausible story. Models optimize for coherence. Attackers now generate coherence on demand. Coherence is no longer evidence.
Replacement talk is a budget story. The 2025 ISC2 Cybersecurity Workforce Study found AI was the most pressing skills need, cited by 41% of respondents, ahead of cloud security at 36%. That is not a survey of people asking to be replaced. It is a survey of people who cannot yet govern, red-team, or securely operate the systems being dropped on them.
Business context does not live in the training set
Insider cases make this concrete. Anomalous data access by a departing executive is not a malware family. It is HR, legal, and a communications plan. A model can flag the DLP hit. It cannot decide whether this is theft, a messy offboarding, or a litigation hold. Treating those as the same class produces either a missed exfil or a wrongful-termination problem.
The next product wave is not a better chatbot. It is agents with tools: ticketing, isolation, identity revoke, cloud CLI. Tool use turns a language model into an actor. An actor needs a control plane that is not the model itself.
Machines are now exceptionally good at processing security telemetry at a scale no human team could match. Correlation across SIEM, EDR, identity, and cloud logs is cheaper than it was. Natural-language interfaces can draft investigation queries and summarize a case file. None of that is controversial among people who have shipped a detection engineering pipeline.
A detection that no one will act on is theater. An action no one will own is a different kind of incident. The industry can automate the middle. It cannot automate accountability. That is the constraint. It is also the reason the expensive humans stay on the loop, not as a courtesy to tradition, but because someone still has to decide.
The workforce gap is skills, not headcount theater
Walk into the average enterprise SOC and inventory what is running there. You will find a stack of detectors, a handful of SOAR playbooks, and a queue that still requires a person to say “this is a true positive and here is the blast radius.” The model does not know which subsidiary is in a regulatory exam. It does not know that the “anomalous” admin account is the break-glass identity used every quarter-end. Context is not a missing feature. It is the job.
Allow-lists on which APIs the agent may call are necessary and insufficient. Prompt injection against a SOC agent is not a party trick. It is an attacker sitting in a ticket, a log line, or an email the agent was told to “investigate.” The model’s job is to follow instructions. The attacker’s job is to become those instructions. Human judgment here is architectural: what the agent is forbidden to do even if it is highly confident. No unsupervised production isolation. No unsupervised identity disable for VIP accounts. No unsupervised changes to conditional access. Dual control for irreversible actions, the same way you already treat wire transfers.
A playbook can isolate a host. A playbook cannot tell you whether isolation of that host takes down a hospital imaging system or a batch job that can wait. East-west traffic after a phish is a graph problem. Whether to burn the session, the device, or the identity is a judgment about containment versus continuity. Breakout time is measured in minutes. The wrong containment is measured in outage tickets and, in regulated environments, in examiner letters.
The Verizon DBIR program continues, year after year, to show the same cluster: people, credentials, vulnerabilities, ransomware. The 2026 edition still lists social engineering, phishing, and stolen credentials among the frequent causes. Models will score more of those events faster. They will not make the human element go away, because the human element is both the target and the last responsible operator.
Agents raise the floor for human governance
You will be asked to show a headcount reduction next budget cycle because a vendor slide had a robot in a chair. Measure something else. Measure missed true positives after the model’s confidence threshold was raised. Measure how often analysts override the copilot, and whether those overrides were correct. Measure time-to-understand, not time-to-summarize. Summaries are cheap. Understanding is the scarce resource.
AI is already inside detection, enrichment, and ticket drafting. That is real work. It is not the same work as deciding whether to isolate a revenue system or not. Those are different jobs.
The last decade’s failure mode was volume. Too many alerts, too few people, too much muscle memory clicking “close.” The next failure mode is the inverse: too much trust in a ranked list. Dropzone (a vendor with a product in this category) warns that if AI is assumed to be flawless, teams miss red flags. Their interest is obvious. The failure mode is still real. Automation bias is well documented outside security. Inside a SOC it looks like an analyst who stops opening the raw telemetry because the summary sounded complete.
False positives from a rules engine were annoying. False positives from a generative layer are persuasive. They arrive as narrative. Narrative is harder to dismiss at 03:00 than a red row in a queue. That is why human review is not a nostalgia play. It is a check on fluency.
Redbud Cyber’s framing is blunt: “An AI system might identify technically optimal responses that prove operationally catastrophic.” Technical optimality is a local maximum. It ignores union rules, customer SLAs, merger quiet periods, and the fact that the crown-jewel database is mid-migration. None of that is in the foundation model. Some of it is not even in the CMDB.




