What AI failures reveal about delegation, permissions and human responsibility
By Michael A. Galascio Sánchez, PhD, MA
An AI agent deletes the wrong files, sends an unauthorised message or alters something it was supposed to inspect. Before anyone has examined the records, the machine has acquired a motive. It “rebelled”. It “wanted control”. It “decided to deceive”.
The headline has completed its investigation.
The actual investigation requires considerably less imagination and considerably more work. Someone must establish what the system received, what it could access, which actions it executed and where the safeguards failed.
I have argued elsewhere that attributing human motives to artificial intelligence can obscure responsibility. Here I want to pursue the practical consequence of that argument: how do we investigate a failure without turning it into either a publicity exercise for the manufacturer or a supernatural event for the audience?
A credible defence of artificial intelligence must withstand an examination of its failures. Enthusiasm does not exempt a technology from scrutiny. Fear does not exempt its critics from evidence.
When language acquires consequences
An answer and an action carry different responsibilities.
An assistant may recommend an incorrect change. An agent connected to tools may carry it out. Depending on its design and permissions, it can modify documents, interact with software and use feedback from those operations to determine its next step. Anthropic’s engineering guidance distinguishes predefined workflows from agents that dynamically direct their own processes and tool use. That distinction concerns how a system operates; it does not settle questions about subjective experience. anthropic.com
The practical boundary becomes clear with an ordinary request: identify duplicate records.
Identifying them, recommending their removal and deleting them are separate operations. A system that proceeds from identification to deletion has exceeded the task, even if every record it removed was a duplicate.
A useful outcome cannot retrospectively manufacture permission.
We understand this when dealing with people. An accountant invited to examine our finances does not thereby acquire authority to transfer our savings. Introducing a conversational interface should not cause this elementary distinction to collapse.
The first question is therefore straightforward: what action was authorised?
Preserve the instruction before explaining the failure
An investigation should begin with the original request and the restrictions attached to it.
Was the system asked to inspect, propose, prepare or execute? Did the user explicitly reserve approval? Did later instructions change the scope? Was there a conflict between the task and some other instruction available to the system?
The wording matters.
“Clean this up” may refer to formatting, organisation or removal. Where different interpretations carry different consequences, a dependable process must resolve the ambiguity before acting.
An explicit prohibition presents a different problem. If the user says “do not send”, an external message is a failure regardless of how helpful the system considered it.
The investigation must preserve that distinction. Otherwise, the language of ambiguity becomes an all-purpose solvent for accountability.
Post-incident summaries deserve particular caution. People tend to describe their earlier decisions in the light of what happened afterwards. Organisations can turn this habit into a department.
Preserve the instructions before everyone becomes remarkably certain about what they meant.
Examine the permissions
The next question is what the system could do.
Could it read, modify, delete or transmit information? Did it need all those capabilities? Were its permissions limited to the relevant files and services? Could it act on a live environment during an exploratory task?
OWASP identifies excessive functionality, excessive permissions and excessive autonomy as underlying causes of “excessive agency”. Its recommended controls include limiting available tools and permissions and requiring human approval for consequential actions. The concern is the authority through which an unreliable output can become a damaging operation. OWASP Gen AI Security Project
There is an important engineering distinction here. An instruction asking a model to avoid an action is a behavioural constraint. A software permission preventing that action is an enforced boundary.
Neither makes the entire system infallible. They operate at different levels, and their failures should be examined separately.
A company cannot provide extensive access, rely on a conversational prohibition and then express astonishment when that arrangement proves inadequate.
The machine may be new. The administrative negligence is thoroughly familiar.
Reconstruct the action, not the story
An incident needs a timeline grounded in observable events.
What material entered the system? Which tools were invoked? What parameters were supplied? What responses came back? What changed outside the conversation?
This is especially important when an agent reads material from an external source. A document, email or webpage may contain instructions designed to redirect its behaviour. OWASP describes this as indirect prompt injection when the malicious instructions reach the model through outside content. Its guidance recommends separating untrusted material from instructions and combining safeguards rather than relying on a single filter. OWASP Cheat Sheet Series
A document being reviewed should not acquire the authority of the person who commissioned the review merely because it contains an imperative sentence.
If that boundary fails, investigators need to establish how it failed. Calling the episode a rebellion supplies no useful diagnosis.
The agent’s own description also requires verification. “I saved the file” is a claim. The file’s existence, location and contents are evidence.
Likewise, “I checked the result” does not establish that a meaningful check occurred. Inspect the recorded operation and the external state.
A fluent explanation can accompany an incorrect action. Certainty is among the least expensive things a language interface can produce.
Distinguish a finding from an interpretation
A failed task can establish several things. It may reveal a misunderstanding, an unreliable tool interaction, an authorisation failure or a misleading report.
Those findings matter in their own right.
They do not automatically establish why the behaviour occurred, how often it will recur or whether it generalises to other systems and environments.
A single demonstration can expose a serious vulnerability. It cannot, by itself, establish its prevalence. Conversely, a hundred successful demonstrations cannot erase a failure that appears under a consequential but neglected condition.
The investigator should ask which conditions produced the behaviour, whether it can be reproduced and what changes alter the outcome.
Controlled evaluations also need their context. A test deliberately constructed to expose a weakness can be valuable without representing an ordinary user session. Removing the setup from the finding may produce a more dramatic story and a less informative one.
We should resist both convenient conclusions: that a successful product cannot have a serious defect, and that a serious defect makes every use of the product indefensible.
Neither position requires much investigation. That is part of its appeal.
Give the human an actual decision
“Human oversight” can describe a functioning safeguard. It can also describe a person sitting near a process they barely understand and cannot meaningfully interrupt.
A reviewer needs to see the proposed action, its scope and its likely consequences. They need time to examine it and the authority to refuse.
A button labelled “continue” offers little reassurance if the user cannot tell whether continuing means generating a draft or sending it.
There is a psychological problem here as well. A polished interface can make an operation feel more settled than it is. Repeated approval requests can become routine. Under pressure to work quickly, a person may begin treating review as an obstacle between the task and its completion.
These are reasons to examine the design of supervision. Blaming the individual reviewer without examining the conditions of review leaves the underlying process conveniently untouched.
If an organisation rewards rapid approval and penalises delay, it should not later present hesitation as the employee’s obvious duty.
Responsibility must be allocated before the failure: who defines the task, grants access, approves consequential actions and checks the result?
Follow institutional incentives too
The technical account is necessary. The institutional account completes it.
Who chose the deployment conditions? Who benefited from greater autonomy? Who accepted the remaining uncertainty? Who could have required additional testing?
A vendor may emphasise the user’s responsibility. An employer may emphasise the vendor’s assurances. A manager may describe the deployment as an unavoidable response to competition.
Each explanation may contain some truth. None should be accepted as a substitute for examining the decisions.
The politically useful phrase “the AI decided” can obscure a long sequence of procurement, configuration and supervision choices. It gives an organisation an actor to blame without giving the public a person who can answer questions.
Accountability should follow the actual distribution of control. The model’s limitations belong in that account, alongside the decisions that exposed others to them.
Demonstrate the repair
A convincing response to failure must identify what changed and how the correction was tested.
Perhaps the tools were restricted. Perhaps an approval step was introduced. Perhaps the system’s instructions, model or operating environment required revision.
Whatever the intervention, test it against the conditions that exposed the defect and against relevant variations. Check whether it reduces the original failure while creating others.
NIST’s Generative AI Profile treats risk management as work spanning the system’s lifecycle, including evaluation and incident-related processes. That is a more useful approach than treating a successful demonstration as permanent evidence of reliability. NIST
An apology in a chat window does not establish that any underlying control has changed.
Software can express regret with excellent manners. The permissions may remain exactly as they were.
A defence of AI worth making
I want artificial intelligence to become more capable, more accessible and more useful. That ambition requires an exacting account of performance.
Which task does the system perform well? Under what conditions? With what supervision? What happens when it fails? Can the affected person detect the error, interrupt the process and obtain a correction?
These questions help distinguish valuable applications from unsuitable deployments. They also give criticism a practical purpose.
We should assess AI against the actual task and the available alternatives, including the limitations of existing human processes. Demanding perfection from one option while overlooking the failures of another produces a distorted comparison. Accepting avoidable harm because the technology is impressive produces a different distortion.
The standard should be evidence, proportionate to the consequences.
When an agent fails, preserve the record. Establish the authorised task. Examine the access, the actions and the external result. Identify the failed boundary and test the correction.
There is enough intellectual work in that sequence to occupy us without inventing a personality for the server.
The machine may have produced the error. We remain responsible for the quality of the explanation.




Deja un comentario