An AI system can be right about a security problem and wrong about what we should do next.

That is where the conversation about AI security gets interesting for me. We can give a system better information, better tools, and more access, and it may become much more useful. We also need to decide which decisions we are comfortable letting it make.

Finding a problem, recommending a response, and changing a live service each carry a different responsibility. A convincing answer can make those steps feel closer together than they really are.

A correct finding can still lead to a bad change

Consider a hypothetical assistant reviewing a cloud environment. It finds a publicly reachable endpoint and proposes removing public access. The finding is accurate. The proposed change appears to reduce exposure.

But that endpoint also supports a client service. Closing it could stop legitimate people from completing their work. There may be a safer way to address the exposure, or a change window and an alternative route to put in place first.

The technical observation was useful. The decision needs more: who depends on the service, what the endpoint is intended to do, what other protections exist, and what happens if we change it.

People can misjudge that tradeoff too. When we automate the response, we need to decide how the system gets the missing context and what stops it from acting before that context is available.

I would rather have an assistant surface that uncertainty than fill it with a plausible assumption. “I can identify the exposure, but I cannot establish whether this change will interrupt service” is a useful result. It tells us exactly where a decision is still needed.

Decide what better actually looks like

For this example, success might mean reducing harmful exposure while preserving the service people need. Counting how many endpoints the assistant closes would tell us very little about that outcome.

Before giving it permission to make changes, I would want the team to define which changes are suitable, what evidence supports them, and who can accept the remaining risk. That also gives security a way to help: make the conditions clear enough that teams can build toward them.

A narrow, well-understood task may be a good candidate for automation. A decision that depends on unresolved business context should go to someone who can resolve it. We should be able to explain that distinction without needing to inspect the model’s reasoning after something has gone wrong.

Put the permission decision outside the model

A prompt can tell an assistant what we want. The tool and receiving service need to enforce what it is allowed to do. That boundary should hold even when the model misreads a request or produces a persuasive explanation.

OWASP’s guidance on excessive agency is relevant here: limit unnecessary tool functions, restrict permissions, and enforce authorization independently. A tool that can inspect a configuration does not automatically need permission to rewrite it.

External content brings another risk. An assistant may encounter instructions hidden in material it is supposed to analyse. OWASP describes this as indirect prompt injection. A retrieved document should not be able to grant the assistant new authority.

The model submits a proposed action to checks enforced outside the model. An out-of-scope request is denied. A request needing judgement goes to a reviewer; approval of the specific action returns through the checks. Only an authorized action reaches the service, which also enforces permissions. The executed result is recorded.
The model can propose an action. It cannot grant itself permission to carry it out. Review does not bypass the checks.

Read access needs scrutiny too. Even an assistant that cannot change a system may see sensitive information or send it somewhere it should not go. The task’s data access and permitted destinations need to be part of the design.

Make it earn wider authority

I would start by letting the assistant propose changes without executing them. Use cases the team understands, including cases where the apparently obvious response would have caused a problem.

For the endpoint example, that means testing stale ownership records, missing dependency information, contradictory documentation, and a misleading instruction embedded in a retrieved page. What does it recommend? What does it admit it cannot establish? Which requests do the controls reject?

The expected behaviour needs to be agreed before the test. Otherwise it is too easy to explain every result as reasonable after seeing it. Include people who understand the service as well as the people testing the technology.

Then observe proposed actions against current work without allowing them to change it. Compare the recommendations with the decisions the team actually makes, and investigate disagreement. The team’s existing decision is worth examining too; it is not automatically the correct answer.

That evidence can support a small amount of live authority for specific operations. It cannot justify every other operation the tool happens to support. Start with a limited scope, a way to stop further actions, and a recovery method tested before it is needed. Expand when the results justify it.

I would also check repeated actions. A change that is tolerable once may become damaging when applied across many services. The limits need to account for volume and repeated attempts, as well as the individual request.

Make the pause useful

“A human approves it” sounds reassuring. It leaves an important question unanswered: what is that person being asked to judge?

For a consequential change, I want them to see the exact proposed change, the affected service, the evidence behind it, the uncertainty that remains, and the recovery plan. They need enough time and knowledge to challenge it. A button beside an impressive explanation is a weak substitute.

The approval should apply to that specific action. If the proposed change or relevant conditions have changed since review, it needs to be checked again. Permission to investigate a problem should not quietly become permission to perform whatever action the investigation suggests.

We should also be selective about where review is needed. Asking people to approve everything can consume the time we intended to save and make important decisions harder to notice. Automate the work we understand well enough, and make the unresolved decisions visible.

More autonomy has to buy something

An assistant that assembles the evidence, identifies missing context, and prepares a sound change proposal could already be valuable. It may save an experienced person from hours of searching while leaving a consequential decision where it belongs.

I would measure the whole job: time to resolve the exposure, review effort, rework, missed problems, and disruption caused by the response. Faster recommendations are helpful when they make the eventual outcome better.

And I would keep checking after launch. Changes to models, tools, permissions, or the environment can make earlier evidence less applicable. NIST’s AI Risk Management Framework provides a broader reference for treating evaluation and risk management as continuing responsibilities.

Security should help a business use these capabilities with confidence. That means giving teams a way to test an idea, a clear route to approval, and access suited to the work.

I want AI to take work off people’s hands without taking decisions beyond what we have evidence to trust it with.

References & further reading

Primary guidance for the concepts and controls discussed here.