The model does the work. The authority sits outside it.
11 October 2026 · Dave Reid · 11 min read
Satya Nadella has published a short essay on X, Models as Insider Risks in the Super Intelligence Era. It is worth reading in full. One of his sentences carries the whole argument:
In other words, we need to separate the supply of intelligence from the authority over it.
We agree, and not only as a matter of opinion: it is the design Deliverd is built on. This post takes his principles one at a time. For each, it says what Deliverd does today and, where the answer is not ours to give, says that too.
Why "insider" is the right word
Nadella starts from an uncomfortable position. We are deploying models with access to sensitive data and the ability to take mission-critical actions. Yet we cannot trace their behaviour to a code path the way we could with traditional software. A model provider's assurances do not move the responsibility off the organisation that deploys it. So he proposes treating frontier models, closed and open weight alike, as insider risks:
Not because they are necessarily malicious, but because any sufficiently capable actor with access to important systems can make mistakes or be compromised, and the architecture of containment and control must account for that.
The framing helps because it makes the problem familiar. Organisations already know how to work with a capable colleague whose thinking they cannot inspect:
- give them an identity;
- give them only the access the job needs;
- write down what they do;
- keep a way to stop them.
Nobody asks a new finance hire to explain their reasoning before letting them issue refunds. They ask for a second signature above a limit.
The control sits outside the model
The most important line in the essay is about where the controls live:
It's why the controls that govern what a model can access and what actions it can take must sit outside the model.
He traces this to a security principle from the 1970s: a program must not be able to bypass or tamper with the mechanisms that enforce its permissions. That is the reference monitor, from James Anderson's 1972 planning study for the US Air Force. It has three properties: it is always invoked, it cannot be tampered with, and it is small enough to verify.
For an agent working with Deliverd, that means four concrete things.
- An agent cannot write policy. The API refuses, with a sentence we wrote for exactly this case: "Policy is what decides whether an agent needs a person, so an agent able to write it decides that for itself." An administrator writes policy, against business actions the organisation has named, such as
finance.refundordeployment.production. Policy never refers to whichever tool happens to carry an action out. - An agent never decides. It can ask for an approval, answer an approver's question and withdraw its own request. A decision needs a person, signed in. Nobody, person or agent, approves their own request.
- No rule means a person decides. An action that no rule mentions goes to a person. Most permission systems default the other way, and for a new kind of actor that is the wrong way round.
- What a person approved cannot change under them. A gate's scope is fixed in the database from the moment the request is written: its destination, whether that destination is inside the organisation, its classification and the hash of the file involved. Each approver's vote also records a digest of exactly what they were shown.
None of this depends on the model behaving well, explaining its reasoning faithfully or telling the truth about itself. Something that is not the model checks it.
Identity and least privilege
In Deliverd every agent can have an identity of its own: its own name, its own keys and its own reach. That reach covers which workspaces it may publish into, which reports it may change and who it may share with. An agent identity can never make anything public; the service refuses that outright rather than hiding the option. So the record says that Finance Reporting Agent did something, not the person who happened to start it.
Above that sits authority. An organisation decides, for each of twelve kinds of action, whether an agent may do it alone, only with a person's approval, or not at all. The kinds run from read and draft to delete, financial action and customer decision. By default, looking and suggesting are the agent's own work. Anything that changes something, reaches someone or decides for them waits for a person. An agent's own setting can tighten the organisation's and never loosen it.
One boundary, said plainly. What a model can read in your other systems is your identity provider's job, not ours. We sit at the moment it wants to do something.
The model does not keep the evidence
Two of Nadella's principles belong together. Every meaningful model action must leave human-readable evidence that can be trusted. And:
No single model should control both a system's behavior and the evidence required to determine whether that behavior is aligned with the original intent.
In Deliverd, we keep the record, not the agent. Every gate decision is written down, including the ones nobody watched. Each records:
- who asked;
- which policy applied;
- who decided, and under what authority;
- what they were shown;
- what happened next.
"What happened next" is where honesty matters. When an agent reports what it did with a permission, the record says how much of that we checked:
| Receipt | What it means |
|---|---|
| Verified | Deliverd carried out the act itself, so we saw it happen. |
| Hash matched | The agent reported acting on exactly the file that was approved, and the hash it reported matches. |
| Reported | The agent's word, with nothing to check it against. |
This applies Nadella's point about not relying on the model to vouch for its own outcome: every receipt states how much was checked. Putting a green tick on all three would look better and be untrue for two of them.
There is one word of his we will not claim, though, and that is tamper-proof. Our audit log is insert-only at the database level, so it refuses an edit even from our own service key. The only deletions are retention pruning, and a legal hold stops even that. An evidence pack, in PDF, JSON or CSV, carries a digest that shows whether it has changed since it left us. That is tamper-resistant and tamper-evident. A database administrator with enough access can always do more than an application can stop, and we would rather say so than let a word promise more than the system does.
Let models check each other, not decide
You can and you should use models to adversarially test and verify each other. However, you can then end up with an opaque model inside an opaque orchestration layer, watched by another opaque model.
Deliverd has exactly one place where a model watches the agent, and we designed it with that warning in mind. An organisation can switch on an ethics screen, in which a second model reads each request an agent makes and may suggest a concern to the person deciding.
- It can only add a flag. It never approves, refuses, escalates or removes anything.
- If it cannot run, nothing changes. The organisation's own written rules still read the request.
- It is off by default. Switching it on sends the request's words to a model provider, so an administrator has to choose that knowingly.
A flagged request is harder to wave through. It needs a member who signs in, and approving it needs a written reason. A guest cannot approve it from an email link.
So the model makes the person look harder, while the authority stays with the person and with rules that are written down, readable and give the same answer every time.
The emergency brake
An authorized person should always be able to pause or shut down a model mid-task.
Deliverd gives you four ways to stop an agent.
- Disable the agent. One button in Admin → Agents. Its key is refused on its next request, so it cannot check a gate, use a permission or ask for anything else.
- Revoke a permission it has not used yet. Revoking needs a reason, the agent is told, and any later report of acting on that permission is refused.
- Rely on narrow permissions. A yes covers one destination and one file, expires, and is usually good for one use, so it cannot be stretched later into something nobody approved.
- Let serious incidents stop the work. Anyone, an agent included, can report an AI incident against a task or an AI system. While a high or critical incident is open, that task's governance gate stays closed. Only a person can resolve it, and they have to say how.
The limit is just as important, so here it is. We cannot reach into the system where your agent acts. The brake stops any agent that asks before it acts, which is the point, but an agent that never asks is not contained by us.
That is why another of Nadella's instructions matters: separate the model from the harness, and keep the controls outside both. The ask belongs in the harness, not in the prompt. When Microsoft Agent Framework or Foundry pauses a tool call for approval, our Python SDK's adapters put the request in front of a named person and answer the framework with their decision. The model does not get a vote on whether the question is asked.
When the system changes
Nadella wants the whole system tested continuously, including its failures, attacks and changes, and not only its successful tasks. We do not red-team your model, and we would be wary of anyone who says one product can. What we can do is notice when the system under governance changes.
Suppose you record a change to the model behind one of your AI systems: a new provider, a new model or a new version. Every governed task that uses that system gets a reassessment, which lists what changed and which controls to look at again. The evidence and approvals that the change makes stale stop counting, and the task's governance gate stays closed until a person completes the reassessment.
Refusals are recorded as carefully as approvals, so the record is something to test against: which rules fired, which flags were overridden and which requests expired with nobody answering.
Incidents
Nadella asks for timely disclosure to the people affected, and for lessons shared across the industry. The second is bigger than any one vendor. The first starts with a record. Deliverd's incident register holds:
- what happened, and against which task or system;
- how severe it was;
- who reported it;
- how it was resolved.
It sits next to the decisions and receipts around the incident, and the compliance evidence pack carries all of it. Telling the people affected is still your job; our part is that you have the facts when you do.
We try to hold ourselves to the same standard. When a security researcher showed us how an app could pass itself off as a well-known AI tool on our consent screen, we fixed it and wrote it up.
What is not ours
Two of Nadella's asks sit inside the model, and Deliverd does not touch them.
- Transparent reasoning. We do not see a model's chain of thought. We record what the agent said it intended to do before it acted, in words a person can read, and what the person decided. That is evidence about the action, not about the model's reasoning.
- Alignment. Nadella sets it aside as the hard problem, and so do we. Containment and governance are the engineering available while it stays unsolved.
In a Microsoft organisation
The essay comes from Microsoft's chief executive, so it is worth saying where Deliverd fits in Microsoft's stack.
- Teams: approvers decide in a direct message, with Approve and Decline.
- Copilot Studio and Foundry: agents connect to Deliverd as an MCP tool.
- Agent Framework and Foundry: tool approvals go to a named person through the Python SDK.
- Microsoft Entra ID: handles sign-in and directory sync.
The details are on Deliverd for Microsoft.
Trust the model least
Nadella ends with this:
The most trustworthy Super Intelligence system will not be the one with the model we trust most. It will be the one that enables us to trust the model the least.
That is the sentence we would put on the wall. Deliverd does not ask you to trust the model. It asks you to trust a rule you wrote and a person you named, and it keeps a record of which one decided.
Your agent does the work. The authority stays with you.
Deliverd is human-in-the-loop infrastructure for AI agents. The gate is documented in the developer docs.
Quotations are from Satya Nadella's post on X. Deliverd is independent of Microsoft, and nothing here suggests Microsoft endorses it.