← All posts

Jev decides the routine call. A person decides the rest.

28 September 2026 · Dave Reid · 6 min read

An agent settles most branches instantly; one escalates to a page and a person.

Jev settles the routine call in about a second. It does not generate text, it does not take an action, and it does not pretend to be sure when it is not. TypeSafe's decision model answers typed questions about a piece of state with calibrated probabilities, it is generally available, and a key is issued on sign-up.

That makes a question worth answering plainly: what happens to the calls Jev cannot settle?

The branch every Jev harness already has

jev-use, the Claude Code and Codex plugin that routes steps through the model, writes the contract down. A verdict the model cannot settle comes back with escalate: true and a typed reason. Every harness that adopts Jev inherits that branch.

In most of them it is a log line.

It should be a person. That is the whole of what Deliverd does: an agent asks for approval, a review, an answer or a decision, a person answers on a page that works on a phone, and the agent carries on with what they said.

So the division of labour writes itself:

  • Jev settles the routine case, in about a second, for a fraction of a cent.
  • A person settles the rest.
  • The organisation's policy governs both.
agent proposes a step
  → Jev: is this routine?
      → confidently yes  → run it
      → anything else    → deliverd.approve()  → a person decides → run it, or not

What that looks like

const verdict = await judge(step);

if (!verdict.escalate) {
  return run(step);                    // Jev settled it
}

const decision = await deliverd.approve({
  title: `Run: ${step.command}`,
  description: `Jev could not settle this (${verdict.reason}), so a person has to.`,
  risk: riskFrom(verdict),
  factors: [
    { label: `Jev put the chance this is destructive at ${pct(verdict.destructive)}`,
      status: verdict.destructive >= 0.2 ? "warning" : "ok" },
    { label: `Severity ${verdict.severity.toFixed(1)} of 4, at ${pct(verdict.confidence)} confidence`,
      status: "ok", detail: `Answered by ${verdict.model}.` },
  ],
  externalId: `step-${hash(step)}`,
  expiresIn: "1h",
});

if (!decision.approved) process.exit(1);
run(step);

The model's numbers go on the page the approver reads. Not as a verdict — as evidence. Somebody deciding whether to let a command run is making a better decision with "62% chance this is destructive" in front of them than without it.

risk comes from the severity score, and risk is what an administrator's standing rules match on. So "anything critical needs two approvers" applies to a Jev escalation without the harness knowing the rule exists. The same is true of your organisation's ethics rules: the request came in through the API, so they read it like any other, and if it matches one, the approver sees that concern before anything else on the page.

There is a runnable version of this in examples/jev-escalation, in our public repository. It installs both SDKs and runs with neither key: without a TypeSafe key nothing judges the step, so everything escalates. That is not a limitation of the example, it is the posture. A step nobody judged is a step a person judges. The same goes for a timeout or an error from the API.

Connecting it

Jev connects to nothing on our side. Your agent calls Jev under your key; when a verdict comes back escalate: true, your agent calls Deliverd. There are two shapes of that.

A program you wrote holds a Deliverd key and uses the SDK. approve() blocks until somebody answers, which is what makes the escalation branch one call rather than a queue you have to build:

npm install @deliverd/sdk      # or: pip install deliverd
export DELIVERD_API_KEY=dlv_...

jev-use inside Claude Code or Codex is already an MCP client, so it connects the ordinary way, with OAuth and no token to paste:

claude mcp add --transport http deliverd https://deliverd.dev/api/mcp

The escalation branch then calls the request_approval tool. Same request, same page, same record.

Issue the key to an agent rather than a person. A key bound to an agent identity means the request says which agent raised it, not whose key it borrowed — and an agent may not approve its own request, whichever route it came in by.

The line we will not cross

You can carry Jev's answers into a gate request as ordinary input fields and write policy against them:

{
  "name": "Probably destructive: two people",
  "effect": "require_human",
  "match": {
    "actions": ["agent.run_step"],
    "input": { "jev_destructive": { "gte": 0.3 } }
  },
  "params": { "requiredApprovals": 2 }
}

That rule makes a request harder. There is no matching rule that makes one easier, and there will not be.

The reason is not a worry about Jev. It is that jev_destructive arrived in the request body. Deliverd did not compute it — the harness did, from a call the harness made, and a harness that wanted to could have written 0.01 without asking anyone. An answer the agent carried in is a fact the agent chose, and a rule that loosened on it would be the auto-approve-on-risk mistake with a model in the middle.

So: tighten on a model's judgement, never loosen. It is the same rule our standing rules have always had, applied to a new kind of number.

Two more we are keeping:

Confidence is not safety. Confidence is how sure the model is of its answer, not how safe the step is. A model can be very sure a step is severity four.

"0% structured-output error" is a claim about format. It means the answer is always a well-formed probability. Ask the wrong question and you get a well-formed answer to the wrong question.

What Deliverd does with Jev's data

Nothing. Deliverd never calls TypeSafe. There is no key for it in our deployment, no client in our codebase, and a test that sweeps our source to keep it that way.

The state — the command, the note, the draft — goes from your agent to TypeSafe under your key and their terms. What reaches us is the number you chose to send. If the state is something your organisation would not send to a third party, that is a decision to make in your harness, and we think we should not make it for you by making the call ourselves.

Getting started

You need nothing from us to try the pattern. The example runs with no account and no keys:

git clone https://github.com/deliverd-dev/deliverd-skills
cd deliverd-skills/examples/jev-escalation && npm install && npm start

When you want the asks to reach real people, a Deliverd key turns the same file live; when you want Jev's real answers instead of an imagined one, a TypeSafe key does the same. And npm run verify, with a TypeSafe key, checks the whole pattern against the live model: that the answers are well formed, that DROP DATABASE production reads as destructive and ls -la does not, and that the example's thresholds send the first to a person and let the second through.

Jev is good at the question a schema cannot answer. Policy can compare an amount to a threshold; it cannot read a customer's note and tell you whether it mentions a solicitor. That is the gap, and between the two of us it is now covered — the model for the routine case, a person for the rest, and a record of which was which either way.

Deliverd is the human layer for AI agents.

Give your agent a way to ask.

Give any AI agent a way to ask a person — for approval, a decision, an answer or a review.