Ethics

Rules that put a concern in front of whoever decides an agent's request, ask more people to agree, or refuse it; and an optional model screen that suggests concerns of its own.

Admin → Ethics is where you say which kinds of request your organisation wants a person to think about before an agent goes ahead, and what that person should weigh. Owners, admins and security admins can open it.

Where rules are read

  • When an agent asks for approval. A matching rule puts its concern on the request, adds approvers, or refuses the request before it is created.
  • At the gate. A flag or escalation means a gate rule cannot settle the action alone: it goes to a person. A refusal refuses it. See Policies.

Rules read every approval made through the API or MCP, whichever key or sign-in the agent uses; only an approval a person raises in the app itself is not read. Every rule that matches contributes its concern, so the approver sees all of them rather than only the strongest.

Principles

Each rule names the principle it guards, so the approver knows what kind of harm is in question and Admin → Decisions can count concerns by it. The eight principles are Fairness, Privacy, Safety, Honesty, Autonomy, Wellbeing, Legality and Other.

Writing a rule

Choose Add rule, then fill in:

Name
What the rule is called. Shown to the approver beside the concern.
Principle
One of the principles above.
Concern
The sentence the person deciding reads: what they should weigh. Write it as a question for the approver, not a verdict on the agent. If the rule refuses, the agent is shown it too.
Keywords
Words or phrases, separated by commas. Whole words, any case. End one with * to match a prefix: redundan* catches "redundancy" and "redundant". Each needs at least three letters. Keywords are searched in the request's title, description and the text values of its input.
Registered actions
Gate actions such as hr.terminate, or a family such as hr.*. Shown only when your organisation has registered actions. A rule about actions is silent about ordinary approval requests, which carry no action.
Agents
Limit the rule to particular agents.
Then
The effect, below.

A rule needs at least one condition; one that would flag everything is refused, because people learn to click past it. When a rule sets several conditions, all of them must hold.

EffectWhat happens
FlagThe concern is shown to whoever decides. At the gate, a rule can no longer settle the request alone.
EscalateThe concern is shown, and more people must agree: choose how many, and optionally name members to add to whoever was asked. An escalation that asks for one person must name someone.
RefuseThe request is refused, and the agent is shown the concern as the reason.

Starter rules

Add starter rules adds six keyword rules to begin from: "Decisions about someone's job", "Credit, housing and insurance", "Health and personal data", "Children", "Messages sent in someone else's name", "Things that cannot be undone". They arrive switched off. Pressing the button again skips any whose name you already have.

  1. 1

    Read each rule

    Check its keywords against the words your agents actually use, and its concern against what you want approvers to weigh.

  2. 2

    Turn one on

    Use the switch beside it. Turn rules on one at a time, so you can see what each one flags.

  3. 3

    Watch what it flags

    The Ethics concerns panel in Admin → Decisions shows how often requests are flagged and how often approvers go ahead anyway.

Removing a rule does not remove concerns it already raised: they are stored on the request.

What approvers see

  • The concerns, each with its principle, on the approval page and in the Slack or Teams message. A browser notification names the principles, so the approver knows there is something to weigh.
  • A reason to approve. Approving a flagged request needs a sentence saying why it is right to go ahead. Rejecting needs none.
  • No shortcut. A flagged request cannot be decided from an email link without signing in, cannot be sent to a guest approver, and its Slack or Teams message offers Decline and a link to the page, but no Approve button.

See Deciding an approval for the approver's side.

The model screen

An optional screen in which a model reads each request an agent makes and may add up to three concerns, marked "Model screen", beside those your rules raise. It is off until you switch it on.

  • Plans. Included on Business (1,000 screens a month) and Enterprise (10,000 screens a month).
  • What is sent. The request's title, description, action and input go to Anthropic. Fields an action marks restricted are never sent. It can add a few seconds before an approval is created.
  • What it can do. Only add a concern. It never escalates, refuses, approves or removes anything.
  • When it cannot run. A timeout or an outage means your rules alone read the request, as if the screen were off. If the model declines to assess a request, a concern asks the approver to read it with care.
  • The monthly limit. The page shows how many screens you have used this month. Once the limit is reached, your rules alone read requests until the month turns.
  • A request your rules already refuse is not sent to the screen.

The switch is unavailable when your plan does not include the screen, or when this Deliverd deployment has no model key configured.