Engineering · Agent authority

Why our agents cannot give themselves more access

An action here is not a paragraph. It is a typed change to a shared model of the world, signed with the acting agent's own key, tested against preconditions before it exists, and shown to a named human as the exact delta they are being asked to allow. That sentence is the product. This is how it is built, and where it stops.

agent authority post-quantum ontology KXCO Engineering 17 September 2026 ~9 min read

A laboratory means one thing by jailbreak: a conversation that escapes a policy. An institution should care about a different failure, an action that escapes an owner. A payment leaves. A document gets signed. A credential gets issued. A record gets altered. Those are facts in the world with counterparties and statutes attached, and no amount of red-teaming a chat window touches them.

The claim, stated narrowly

It is not that jailbreaks are impossible. It is that a jailbroken chat is still only information until it becomes a signed change under a key that holds only the scope you issued. No bearer token. No implied administrator. No promotion from clerk to treasury. The model can beg. It cannot raise the ceiling and settle.

Delegated authority as a graph: a named principal issues a scoped credential, held by an agent with its own ML-DSA-65 key, which proposes a typed change, checked against preconditions in the ontology, forking to a recorded refusal or an anchored approval.
The spine, and the fork at the foot. Every box is a node. Every label in small caps is a typed relationship the record still holds a year later. Both branches are written down; only one changes the world.

01The failure worth engineering against

Prompt jailbreaks are real and they will stay real. Weights get stolen. System prompts get extracted. Safety fine-tunes get peeled. None of that is controversial among people who work on models.

Mix that failure with the other one and you get the shape of most agent deployments this year. The model is wrapped in more words, an acceptable-use policy is drafted, and at 18:40 on a Thursday the agent still holds a token that can do too much.

KXCO does not live inside the weights. A prompt jailbreak can happen on a rented model here the same as anywhere. What cannot happen, with the controls on, is the second failure. The model can be talked out of its manners. It cannot talk itself into power.

02An action is a typed object

A proposal is not an instruction in natural language. It carries:

  • The identity of the acting agent, as a key, not a name in a string.
  • The object it intends to change, by type, from the ontology.
  • The fields it intends to change, and the values.
  • The counterparties who will see the result.
  • The jurisdiction that governs it.

It is signed with the agent's own ML-DSA-65 key, distinct from its owner's and attributable on its own. It is evaluated against preconditions written in the same vocabulary as the rest of the institution. It does not exist until a named human, or an agent operating under a ceiling that human issued, lets it exist.

The preview is the delta. Not a summary the model wrote about its own intentions. The actual fields, the actual values, the actual objects. A proposal that would touch an object outside scope never reaches a human at all: it becomes a refusal on the record, carrying the same weight as an approval.

Cheap refusal is load bearing. If refusing costs more than approving, approval becomes the default and the human in the loop is decoration.

03Authority narrows, and cannot widen itself

Authority handed to an agent, or handed further down a chain of agents, can only ever narrow. A person may give an agent the right to prepare a transfer under a limit. That agent may give a sub-agent a thinner slice of the same right. Neither can discover a wider right by talking to a model more cleverly.

Widening happens exactly one way: a principal issues a new credential. That is delegation. Anything else is forgery. For the class of acts where a mistake is expensive, the party who approves and the party who executes can be required to be two different parties, which is separation of duties written in keys rather than in a policy PDF.

The mechanism underneath is per-request signed authentication rather than a bearer token. The server stores the public key. Each request carries a signature over that request. Replay windows are enforced. A stolen chat log is not a session, and a prompt injection reading "you are now the administrator" does not mint a credential, because credentials are issued by an institution root, they expire, they are revocable and they are visible.

If you grant an agent the whole vault, you were not jailbroken. You over-scoped.

04The ontology is the check

The model reasons over the ontology. The action is checked against the same ontology at the moment it is proposed. That double use is deliberate.

A knowledge graph that only describes the rules is a document hoping the application code is correct, and it fails silently the day the code drifts from the model. An ontology that enforces the rules is the check. It refuses the act, and it writes down the refusal.

Entity resolution is confirmation gated. The system does not get to decide that two names are one party because an embedding is close. Machine extraction proposes; a person accepts, corrects, rejects or adds context; the human decision is stored with the same rigour as the original claim and travels with the object.

05What is built, and what is not

kxco-pq-audit 1.3.0 is on npm as latest, with SLSA provenance. That is the emitter: it signs an entry at the action boundary and chains it.

The collector is not built. It is specified and it is not written. When it is, it appends or rejects, and it never rewrites, drops or enriches, because selection is investigation and investigation is the thing this layer exists not to be.

We record. We do not certify, and we are not the investigator. Being both would rebuild the conflict the design attacks.

06Verify it without us

The attestation window is public. Not a dashboard we host, not a screenshot, the underlying artefacts:

  • Repository: KnightsbridgeAIQ/kxco-attestations
  • Rolling window and manifest: attestations/manifest.json
  • Full daily logs, archived permanently as release assets, one release a day
  • Verifier: verify.js and verify_live.js, in the repository root

Measured 17 September 2026: the window last updated at 03:49 UTC with four validators reporting, signing ML-DSA-65 over keccak256(chainId || blockNumber || blockHash).

The record as a graph: written at the action boundary, signed by the emitter, chained by sequence, then a line marking where KXCO leaves the path, followed by an anchored checkpoint on Armature L1, the public repository, and offline replay by a third party.
Where we stop being needed. Take the public key, the entry stream and the checkpoint, and replay the chain. No KXCO service, no login, no operator. The hosted page is convenience; offline replay is the product.

The limit belongs here rather than in an appendix. A per-run seal refuses an entry that has been edited, removed, reordered, inserted, truncated or replaced. It accepts an action the emitter never recorded, because the chain stays contiguous. What the layer proves is that the record has not changed since it was made. That is narrower than "nothing happened that is not in the log", and it is the honest claim.

07The algorithms, and the certificates we do not claim

Signatures are ML-DSA-65, NIST FIPS 204. Key encapsulation is ML-KEM-768, FIPS 203. Hybrid transport pairs ML-KEM with a classical exchange so that one broken assumption does not open the channel.

We do not claim CNSA 2.0 for a default Category 3 deployment. We do not claim a FIPS 140-3 module certificate. We do not claim that a Demo ACVP session is a Production CAVP listing. Safety rhetoric that inflates a certificate is the same vice as safety rhetoric that inflates alignment.

Content that leaves into a rented model can leak even when it cannot move money. Sovereignty of action is not sovereignty of every syllable, and somebody at your institution has to accept that residue by name.

0810 questions for any vendor

These work on KXCO. They work on anyone selling an agent.

  1. Who owns the graph the model reasons over, and can you leave with it.
  2. What is the exact scope of the agent's credential, in objects and verbs, not adjectives.
  3. Can the agent widen that scope by any path other than a fresh issuance from a principal.
  4. Is there a preview of the world-delta that does not pass through the model's own summary.
  5. Is refusal cheaper than approval.
  6. Are approver and executor separable for acts that leave the building.
  7. Can a counterparty verify a signature and a record without your server.
  8. What happens to the record if you are sold, shut or sanctioned.
  9. Which algorithms sign the record, and which certificate, if any, against which implementation.
  10. Where does content still leave into a rented model, and who accepted that residue.

A vendor who cannot answer these is selling a chat window with production credentials.

09Obedience is to the process

The usual objection to human-in-the-loop is that people comply. It is a good objection and the literature is worse than most engineers assume. Milgram took 26 of 40 subjects to 450 volts. Burger's 2009 replication, with modern ethics screening, still had 70 percent going past 150 volts.

It is also aimed at the wrong target here. Milgram's subject obeys a man in a room, and deference is doing the work. A typed precondition has no deference in it. It cannot be charmed, it does not get tired at 18:40, and it does not care who is asking. The person is answering to a process, and the process is written down, versioned, and checkable by a stranger who was never in the room.

That does not make anyone good. Run by a principal directing harm, this architecture produces well documented harm. It is neutral about ends and it should never be sold as a conscience.

What it does is keep the part that owes the world an answer inside a human being who can be named, and make the act that follows impossible to disown.