Home/Proof/The Verified Operation Spec
The Verified Operation Spec · v1
Define the outcome. Know what counts as evidence.
Use this as a buyer's test for governed AI work, not a promise that a page can settle for you. Six questions, each with a factual answer you can check inside a ten-minute demo. Bring the sheet to us and to every other platform you evaluate.
Take this into the room
Score us here. Score everyone else in the room.
Six rows, three columns, nothing filled in — ours included. The difference is that every row in our column carries a link to the page that settles it, so you can score us in about two minutes without booking anything. Then take the same sheet to whoever else is on your shortlist.
| The test · and what counts as a pass | Seraph | ||
|---|---|---|---|
| 01Is privilege leased per operation, or standing?Pass = scoped to one operation and one target, revoked on a clock | see it | ||
| 02Can the model ever hold a credential value?Pass = the model gets a reference; the value never reaches it | see it | ||
| 03Who asserts the outcome — the tool, or the machine?Pass = the changed machine answers an independent check | see it | ||
| 04Was the rollback written down before execution?Pass = the undo exists before the run, not after it | see it | ||
| 05Is the approval in the same record as the work?Pass = one reference number resolves to both | see it | ||
| 06Can the record be edited without anything noticing?Pass = editing one record breaks every record after it | see it | ||
| Passes out of six A row only counts if they show you the artifact. A confident yes is not a tick. | — | — | — |
The six tests
Each one has a factual answer. That is the whole design.
No test here asks how a platform feels, how modern it is, or how many integrations it has. Each asks something a competent engineer on the vendor's side can answer with a yes, a no, or a screen.
Does the AI hold standing privilege, or is privilege leased per operation and revoked on a clock?
An AI running under a permanent administrator account is a permanent administrator, and everything else in the demo is downstream of that one fact. Ask them to show you the identity the action ran as, then ask whether that identity can do the same thing on a different machine right now. A pass looks like a grant scoped to one operation and one target, with an expiry that lands on the clock rather than on the work finishing. A dodge sounds like "it runs under our secure agent" — which names a product, not an identity.
Can the model retrieve a credential value at any point in the flow?
Redaction happens after the value has already been in the model's context. Masking happens after that. The only answer that survives contact with an auditor is that the value never reaches the model at all. Ask what the model receives when a step needs a password: a reference it cannot resolve, or the thing itself. A pass looks like a reference token crossing the boundary and a value that stops at it, with the connection authenticated before any generated script exists. A dodge sounds like "encrypted at rest and redacted from the logs" — a true statement about storage, offered in place of an answer about delivery.
Who asserts the outcome — the tool, or the machine that was changed?
"Job completed successfully" is the tool grading its own homework. The question is whether anything independent of the tool was asked, after the change, whether the thing is actually fixed. Ask them to break the fix again while you watch and re-open the same record. A pass looks like a fresh check executed on the target, whose answer is the field in the record — so a failure writes itself in without anyone deciding to report it. A dodge sounds like a status chip going green, or a summary paragraph written by the model about its own work.
Was the rollback declared before execution, or reconstructed afterwards?
Every platform can describe an undo once you ask for one. The test is whether it existed before the change did. Ask to see the reversal on screen before the run starts, and ask who wrote it — the person who certified the procedure, or the model, at the moment you asked. A pass looks like a reversal that is part of the procedure and a precondition of it being allowed to run at all. A dodge sounds like "you can always restore from backup", which is a different product answering a different question.
Is the human's authorization in the same record as the work?
Approval in a ticketing system and execution in an automation platform, correlated by timestamp, is not a chain of custody — it is two stories that agree. Ask for one reference number that resolves to both the person who said yes and the job that ran. A pass looks like a single operation id carrying the named human, the time, the privilege, the job, and the result. A dodge sounds like "it is all in the audit log", at which point the only useful reply is: which one.
Can that record be edited afterwards without anything noticing?
Evidence that an administrator can quietly correct is documentation, not evidence. Ask who can edit a record, who can delete one, and what visibly breaks when they do. A pass looks like records that reference each other, so changing one leaves every later record disagreeing with it, and a recipient can check that themselves without the vendor's help. A dodge sounds like "only administrators can modify the audit log" — which confirms the record can be modified.
Where the answers live
Every test is answered by a stage that already has to happen.
This is why the spec is architectural rather than a feature checklist: none of these six answers can be bolted on later. Each one is produced by a step in the operation, or it is not produced at all.
Our own column
How we answer, so you can check us against it.
Publishing a test you pass is easy and worth very little. What follows is only useful because each row names the mechanism and the page that explains it, so you can go after the mechanism rather than the answer.
| The test | What we do | Where to push on it |
|---|---|---|
| 1 · Privilege | Leased per operation, scoped to the operation and the target, revoked on expiry | Just-in-Time Elevation |
| 2 · The model | The model receives a reference; the executor resolves it and authenticates first | Credential Vault |
| 3 · The outcome | The target machine answers an independent check, and that answer is the field | Outcome Receipts |
| 4 · The undo | Declared in the procedure before it is eligible to run | Automation & Runbooks |
| 5 · The approval | Same record, one operation id, covering command and outcome | The operation loop |
| 6 · The record | Each record carries the fingerprint of the one before it | Security & architecture |
We answer yes to all six — and every one of them is written so you can check it on a machine you control rather than take the answer from a website. Two things are worth knowing as you weigh it. SOC 2 Type II is on our roadmap and not yet in hand; where we stand, and the evidence we supply in the meantime, is on the security page. And this spec covers the architecture of an operation — it says nothing about breadth of coverage, which is a separate and fair question to put to us.
Two different instruments
The spec tests the architecture. The five questions test the answer.
They are used at different moments and they catch different things. Bring both.
This page — the architecture
Vendor-neutral, written down, and scored the same way for everyone. It asks what a platform is: where privilege lives, what the model holds, who asserts the outcome. The answers do not change with who is presenting.
Use it to compare a shortlist after the demos, and to write a requirement into an RFP that a vendor cannot satisfy with a paragraph.
Five questions — the sales answer
Sharper, conversational, and built for the live call. It tests how a vendor answers: what a good reply sounds like, what a deflection sounds like, and the one follow-up that turns a posture into a yes or a no.
Use it while the demo is running, in the room, when you only get one shot at the question.
Five questions for your RMM's AI →Bring the sheet
Score us on the call, not from this page.
Fifteen minutes: a real ticket resolved end to end, the record opened, the rollback executed. Bring the hardest question from your last client audit and we will answer it on the call.