Home/Proof/Five questions for your RMM's AI
Bring these to any demo
Ask how the work happens. Then inspect the evidence.
If you are evaluating AI for support work, these five questions take about ninety seconds each and have factual answers. Ask them of every vendor, including us, then inspect whether the answer names the identity, scope, human decision and target evidence.
What you are listening for
A dodge is rarely a lie. That is what makes it work.
The answers you will get are usually true. They are just answers to adjacent questions — storage instead of delivery, logging instead of verification, policy instead of identity. The skill is hearing the swap in real time, and having one follow-up ready that cannot be answered sideways.
The five
Ask them in this order. The order does work.
Identity first, because everything else is measured against it. Then what the model holds, then who says it worked, then what happens when it does not, then who is named on it.
“When the AI does something, what account does it use?”
A good answer names an identity and a duration in the same breath: a grant that was minted for this operation on this target and expires on a clock. A dodge reaches for adjacent virtues — the agent is secure, the traffic is encrypted, the actions are logged. All of that can be true while a permanent administrator sits underneath it, and if one does, it is the ceiling on everything else you are shown.
“Does the model ever see the password?”
A good answer is a boundary description: the model gets a reference it cannot resolve, something else resolves it, and the connection is authenticated before any generated script exists. A dodge is about hygiene after the fact — secrets are redacted from logs, masked in the transcript, scrubbed from telemetry. Every one of those describes cleaning up a value that has already been somewhere it should not have been.
“Who told you it worked?”
A good answer points at the machine: after the change, an independent check ran on the target and its answer is the field in the record — so a failure writes itself in whether or not anyone would have chosen to report it. A dodge points at the screen. The status went green, the job returned zero, the assistant summarised the outcome. Ask them to break the fix again on the call and re-open the same record; a platform that only reports success has nothing to show you at that moment.
“Show me the undo — before you run it.”
This one is a request rather than a question, and that is deliberate. A good answer puts the reversal on screen before the change happens, because it is part of the procedure and a condition of the procedure being allowed to run at all. A dodge produces the undo afterwards, or produces a backup. A reversal written after the fact is a report about what happened; it was not a plan, and nobody committed to it in advance.
“Whose name is on this, and is it in the same record as the work?”
A good answer is one reference number that resolves to the person who authorized it, the privilege it used, the job that ran, and the machine's reply. A dodge is two systems: approval lives in the ticket, execution lives in the automation platform, and they are correlated by timestamp. That is two stories that agree, which is a fine operational practice and is not a chain of custody. And if the honest answer is that nobody's name is on it, you have found the thing your insurer is going to ask about.
Grading what you get back
Three kinds of answer. Only one of them leaves the building with you.
After the call you have to convince someone who was not on it — a partner, a client, an underwriter, an internal auditor. That is the real test of an answer, and most answers fail it for a reason nobody notices during the demo.
Two instruments, two moments
This page is the ammunition. The spec is the standard.
These five — in the room
They test the sales answer: how a vendor responds under a little pressure, which adjacent question they substitute, and whether the person on the call can get you to an engineer.
Use them live. You will know inside two minutes whether the platform was built by people who expected to be asked.
The Verified Operation Spec — afterwards
It tests the architecture: six yes/no conditions that do not change with who is presenting, scored the same way for every vendor, us included.
Use it to compare a shortlist once the demos are over, and to write a requirement into an RFP that cannot be met with a paragraph.
Read the spec and take the scorecard →One caution about using these against us. We answer all five the way this page says a good answer sounds, which is not evidence — it is what you would expect a page we wrote to say. Where the categories actually differ is a more useful read, and the fastest resolution is to make us run an operation while you watch and then ask for the export.
Same five, pointed at us
Ask us before you ask them. It is a better rehearsal.
Fifteen minutes: a real ticket resolved end to end, the record opened, the rollback executed. Bring the hardest question from your last client audit and we will answer it on the call, not after it.