PlatformDelivery BoardAutomation & RunbooksOutcome ReceiptsJust-in-Time ElevationCredential VaultGoverned SessionsDevices & DiscoveryPatch ManagementReporting & ExportsRoles & Multi-Tenancy
Verified AI OperationsThe Operation LoopCommanded AutonomyCompare the operating modelThe Verified Operation Spec
SolutionsFor MSPsFor Enterprise & Internal ITHealthcareLegalFinancial servicesMunicipal & Education
ProofOperation walkthroughSecurity & architectureVerified Operation SpecFive questions for your RMM's AIChangelog
CompanyAboutFounder's noteContact
PricingBuy 1–20 technician licenses onlinePlans — from $499 per monthCustom requirementsFoundation Circle
Log in

Home/Proof/Five questions for your RMM's AI

Bring these to any demo

Ask how the work happens. Then inspect the evidence.

If you are evaluating AI for support work, these five questions take about ninety seconds each and have factual answers. Ask them of every vendor, including us, then inspect whether the answer names the identity, scope, human decision and target evidence.

a vendor demo · eleven minutes in
The moment that decides it
Y
You
When the AI restarted that service, what account did it use?
V
The vendor
It runs under our secure agent — encrypted at rest, and every action is logged.
three true statementsnone of them the answer
Y
You
Does that account exist right now, and can it do that on every machine?
the follow-up that ends it
illustrative — not a transcriptThe written test →

What you are listening for

A dodge is rarely a lie. That is what makes it work.

The answers you will get are usually true. They are just answers to adjacent questions — storage instead of delivery, logging instead of verification, policy instead of identity. The skill is hearing the swap in real time, and having one follow-up ready that cannot be answered sideways.

What a dodge sounds like, and the follow-up that ends itillustration
A DEMO EXCHANGE, DECODED QUESTION 1 · THE CREDENTIAL YOU “What account did that action run as?” WHY ASK IT THIS WAY Not “is it secure”, which invites a posture. “Which identity” has exactly one factual answer. THE VENDOR “It runs under our secure agent. Everything is encrypted at rest and every action is logged.” THREE TRUE STATEMENTS, NO ANSWER Encryption describes storage. Logging describes afterwards. Neither says which identity held the rights at the moment the service was restarted. YOU · THE FOLLOW-UP “Does that account exist right now? And can it do that on every machine, or only the one we fixed?” WHY THIS ONE ENDS IT “Right now” and “every machine” are the two phrases that convert a posture into a yes or a no. There is no third answer. THE VENDOR “…yes. It's a service account with local admin.” WHAT YOU JUST LEARNED A permanent administrator on every endpoint, held by software. That is the finding. It is not a bug, it is the design. ILLUSTRATIVE — THE POINT IS THE SHAPE OF THE ANSWER, NOT ANY PARTICULAR VENDOR
Nothing the vendor said here was untrue, and a standing service account is a perfectly normal way to build an RMM. The finding is not that they were caught out — it is that you now know the ceiling on every other claim in the demo, because nothing above that account can be tighter than the account itself.

The five

Ask them in this order. The order does work.

Identity first, because everything else is measured against it. Then what the model holds, then who says it worked, then what happens when it does not, then who is named on it.

  1. “When the AI does something, what account does it use?”

    A good answer names an identity and a duration in the same breath: a grant that was minted for this operation on this target and expires on a clock. A dodge reaches for adjacent virtues — the agent is secure, the traffic is encrypted, the actions are logged. All of that can be true while a permanent administrator sits underneath it, and if one does, it is the ceiling on everything else you are shown.

    FOLLOW-UP · “Does that account exist right now, and can it do that on every machine?” — see Just-in-Time Elevation

  2. “Does the model ever see the password?”

    A good answer is a boundary description: the model gets a reference it cannot resolve, something else resolves it, and the connection is authenticated before any generated script exists. A dodge is about hygiene after the fact — secrets are redacted from logs, masked in the transcript, scrubbed from telemetry. Every one of those describes cleaning up a value that has already been somewhere it should not have been.

    FOLLOW-UP · “So could I recover the plaintext out of your prompt history?” — see Credential Vault

  3. “Who told you it worked?”

    A good answer points at the machine: after the change, an independent check ran on the target and its answer is the field in the record — so a failure writes itself in whether or not anyone would have chosen to report it. A dodge points at the screen. The status went green, the job returned zero, the assistant summarised the outcome. Ask them to break the fix again on the call and re-open the same record; a platform that only reports success has nothing to show you at that moment.

    FOLLOW-UP · “Which line here came from the target, and which line came from you?” — see Outcome Receipts

  4. “Show me the undo — before you run it.”

    This one is a request rather than a question, and that is deliberate. A good answer puts the reversal on screen before the change happens, because it is part of the procedure and a condition of the procedure being allowed to run at all. A dodge produces the undo afterwards, or produces a backup. A reversal written after the fact is a report about what happened; it was not a plan, and nobody committed to it in advance.

    FOLLOW-UP · “Run it, then undo it, on this call.” — see Automation & Runbooks

  5. “Whose name is on this, and is it in the same record as the work?”

    A good answer is one reference number that resolves to the person who authorized it, the privilege it used, the job that ran, and the machine's reply. A dodge is two systems: approval lives in the ticket, execution lives in the automation platform, and they are correlated by timestamp. That is two stories that agree, which is a fine operational practice and is not a chain of custody. And if the honest answer is that nobody's name is on it, you have found the thing your insurer is going to ask about.

    FOLLOW-UP · “Give me one number I can quote back to you next quarter.” — see Commanded Autonomy

Say this out loud at the start“I am going to ask five questions and write the answers down.” It is not adversarial and it changes the call — the person demoing will usually route you to an engineer, which is exactly where you wanted to be. Vendors who build this way are pleased to be asked; that reaction is data too.

Grading what you get back

Three kinds of answer. Only one of them leaves the building with you.

After the call you have to convince someone who was not on it — a partner, a client, an underwriter, an internal auditor. That is the real test of an answer, and most answers fail it for a reason nobody notices during the demo.

What survives the callthe evidence ladder
WHAT YOU CAN STILL USE ON MONDAY TIER 1 · SAID A claim “It's fully audited, and the AI can't do anything a technician couldn't.” AFTER THE CALL YOU HAVE Your notes. Nothing anyone else can read. TIER 2 · SHOWN A screen A log view inside their console, scrolled past at demo speed. AFTER THE CALL YOU HAVE A screenshot, which proves their interface exists. TIER 3 · HANDED OVER An artifact An exported record carrying the target machine's own answer. AFTER THE CALL YOU HAVE Something a third party can check without them. EVERY QUESTION ON THIS PAGE IS REALLY ONE QUESTION: CAN I GET TO TIER THREE?
Tier two is where most demos stop, and it is genuinely persuasive in the room. The test is whether you can leave with the thing itself — ask for an export of the operation you just watched, sent to you after the call, and see whether the reply is a file or a follow-up meeting.
The one-line versionAsk for the artifact, not the answer. Every question here is a way of finding out whether an artifact exists at all — and what one looks like when it does.

Two instruments, two moments

This page is the ammunition. The spec is the standard.

These five — in the room

They test the sales answer: how a vendor responds under a little pressure, which adjacent question they substitute, and whether the person on the call can get you to an engineer.

Use them live. You will know inside two minutes whether the platform was built by people who expected to be asked.

The Verified Operation Spec — afterwards

It tests the architecture: six yes/no conditions that do not change with who is presenting, scored the same way for every vendor, us included.

Use it to compare a shortlist once the demos are over, and to write a requirement into an RFP that cannot be met with a paragraph.

Read the spec and take the scorecard →

One caution about using these against us. We answer all five the way this page says a good answer sounds, which is not evidence — it is what you would expect a page we wrote to say. Where the categories actually differ is a more useful read, and the fastest resolution is to make us run an operation while you watch and then ask for the export.

Same five, pointed at us

Ask us before you ask them. It is a better rehearsal.

Fifteen minutes: a real ticket resolved end to end, the record opened, the rollback executed. Bring the hardest question from your last client audit and we will answer it on the call, not after it.