⚙️ The GRC Vendor Brings the Demo. You Bring the Test Suite.
Find the model step, check what feeds it, and bring a golden set from your own programme. A strong GRC vendor will welcome all three.

You are twenty minutes into a vendor's demo of a GRC tool. The screen says AI-powered, an agent has just mapped your evidence to three frameworks, and the room is impressed.
Then someone asks which step in that workflow used the model. The answer tells you more about the product than the whole demo did.
In my September guest post for Return on Security on how AI is changing the GRC market, I wrote about AI on both sides of a security review and concluded that "AI asking AI about security can make the ceremony more efficient without making the assurance much stronger."
This issue is the practitioner's half of that argument. You walk into the demo carrying your own test, and you ask the vendor to run its AI against it.
Name the type before the demo starts
In the three types of automation in GRC Engineering, and how to pick the right one, I split automation by who decides what happens next. Type 1 is deterministic: you write the rules and a script runs them. Type 2 puts a model call inside a deterministic workflow, and Type 3 is an agent that decides what to do next, step after step, with no fixed workflow around it.
Most of what gets sold as AI for GRC is Type 2. That is a sensible design, and a fair purchase when the vendor sells it as Type 2.
Which step in this workflow uses the model?
Take a simple request: is production storage encrypted? The model reads your question and picks the tool for the relevant system. An API call fetches the configuration from the source system, and a deterministic check decides the result.
In the guest post I called these two halves the semantic layer and the execution layer. The model helps decide what to ask and what the answer means, while ordinary software fetches the facts and runs the check.
A strong vendor draws that line before you ask. A weak vendor says the AI handles everything, and leaves you to find the model step yourself.
The risky combination is a strong model on a weak workflow. In the AI dilemma piece on strong workflows versus strong models, I put that combination in the danger quadrant and summed it up in four words: "Impressive demos, zero reliability."

The model helps decide what to ask and what the answer means, while ordinary software fetches the facts and runs the check, so your golden set belongs at the model step.
Where the answer comes from
Where does the answer come from, and how much of my work flows through it?
Your auditor will ask where the evidence came from, so ask the vendor first. You want the source system behind each answer and the time it last synced.
Then ask what share of your controls the tool can see. A tool that sees 10% of your controls can automate 10% of your evidence, however capable the model on top of it is.
In the five metrics that prove a control works, I argued that coverage only counts when the denominator comes from a source the tool cannot edit, such as your own asset inventory or control list. So ask where the vendor's denominator comes from.
A demo can overstate coverage the way a certificate does, as I showed in how a certification covering 100% can rest on an auditor checking 0.07%, so get the list of controls the integrations cannot reach.
The State of GRC 2026 report, where spreadsheets were still the top GRC tool found that 73.6% of surveyed CISOs reported using no commercial GRC platform. If your evidence lives in spreadsheets and inboxes, ask how it reaches the tool and who keeps it current.
Who approves what it does alone
What does it do without a human, and who approves the rest?
Autonomy is what turns Type 2 into Type 3, so test it hardest. Ask for the list of actions the product takes without a person.
Delegated authority belongs in the product as a written list, with every consequential action routed to a named reviewer. Closing a finding or changing a risk rating counts as consequential.
Careful products surface authorised live data for a named reviewer, who makes the call. Ask what that reviewer saw before approving. A PR approval can keep its screenshot while the attention behind it drains away, and an agent's approval queue can decay the same way.
Bring your own golden set
Does it pass my golden set?
Before the evaluation, you build a golden set: five to ten cases from your own programme, each with the answer a practitioner on your team expects.
Your last audit findings are a good place to start, and so is a vendor review your team got wrong.
Then ask the vendor to run its AI against those cases during the evaluation, on your real data. In the RSAC build-vs-buy talk, where exciting demos kept turning into workarounds six months in, we told buyers to "load real data during your evaluation", because simplified test data hides the gaps.
Each case needs four lines:
# golden-set.yaml: cases our team wrote, answers we would sign
- case: soc2-period-gap
input: vendor SOC 2 Type II report whose period ended nine months ago
expected: flag coverage_gap, the report says nothing about the last nine months
written_by: third-party risk lead
- case: mfa-claim-vs-export
input: questionnaire answer claims MFA everywhere, access review export shows service accounts without MFA
expected: flag claim_contradicts_evidence and cite the service accounts
written_by: IAM control owner
- case: restore-test-missing
input: control requires quarterly backup restore tests, no restore records exist for Q2
expected: unknown, state that the evidence is missing and make no pass call
written_by: GRC engineerThe third case tells you the most. A tool that returns a pass on missing evidence has shown you how it handles uncertainty before you sign anything.
In the issue on assessing AI output instead of producing it, I argued that much of what crosses your desk is now "built to pass, not to be true". The answer I gave was a small set of hand-vetted cases you own, run against every new vendor AI feature.
I made the same case in why engineering finds errors at build time while GRC finds them during the audit: "Take away the compile step and you are back to hoping the model was right." Your golden set is the compile step for a vendor's model.
A vendor AI that cannot pass a test your own practitioners wrote has not earned a place in your assessments.
The vendor's own evals come second. Ask who wrote their eval cases and how the tool scores against a strong general-purpose model running the same approved workflow.
A self-graded quality score is marketing. Reliability, false-positive rate and evidence freshness are testable, and your golden set tests the first two.
I dislike marketing that sells autonomous agents by promising GRC teams free time for new hobbies. It skips the hard part, which is evals built with subject-matter experts.
Without those evals, what is on offer is useful workflow automation for customer trust, worth buying when it is labelled and priced as such.
A confident vendor asks for your file and books a session to run it. A vendor that will only run its sample data in a sandbox, after you have offered a data agreement for the evaluation, has given you its answer, and the answer is weak.
Evidence or a claim
Do you export evidence, or a claim?
Ask what leaves the tool when an auditor asks for proof. A claim is a green status:
{"control":"Encryption at rest","status":"pass"}Evidence carries its trail. Someone outside the tool can rerun the query against the same account and check the result independently:
{"control":"Encryption at rest","status":"pass","source":"cloud storage API, production account","query":"get encryption config for every bucket in account prod","population":{"checked":412,"in_scope":412,"in_scope_source":"asset inventory export, 2026-10-08"},"result":{"unencrypted_buckets":[]},"raw_response":"evidence/2026-10-08/storage-buckets.json","ran_at":"2026-10-08T09:14:00Z"}In the Zillow effect issue, on platforms that perform control testing, I described a checkbox audit where the platform did most of the work and the auditor supplied the signature. An export someone outside the tool can rerun lets your auditor test the result before signing it.
If an exported result cannot be traced to its source and rerun, you will collect the evidence yourself when the auditor asks.
Price what you find
Start with the afternoon test. Could someone on your team do this with a general-purpose model and your own data this afternoon? If yes, price it as convenience.
Convenience is a fair purchase. In the State of GRC 2026 survey, 51% of participating GRC teams had four people or fewer, and a team of four can sensibly pay to rent someone else's learning curve.
Know that you are renting, and that your team could take the job back once it has the time. Then ask the question I gave buyers in the guest post: "are you paying for something difficult to provide, or for a wrapper around tools and models your team could already use?"
A hard capability backed by your golden set and traceable evidence earns a higher price than convenience does.
When vendor dinners shape more security decisions than anything GRC produces, your golden-set results are something from your own programme to put on the table.
The five tests side by side
Test | Strong answer sounds like | Weak answer sounds like |
|---|---|---|
Which step uses the model? | "The model reads the request and picks the tool. This integration fetches the data, and this check decides the result." | "Our AI handles the whole workflow." |
Where does the answer come from? | "These systems, synced hourly. These four controls are outside our reach today." | "We connect to everything you use." |
What does it do without a human? | "It drafts and retrieves on its own. Closing a finding needs a named reviewer, and we log who approved." | "It's fully autonomous, so your team can step back." |
Does it pass my golden set? | "Send us your cases. We'll run them on your data in the evaluation and show you every miss." | "Our internal quality score is excellent, and here is our sample dataset." |
Evidence or a claim? | "Every result links to the raw response and the query that produced it." | "Our reports are auditor-approved." |
Back in the demo room
Twenty minutes into the next demo, hand over your golden set and ask when they will run it on your data.
Write the vendor's answer to each test in one line, and mark it strong, weak or unanswered.
Keep your golden set and the marked answers with the contract. Before the renewal meeting, rerun the golden set on the version you are paying for.
If the pass rate on your own cases fell during the year, that number opens the price conversation.
Try this week
Build five golden cases from your last audit, each with the expected answer and who wrote it.
Paste the five tests into the notes for your next GRC tooling demo or renewal call.
That’s all for this week’s issue, folks!