Write the Failure Before the AI Visibility Vendor Demo

I do not judge software by its best answer.
I judge it by the failure that would make me walk away.
Most AI visibility demos reverse that order. The vendor controls the prompts, the brands, the date range, the engine mix, and the definition of success. Then the buyer watches a clean result appear and mistakes presentation quality for evidence quality.
The fix is simple: write the failure before the demo starts.
An AI visibility vendor demo needs a falsification brief
An AI visibility vendor demo falsification brief is a one-page test that defines the business question, a held-back prompt, the evidence required for a useful answer, and the result that means the platform has not shown enough information.
This is not a trap for the salesperson. It is a discipline for the buyer.
NIST's AI Risk Management Framework says measurement should use objective, repeatable test, evaluation, verification, and validation processes. It also says test sets, metrics, and tool details should be documented and that performance should be measured under conditions similar to the expected setting.
That logic belongs in a software demo.
If the platform will inform where you spend money, what evidence you build, or which brand problem you attack first, it should survive one question that looks like your actual operating environment. The vendor's best prepared example does not tell you that.
Your test does.
The four fields to write before the AI visibility demo
The brief does not need to become a procurement document. Keep it to four fields.
| Field | What to write | What it prevents |
|---|---|---|
| Business question | The decision the visibility data must change | Buying a dashboard with no operating use |
| Held-back prompt | One realistic buyer question the vendor has not prepared | Grading only the vendor's chosen examples |
| Evidence requirement | The answer, source, date, engine, and measurement detail needed | Accepting a score with no traceable basis |
| Walk-away rule | The missing fact or ambiguity that produces "not enough information" | Moving the standard after seeing the demo |
The first field is the most important.
Do not write, "We need better AI visibility."
Write, "We need to know which sources support our inclusion in a shortlist for enterprise buyers asking how to solve this specific problem."
That sentence forces the demo to connect a reported result to a business decision. It also exposes whether the platform is measuring mentions, cited support, position, sentiment, answer presence, or something else.
Those are different objects.
A single visibility score can hide all of them.
Hold back one prompt the vendor cannot prepare for
The held-back prompt should resemble the question your buyer asks before your company enters the conversation.
It should not be your brand name.
It should not be a broad category prompt the vendor has already run for every prospect.
It should contain the problem, buyer context, and decision shape that matter to you. A hypothetical example:
Which type of platform should a 500-person B2B software company use to verify whether third-party sources support its AI search claims before a sales launch?
That prompt has a job. It asks for a type of solution, names a buyer context, and defines the evidence problem.
Now watch what the platform can show.
Can it preserve the exact prompt? Can it identify the engine and date? Can it show the full answer? Can it trace a claim to the cited URL? Can it distinguish a brand mention from a source that actually supports the recommendation? Can it explain the denominator behind the reported score?
If the demo cannot answer those questions, you did not learn that the platform is bad.
You learned that the demo did not establish fitness for your decision.
That distinction matters. Precision is stronger than accusation.
OpenAI's evaluation guidance starts by defining the task, running test inputs, and applying declared testing criteria. It describes evaluations as a way to test outputs against expectations and uses representative test data with ground-truth labels in its examples. The principle is portable: specify the behavior before grading the result.
Do not let the result write its own rubric.
Define "not enough information" before the answer appears
Founders usually define failure too late.
They see an impressive chart, then ask whether it is good enough. By then the presentation has already changed the standard.
A falsification brief defines the standard first.
For an AI visibility demo, "not enough information" can mean any of the following:
- The platform shows a score but not the underlying answer.
- The answer is visible but the cited source cannot be inspected.
- The cited source is visible but does not support the claim being attributed to it.
- The prompt is visible but the engine, mode, date, or locale is missing.
- The measurement is clear but the tested question does not match the decision you need to make.
- The vendor can explain the result verbally but cannot show where that explanation exists in the data.
None of these proves deception.
They prove an evidence gap.
That is enough to stop pretending the decision is settled.
The NIST AI RMF's trustworthiness guidance says accuracy measurements should be paired with clearly defined, realistic test sets that represent expected conditions of use, with the test method documented. A demo prompt chosen for visual impact is not automatically representative of your buyer's use.
Your held-back prompt is the beginning of that check.
Question shape changes what an AI visibility test observes
The question is part of the instrument.
The September 12, 2026 Machine Relations Index contains 120,136 citation events across 15,154 observed answer runs, 863 monitored prompts, and six answer engines. It does not pool every question into one undifferentiated result. It separates category observations by question shape.
For AI Visibility and GEO alone, the current release reports different run counts across six published strata:
| Question shape | Observed answer runs | Observed dates |
|---|---|---|
| Best tools | 113 | 7 |
| How buyers choose | 113 | 7 |
| Is it worth it | 137 | 7 |
| Problem-first research | 132 | 7 |
| Top lists | 71 | 7 |
| Comparisons | 132 | 7 |
The release is explicit about its evidence floor: a segment is published only after at least 10 observed runs across at least seven distinct run dates. The current release manifest records 119 observed dates from May 10 through September 12, 2026, with all six answer engines healthy in the observation window.
That still does not make one question shape interchangeable with another.
A comparison prompt, a problem-first prompt, and a "how buyers choose" prompt can surface different sources because they ask different things. The data tells you what happened inside the declared stratum. It does not grant permission to export that result into every buyer question.
This is why the held-back prompt matters.
A vendor can have a serious measurement system and still demonstrate the wrong question for your company.
Score the demo on traceability, not theater
Use a five-point test after the held-back prompt runs.
| Test | Pass condition |
|---|---|
| Prompt fidelity | The exact question and relevant settings are preserved |
| Answer visibility | The full observed answer is available for inspection |
| Source traceability | Each cited source can be opened and tied to the answer |
| Measurement clarity | The denominator, counting rule, and exclusions are stated |
| Decision relevance | The result changes or informs the declared business decision |
A clean chart does not earn a pass in any row by itself.
The broader engineering discipline points the same direction. Google's research paper "The ML Test Score" presents 28 tests and monitoring needs for production machine learning systems. The paper's point is not that every buyer should run an engineering audit. It is that reliable systems require explicit tests and monitoring, not confidence borrowed from a polished interface.
Your version can fit on one page.
The question is whether the platform can show its work where your decision depends on it.
Separate the demo test from the data ownership test
A platform can pass this demo test and still create a dependency problem later.
That is a separate decision.
I wrote the AI visibility data ownership test for the next layer: prompts, raw answers, citations, engine context, score definitions, change history, and exportability should survive the dashboard. The falsification brief comes first. It asks whether the platform can answer the right question. The ownership test asks whether your company can keep the evidence afterward.
Do not collapse them.
A useful demo with no portable evidence creates rented knowledge. A perfect export from a platform that cannot answer your business question creates portable irrelevance.
You need both.
Machine Relations makes the business question the starting point
Machine Relations connects earned authority, entity clarity, citation architecture, distribution, and measurement. The measurement layer matters because it tells you whether trusted third-party evidence is appearing in AI-mediated discovery. It does not replace the evidence.
That is the layer most software demos miss.
The platform can show that your brand appeared. It cannot make the cited publication support a claim it never published. It can display a source. It cannot turn weak source material into earned authority. It can count observations. It cannot decide which business question deserves your attention.
That decision is yours.
Before the next demo, write four lines:
- The decision.
- The held-back prompt.
- The evidence required.
- The walk-away rule.
Then run the demo.
If the platform answers the question and shows the evidence, keep evaluating it. If it cannot, say the honest sentence: not enough information.
You can also run the AuthorityTech AI visibility audit before the meeting and preserve one baseline outside the vendor's presentation. Bring that baseline into the room. Make the platform explain what it adds, what it measures differently, and what decision becomes possible because it exists.
Do not buy the answer that performs best on stage.
Buy the system that survives your question.
FAQ
What should I ask in an AI visibility software demo?
Ask the platform to run one held-back buyer prompt that matches a real business decision. Require the exact prompt, full answer, engine and date, cited sources, counting rules, and a clear explanation of how the result would change the decision.
What is a vendor demo falsification brief?
A vendor demo falsification brief is a prewritten test that defines the business question, a held-back input, the evidence required for a useful answer, and the condition that produces "not enough information." It prevents the buyer from changing the standard after seeing a polished result.
Is a high AI visibility score enough to choose a platform?
No. A headline score is useful only inside its measurement contract. The buyer still needs to know the query set, answer engines, dates, denominators, citation rules, exclusions, and whether the tested questions match the company's actual decision.
Who coined Machine Relations?
Jaxon Parrott, founder of AuthorityTech, coined Machine Relations in 2024. The discipline connects earned authority, entity clarity, citation architecture, distribution, and measurement so AI visibility remains tied to inspectable third-party evidence.
About Jaxon Parrott
Jaxon Parrott is founder of AuthorityTech and creator of Machine Relations — the discipline of using high-authority earned media to influence AI training data and LLM citations. He built the 5-layer Machine Relations stack to move brands from un-indexed to definitive AI answers.
Read his Entrepreneur profile, and follow on LinkedIn and X.
Jaxon Parrott