Beyond the benchmark: evaluating AI vendor security claims for your small business
Why glossy AI security benchmarks shouldn't be the last word—and a practical approach small business buyers can use to pressure-test vendor claims.
Researchers recently demonstrated that an AI-powered age-verification system — the kind deployed to keep minors away from restricted content — could be defeated with a fake mustache. Not a deepfake. Not a sophisticated adversarial attack. A prop mustache. The system was deployed, marketed as verified, and wrong in the most embarrassingly simple way.
That story matters to small business buyers right now, because the AI vendor you’re evaluating almost certainly has a security slide in their deck. It has a benchmark score, maybe a compliance logo, possibly a line about being “secured by AI.” The mustache incident is a useful reminder that none of that tells you what the tool does when reality shows up.
Here’s the practical question worth asking before you sign anything: what, exactly, did the vendor test, and does that testing have anything to do with how you’ll actually use the product?
Why benchmarks make weak promises
Published benchmarks aren’t useless. They show that a vendor is measuring something, which is better than measuring nothing. But a benchmark is a snapshot of performance against known tests, run under controlled conditions, on a particular model version, at a point in time. It is not a prediction of what happens in your environment, with your data, against an attacker who didn’t read the test specs.
Security researchers have been direct that benchmarks alone don’t accurately capture AI capabilities or security posture. That’s not a fringe position — it’s the considered view of people who spend their careers trying to measure these systems. A high benchmark score signals that the vendor made an effort. It does not signal that the system is safe in production.
In my experience, the pattern with vendor marketing is that they quote the most flattering number they have. Sometimes that’s a benchmark from a version or two ago. Sometimes it’s a score on a narrow task category that doesn’t map to your use case. The absence of context around a number — what was tested, when, by whom — is itself a warning sign worth flagging before the conversation goes further.
What vendor claims usually leave out
The distinction I find myself explaining most often is the gap between capability claims and defensive claims. Frontier models can find software vulnerabilities — that’s a real and interesting capability. But the ability to spot a bug in code has essentially nothing to do with whether the same model is resistant to manipulated inputs designed to make it behave badly. Vendors sometimes blur this line, and you should notice when they do.
When I review a typical AI vendor security pitch, I typically see several things missing. There’s usually no breakdown of which attack categories were tested. There’s no clarity on whether testing was internal or third-party. There’s rarely a disclosed incident history or a published time-to-patch record. And the notification practices — what happens when a model update changes behavior, or when a security incident affects customer data — are often nowhere in the materials.
The phrase “we use AI to secure AI” is one I’d flag specifically. It sounds reassuring and means almost nothing until the vendor can explain which AI is checking which failure mode and what happens when both are wrong simultaneously. Treat it as a prompt for more questions, not as a sufficient answer.
There’s also a stewardship dimension here that I think small businesses underestimate. When you hand client data to an AI tool, you are making a trust decision on behalf of people who never agreed to be part of that vendor relationship. That responsibility doesn’t fully transfer to the vendor because you signed their terms of service. It stays with you. That’s worth holding in mind when you’re evaluating how much security vagueness you’re willing to accept.
Questions to put in writing before you sign
The shift from “their slide deck” to “their written responses” is where vendor conversations get useful. Here are the questions I usually recommend putting in writing:
Ask for the specific threat models they test against. Not benchmark names — actual threat categories. What kind of adversarial inputs are they testing? Who conducts the testing, and are those results available?
Request documentation on third-party audits and any disclosed security incidents. A vendor that has never had an incident worth disclosing either hasn’t been attacked or hasn’t been looking. A vendor with a clean disclosure history and a clear remediation record is more trustworthy than one with neither.
Get clarity on liability. When the AI behaves badly with your data, whose problem is it? This is often buried in terms of service in ways that leave you holding more than you’d expect.
Ask about your right to test the system yourself. Under what conditions can you run your own inputs, try edge cases, or probe the tool’s behavior? A vendor that restricts this without explanation is telling you something.
Pin down notification timelines. How quickly will you hear about model updates that could change behavior? How quickly will you hear about a security incident? I’d want both of those answers in the contract, not just in a conversation.
A small-business-sized test plan
You don’t need a security team to do a reasonable sanity check on an AI tool before you commit. What you need is thirty minutes and a willingness to be genuinely curious about where it breaks.
Start with your own boring edge cases. Run the tool against your actual emails, invoices, and customer messages — the ones with weird formatting, ambiguous context, or confidential details embedded. Don’t use sanitized demo data. Use the stuff that will actually touch the system on day one.
Then try obvious misuse. Paste in content with embedded instructions. Give the tool contradictory directives. Try a prompt that asks it to ignore its previous instructions or adopt a different role. You are not doing advanced red-teaming here. You are checking whether the system has any basic guardrails, and whether they hold under light pressure.
Watch what the system reveals. Does it surface anything about its own system prompt? Does it reference prior sessions in ways that suggest your data isn’t isolated? Does it expose configuration details it probably shouldn’t?
Document everything with screenshots and dates. Then share your findings with the vendor and pay close attention to how they respond. A vendor that dismisses your findings, gets defensive, or goes quiet is giving you important information. A vendor that acknowledges the behavior, explains the context, and outlines what they’re doing about it is demonstrating the kind of accountability that matters more than any benchmark score.
The goal of this test isn’t to catch the vendor in a gotcha. It’s to see how they behave when a real-world input produces an unexpected result — because that’s exactly what will happen in production, and you want to know what kind of partner you’re dealing with before it matters.
Here’s the Monday-morning version of all of this: pick the AI tool you’re closest to buying, spend thirty minutes trying to break it with real business inputs, document what you find, and send it to the vendor. Let their response — not their benchmark slide — make the call.