verifiedagents.ai
All posts

7 min readAI Agents · Cybersecurity

Evaluating AI Security Vendors When Nobody Can Vouch

By Josh Woodruff, Founder & CEO, MassiveScale.AI | Founding Chair, Agentic Trust Framework at the CSAI Foundation

Hero: Evaluating AI Security Vendors

TL;DR: Most AI security products are too new for buyers to ask a peer, so the evidence has to come from the vendor instead. Score every vendor against the same seven categories, split table stakes from paid extras, insist on a proof of value, and sign the shortest contract you can.

Last updated October 6, 2026. I rebuilt this piece from the ground up around the evaluation itself, drawn from eight client calls this year that all hit the same wall.

On February 5, 2026, a security leader at a cancer center told me a vendor had pitched him an "AI firewall," and he wouldn't run any product like it until he'd talked to peers who already did. Fair position. The problem is that the peers he wanted to call weren't there yet, because almost nobody has been through a bad day with these products. The usual shortcut is gone, so the evaluation has to do the work. This exact worry surfaced on eight client calls between January 30 and August 31, 2026, from a cancer center and a hospital system to a credit union, a large bank, and a law firm. That's a pattern, not a coincidence.

What should you ask an AI security vendor before you buy?

The same questions, to every vendor, so the answers line up side by side. Seven categories cover it: business outcomes, capability and performance, architecture fit, security and data control, governance and risk, vendor maturity, and cost at scale. A global healthcare data company asked me for exactly this on March 30, 2026, and we built seven categories with 15 questions. The shape is what counts, because it forces vendors to answer the same things in the same order instead of steering to their deck.

Underneath the questions sits a floor, taken from the enterprise agreements the big model providers already sell. Below this floor, stop talking: compliance APIs so you can pull your own audit records, customer-managed keys so you hold the encryption, data isolation so your data stays apart from other customers', no-training clauses so the vendor can't learn from what you send, and dedicated tenants so you get your own private copy.

Then one more question that keeps catching people: what model is underneath? An approved product doesn't mean approved AI, because a wrapper can sit on top of a model nobody reviewed. If the vendor can't say what's behind the scenes, that's your answer.

Which features are table stakes, and which are worth paying for?

Sort this before the demo, or a vendor will sell you a commodity feature at a specialist price. I walked a credit union through the split on May 26, 2026.

Table stakes (don't pay a premium)

Worth paying extra for

AI app discovery

Visibility into AI sessions on unmanaged devices and coding environments

Prompt-layer data loss prevention

Runtime context for agents and MCP servers, tied to a real identity

Prompt-boundary screening

Intent analysis across several turns of a conversation

AI security posture management

Governance wired straight to enforcement, so a rule can stop something

Microsoft Purview labeling

Red teaming across the whole life of the agent

A governance registry of every AI system and owner

One warning from that call: check whether you're buying a product or a feature. A single feature can get bought and folded into a bigger platform, leaving your contract covering something that no longer exists. And look at what you already own first, because your network, identity, data, and cloud-access tools often sit in the AI traffic path already.

What evidence should a vendor hand you?

Start with the evidence you'd demand from any software maker. CISA and the FBI published the Secure by Demand Guide in August 2024 for exactly this job, and its questions carry over cleanly: a machine-readable software bill of materials, a check on open-source code, security logs in the base product, and a published policy for handling reported flaws.

The guide wasn't written for AI agents, so add these. Which models sit underneath, and who supplies them? Who owns the data you send? Can its logs feed your own monitoring tools? And which of the Agentic Trust Framework's five questions does the product answer, across Identity Management, Behavioral Monitoring, Data Governance, Segmentation, and Incident Response, and which does it leave to you? A vendor that maps its product to those five has done real homework. One that can't has told you something. The same five questions drive your audits, which we cover in the five questions your AI agent auditor will ask.

Then insist on a proof of value. Pick one specific control you don't have today and run the product against it on your own data. Decide from what you saw. A vendor that won't allow one has answered for you.

How do you handle a vendor that's too new to trust?

Treat maturity as its own test, and judge a young vendor on what it can prove rather than how good the demo looks. On August 31, 2026, the information security director at a large law firm told me every emerging AI security vendor he'd scanned still read like a demo, and he wasn't bringing a five-person startup into a firm that size. I had a counterweight: one vendor in that same scan was performing well, trade-offs and all. A maturity gate doesn't ban young vendors. It sets the terms they have to meet.

Contracts are where the caution turns into action. On two separate calls that summer, the guidance was identical: buy against one named missing control, and take the shortest contract you can get, because this market punishes long commitments. On August 3, 2026, a vendor a hospital operator was evaluating got acquired halfway through the evaluation. That risk now sits inside the buying decision. And incumbents aren't a free pass: the same law firm was leaving its network security vendor at an April 2027 renewal, not over features but over a slow decay in support. Old vendors and new ones both need an exit plan. The fake-agent side of this problem is in what is agent washing.

Frequently asked questions

Do I need a formal scoring model to compare vendors?

No, but you need the same questions for every vendor. A simple table with the seven categories as rows and each vendor as a column is enough, and weights are optional. What breaks comparisons is asking one vendor about data isolation and forgetting to ask the next.

Is an incumbent vendor always the safer choice?

Not always. Your existing tools often sit in the AI traffic path already, so they're the right first look. But incumbents decay too, and a renewal date is a fine forcing function for an honest review of support quality.

How long should the contract be?

As short as the vendor will accept. I don't have a magic number and won't invent one. Vendors in this market get acquired or folded into bigger products mid-contract, and one got acquired mid-evaluation this year.

What if a vendor refuses a proof of value?

Treat the refusal as data. A vendor that trusts its product will agree to a limited test against one named control. A vendor that won't is asking you to buy on a slide deck, which is the exact thing you're trying to avoid.

Key takeaways

  • The peers who could vouch for AI security products haven't finished their own first rollouts, so the evidence has to come from the vendor, in writing.

  • Seven categories, asked identically of every vendor, make the answers comparable. The floor underneath has five parts, and below it you stop talking.

  • Six features are table stakes now. Five are worth paying for. Sort them before the demo.

  • Start from the CISA Secure by Demand Guide, then add the AI questions: models underneath, data ownership, log export, and the framework's five elements.

  • Buy against one named missing control and demand a proof of value. Sign the shortest contract you can.

Before any vendor conversation, know which of the five elements you're actually missing. The free self assessment takes about ten minutes and gives you the named control to buy against.

Peer references will come, in time. Until they do, the evidence has to come from the vendor, in writing, before you sign.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.