Key takeaways
✓Most AI vendors over-promise. A structured set of questions asked before you sign anything will surface the gap between demo and reality faster than any reference check.
✓The evaluation should cover four areas: data handling, integration, support, and the vendor's own roadmap. Skipping any one of them creates risk.
✓Red flags are as important as green lights. A vendor who cannot answer a direct question about data residency or model versioning is telling you something important.
✓Procurement is not just an IT problem. Legal, compliance, and the teams who will actually use the tool all need a seat at the table before the contract is drafted.
✓Good AI procurement is repeatable. The questions in this article can be turned into a standing evaluation template your organisation uses for every new tool.
Why AI vendor procurement is harder than it looks
Most enterprise software purchases are hard to reverse. A bad CRM costs you a painful migration and some lost productivity. A bad AI vendor can cost you considerably more: compromised data, a workforce that has lost trust in the technology, and a governance mess that takes years to untangle.
The problem is that AI demos are exceptionally well-designed to impress. Vendors run them in controlled conditions, on clean data, with outputs that have been curated to show the tool at its best. What you see in a 45-minute product demo rarely resembles what the system does when it hits your actual data, your edge cases, and your compliance requirements.
There are a few structural reasons procurement teams struggle here.
First, AI capabilities are genuinely novel. Procurement professionals who can rigorously evaluate a logistics platform or a payroll system may have no prior reference point for judging whether a large language model's accuracy claims are credible, or what "hallucination rate" means in practice for a legal document workflow.
Second, the category is moving fast. A vendor's product roadmap from six months ago may bear little resemblance to what they can actually deliver today, in either direction. Some things that were promised have quietly disappeared. Some capabilities are real but require a higher pricing tier, a separate integration, or a level of technical setup your team is not resourced for.
Third, cost structures are opaque. Many AI tools are priced per token, per API call, or per user with consumption-based overages. Without a realistic estimate of your usage volume, the initial quote is nearly meaningless. A finance team that runs the tool lightly in a pilot will face a very different invoice once three departments are using it at scale.
The demo is not the product
Vendors build demos to eliminate uncertainty. Your evaluation process needs to reintroduce it, deliberately, by testing against your own data, your own workflows, and your actual edge cases.
For Australian organisations, there is also a data residency dimension that adds genuine complexity. Many AI platforms process data through infrastructure hosted overseas. That may be acceptable depending on your industry, your data classification, and your interpretation of the Privacy Act 1988. But you need to ask the question explicitly, because vendors rarely raise it unless you do. The data leakage scenarios that matter most are rarely the ones on the vendor's FAQ page.
None of this means AI purchases are too risky to make. It means the evaluation process needs to be more disciplined than it would be for conventional software.
How should you structure an AI vendor evaluation?
A good evaluation starts before the first vendor call. If you wait until a demo to decide what matters, the vendor controls the agenda. Their polished walkthrough will show you the best-case scenario, and you will walk away impressed without having learned anything useful.
The structure that works is straightforward: define your criteria first, then bring the right people into the room, then request documents before any live demonstration.
Who needs to be involved
Procurement and IT are the obvious seats at the table. But AI vendor evaluations routinely go wrong because they stay too technical for too long. The people who will actually use the tool, a team lead from finance, a senior analyst, whoever owns the process being changed, need to be present early. They will spot integration problems that no architect would think to ask about.
You also want a representative from legal or risk, particularly if the tool will process personal data or connect to internal systems. Discovering a data residency problem after you have shortlisted two vendors is a waste of everyone's time. For government or regulated-sector procurement, this is non-negotiable. The AI in the Australian public sector context makes early legal involvement even more important.
What to request before the demo
Send a short document request to every vendor at least a week before any live session. Ask for:
A one-page architecture summary, showing where data flows and where it is stored
Their most recent third-party security audit or SOC 2 report (or equivalent)
Documented uptime figures for the past 12 months, not projected availability
A list of current enterprise customers in a similar industry or scale, with references available on request
Pricing documentation that covers overage costs and contract exit terms, not just the headline figure
Vendors who push back on providing these documents before a demo are telling you something. The ones worth talking to will send most of this without hesitation.
The sequence matters
Set your evaluation criteria before the first vendor call. Once you have sat through a compelling demo, anchoring bias makes it genuinely harder to apply the same standard to everyone.
How many vendors to evaluate in depth
Shortlist ruthlessly. Three vendors evaluated properly is worth more than seven evaluated superficially. An initial screening round based on written responses and document review can eliminate half your list without a single call. Save the structured, multi-stakeholder sessions for the vendors who clear that bar.
The 10 questions that filter hype
Ask these in any vendor conversation, whether it is an initial demo, a formal RFP response, or a proof-of-concept review. A confident vendor will welcome the specificity. A vendor that deflects, generalises, or talks over you is telling you something important.
1. Which specific task will your tool perform, and how is performance measured?
AI vendors often describe capabilities in terms of categories ("it does document processing") rather than tasks ("it extracts line items from invoices and matches them to purchase orders with X% accuracy"). Push for the task level. Ask how they measure whether the tool is doing that task well, and what the benchmark is based on: their own tests, independent evaluation, or customer production data.
A good answer names a metric and a methodology. A weak answer describes the category and changes the subject.
2. Where does our data go, and who can access it?
This is not a technical question; it is a governance question. You need to know whether your data is used to train shared models, whether it is retained after a session ends, which cloud region it sits in, and who at the vendor has access. For Australian organisations, data residency often has compliance implications, particularly in health, finance, and government.
Ask for this in writing. A vendor that cannot answer it clearly before the contract is signed will not answer it clearly after.
Data handling is a procurement decision, not just a security review
If your AI vendor cannot tell you exactly where your data goes and who can access it, the conversation should pause until they can. This is due diligence, not a technical detail to sort out later.
3. What does the integration actually require?
"Easy to integrate" covers a wide range, from a browser extension your team installs in an afternoon to a six-month API build that requires a dedicated engineering team. Ask for the actual integration architecture: what systems it connects to, what APIs or connectors exist, what your IT team needs to provide, and what the vendor provides versus what you build yourself.
If they point you to a partner ecosystem to answer this question, ask who leads the integration and what the typical cost and timeline is.
4. What happens when the model is wrong?
Every AI system makes errors. What matters is whether the vendor has thought about this seriously. Ask how the tool signals low-confidence outputs, what the failure mode looks like in production, and how your team is expected to catch and correct errors. For high-stakes decisions, a verification protocol is essential, and the vendor should support one rather than resist it.
A vendor that claims the model is accurate enough that errors are not a concern is a vendor whose risk model does not align with yours.
5. Who else in our industry is using this, and what specifically are they doing?
Reference customers are expected. What matters is whether the references are relevant. A logistics reference does not help a legal team evaluate a contract review tool. Ask for customers in your sector, using the same feature set, at a similar scale. Then actually call them and ask what did not work, not just what did.
6. How do you handle model updates, and how are we notified?
AI models change. A vendor pushing a model update can change the behaviour of a tool your team relies on without any visible change to the interface. Ask whether model updates are automatic or opt-in, how breaking changes are communicated, and what the rollback options are. This matters especially if the tool is embedded in a regulated workflow.
7. What does your pricing look like at scale?
Early pricing is often designed to be easy to say yes to. What you need is the pricing curve: what happens when you go from 50 users to 500, from 1 million tokens to 100 million, or from one use case to five. Ask for a worked example at the scale you expect to reach in 18 months, not the scale you are at today.
Hidden consumption-based costs are one of the most common sources of budget overruns in AI deployments. Ask the vendor to show you what a bad month looks like.
8. How do you support our team's adoption, not just the technical rollout?
A tool that no one uses is a cost, not a capability. Ask what the vendor provides beyond implementation: training materials, change management resources, user onboarding, and ongoing support. If the answer is a link to documentation and a support ticket queue, you are responsible for the adoption work yourself. That is not necessarily a problem, but you should price it in. Change management for AI rollouts takes real effort, and some vendors actively support it while others do not.
9. What are the security certifications, and what is in scope?
ISO 27001, SOC 2 Type II, and ASD Essential Eight alignment are common reference points in Australian enterprise and government contexts. Ask which certifications they hold, when they were last audited, and, critically, what is actually in scope. A vendor can be SOC 2 certified for their billing infrastructure while their AI inference layer sits outside the scope entirely.
If security certifications are relevant to your procurement, ask for the certificate and the scope boundary, not just a yes or no.
10. What does the exit look like?
Vendor lock-in is a real risk in AI, where proprietary data formats, embedded integrations, and trained workflows can make switching painful. Ask what happens to your data if you leave, how long it is retained, what export formats are available, and what the contract terms are around termination. A vendor that makes exit easy is a vendor that is confident in the ongoing value of what they are building.
This question also reveals how the vendor thinks about the relationship. A long or evasive answer is worth noting.
What red flags should stop the conversation?
Most vendors will say the right things in a first meeting. The signals that matter come later, when you push for specifics.
They can't show you a real deployment. A polished demo built on curated data is not evidence. If a vendor cannot connect you with a reference customer in a comparable industry, or show you a live production environment (even anonymised), treat the product as unproven. A demo proves the best case. You need to see the ordinary case.
The security answers are vague or deferred. "Our legal team handles that" or "we can share our SOC 2 report after signing" are not acceptable responses during evaluation. Data residency, encryption standards, and model training practices should be documented and available before you invest further time. If the vendor can't answer your data questions without a lawyer in the room, that tells you something about how they handle compliance operationally. The data leakage risks are real enough that this is not a box to tick at contract stage.
Pricing only makes sense at scale they're defining for you. Watch for proposals where the cost model is opaque until you commit to a minimum seat count or a platform bundle. Responsible vendors can model your actual usage and give you a realistic number. If the pricing conversation keeps steering back to enterprise packages before you've agreed on scope, that's a commercial structure designed to obscure value.
They promise accuracy without defining it. "95% accurate" means nothing without knowing what task, what data, and what the failure mode is on the other 5%. If a vendor quotes accuracy figures without qualifying them, ask directly: accurate on whose data, measured how, and what happens when the model is wrong? Evasion here is a strong signal to pause.
The contract locks in the relationship before the outcome is demonstrated. Multi-year commitments, auto-renewal clauses, and restrictive data portability terms all shift risk onto you before the vendor has earned that trust. A vendor confident in their product will accept a pilot with defined exit terms. One who resists that arrangement is protecting their revenue, not your outcome. This connects directly to the pilot structure: if you haven't built a way out, you've already lost negotiating leverage. The bridge from pilot to production should be something you define, not something the vendor defines for you.
The question that cuts through most
Ask the vendor what a failed deployment looks like with their product, and how they handle it. A good vendor has a clear answer. A vendor who can only describe success scenarios has not thought seriously about your risk.
They're dismissive of your governance requirements. If you raise AI governance questions and the vendor treats them as bureaucratic friction, that attitude will not improve after the contract is signed. Governance requirements are not negotiable obstacles; they're operating conditions. A vendor who understands enterprise deployment understands this.
None of these flags are automatic disqualifiers in isolation. A new vendor may have limited reference customers. A startup may still be building out its compliance documentation. Context matters. But multiple flags together, or any single flag that touches data security or contractual lock-in, should put the deal on hold until you have clear answers.
Frequently asked questions
How many vendors should we evaluate in an AI procurement process?
Three to five vendors is a practical range for most enterprise AI procurement processes. Fewer than three gives you no real comparison; more than five and the evaluation becomes difficult to run rigorously, and vendors sense they are in a cattle call and disengage. Shortlist based on a written requirements document, then run your full question set against the shortlist only.
Should we run a proof of concept before signing a contract?
Yes, and you should negotiate the terms of the proof of concept before you start it. Define what success looks like in writing, agree on the dataset or workflow you will test against, and set a time limit. A vendor who resists a scoped POC on your actual data is telling you something. A POC that drifts into a months-long pilot with no exit criteria is not a POC, it is a soft lock-in.
What data residency requirements apply to Australian enterprises buying AI tools?
Australian enterprises subject to the Privacy Act 1988 must ensure personal information is handled in accordance with the Australian Privacy Principles, which include obligations around cross-border disclosure. Some industries, including health and finance, carry additional sector-specific requirements. Ask every vendor to state in writing where your data is stored, processed, and backed up, and whether it is used to train shared models. If the vendor cannot answer that in plain language, the answer is probably not in your favour.
How do we evaluate AI vendors when our own team does not have deep technical expertise?
Focus your evaluation on outcomes and accountability rather than architecture. Ask vendors to demonstrate the tool on a real workflow you own, not a polished demo environment. Bring in a third-party adviser for the technical due diligence, or use a structured question framework like the one in how to evaluate agentic tools. The questions around data handling, SLA commitments, and escalation paths do not require deep technical knowledge to ask or to score.
Is it worth buying from a large platform vendor versus a specialist AI vendor?
It depends on what you are trying to solve, and both have genuine trade-offs. Large platform vendors offer integration with tools you likely already use, clearer enterprise support paths, and more predictable pricing models. Specialist vendors often provide deeper capability in a narrow domain and more flexibility to customise. The risk with specialists is vendor viability: an AI startup that looks impressive today may be acquired, pivoted, or wound down within two years. Build contractual protections around data portability and transition assistance regardless of which type you choose.
Ready to move from evaluation to implementation?
Getting the vendor selection right is only half the challenge. The harder work is integrating that tool into real workflows, upskilling the people who will use it, and building the governance layer that keeps it honest over time.
If you are working through an AI procurement process and want a clearer view of what good implementation looks like before you sign anything, that conversation is worth having early.
Not sure what to ask your AI vendor next?
We work with Australian enterprises at every stage of the AI journey, from scoping and vendor evaluation through to training and adoption. A short call can help you pressure-test your shortlist before you commit.
