Measuring AI accuracy before you buy
A vendor's accuracy slide is marketing until you test it on your own documents, calls or images.
· 2 min read · For anyone evaluating an AI vendor's claims

Every AI vendor shows a slide with an accuracy number — 95%, 98%, whatever sounds impressive. That number was measured on their test data, in their conditions, and tells you almost nothing about how it will do on your documents, your call recordings or your factory's camera footage. The only accuracy number worth trusting is one measured on your own data.
What accuracy actually means
A single percentage hides a lot. A system can be very accurate on easy cases and poor on hard ones, and still report a high average. Ask what was measured — is it how often the top answer is right, or how often the right answer is somewhere in a list of options? For anything where missing something matters more than a false alarm (a defect on a production line, for instance), ask for both: how often it catches a real problem, and how often it flags something that is not actually a problem.
How to test it properly
- Use a sample of your own real data — documents, calls, images — not a vendor demo.
- Include the messy, ordinary cases, not just the clean examples that show the product off well.
- Have someone who is not selling you the product check the results against the correct answer.
- Ask what happens on the cases it gets wrong — does it say "not sure", or does it guess confidently?
What it costs you
A proper test takes time and a real sample of your data, which some vendors resist providing for because their number looks worse on it than on their own. Treat that resistance itself as useful information. A short, honest test run on maybe a hundred real cases tells you more than any slide, and it is worth insisting on before any long-term commitment.
Questions to ask any vendor
- Can we test accuracy on our own sample data before we commit?
- What does the system do when it is not confident — guess, or say so?
- Is the accuracy number the same on hard cases as on easy ones?
How we can help
We test our models — Vizhi, Sol, Kural, Thedal and Thudippu — on your own data before you commit to anything, because that is the only number that means something to you. See our approach at /ai, and send us a sample at talk to us.


