choosing an AI model

Choosing an AI Model: How to Read Benchmarks Honestly

August 30, 2026 · 7 min read · Autana Solutions, Vancouver
Choosing an AI Model: How to Read Benchmarks Honestly — Autana Solutions

Every few weeks a new model tops a chart, someone forwards it to us, and the message is one line: should we switch? Fair question. Wrong first question. Choosing an AI model starts with what your business does all day, not with a ranking.

The stakes are ordinary and real. Statistics Canada reports that 19.2% of businesses used AI to produce goods or deliver services in the second quarter of 2026, up from 12.2% a year earlier and 6.1% in the second quarter of 2024 (Statistics Canada). The leading uses were data analytics (36.6%), text analytics (34.5%) and virtual agents or chatbots (28.2%). Those are front desk and back office tasks. No public leaderboard scores them.

A benchmark measures the benchmark

A benchmark is a fixed set of questions scored under controlled conditions. That's genuinely useful, and it's also the whole limit. NIST puts it bluntly in its AI Risk Management Framework: measurement approaches "can be oversimplified, gamed, lack critical nuance, become relied upon in unexpected ways, or fail to account for differences in affected groups and contexts" (NIST AI 100-1, 2023). The same framework warns that measuring a system "in a laboratory or a controlled environment" can yield results that "differ from risks that emerge in operational, real-world settings," and its MEASURE 2.5 category asks organizations to document "limitations of the generalizability beyond the conditions under which the technology was developed."

Translated for a shop in Burnaby: a model that aces graduate-level exam questions has told you nothing about whether it can read your messy intake form, spot that the customer already called twice, and book the right service.

Three things a published score hides

  • Selective disclosure. The Leaderboard Illusion study documented undisclosed private testing on a major public arena, including 27 private model variants tested by one provider ahead of a single release. Providers can test many candidates and publish only the winner (Singh et al., 2025).
  • Unequal practice data. In the same study, two proprietary providers received roughly 20.4% and 19.2% of total arena data, while 83 open-weight models combined received 29.7%. The authors estimate that additional arena data can produce up to 112% relative performance gains on the arena's own distribution. Part of a score reflects who got to practice on the test.
  • No error bars. Evan Miller's paper "Adding Error Bars to Evals" argues that evaluations are experiments and should report standard errors and confidence intervals like any other experiment (Miller, 2024). Most headline comparisons don't. Two models a point apart are often just tied.

The numbers you can actually verify are the prices

Capability claims are slippery. Published prices are not. They sit on vendor documentation pages with a date, and you can check them yourself before you commit.

Checked on 30 August 2026, Anthropic's own pricing documentation lists Claude Haiku 4.5 at $1 per million input tokens and $5 per million output tokens, Claude Sonnet 5 at $2 and $10, and Claude Opus 5 at $5 and $25, with a 50% discount on batch processing and cache reads billed at 0.1x the base input rate (Anthropic pricing docs). On the same date, OpenAI's pricing page lists gpt-5.6-luna at $0.20 input and $1.20 output per million tokens at short context, and gpt-5.6-sol at $4.00 and $20.00 (OpenAI pricing docs).

Notice the spread. Inside a single vendor's lineup the cheapest and priciest options differ by roughly twenty times. That gap is usually a bigger lever on your bill than the gap between vendors.

For scale, Anthropic's documentation works through its own illustrative example: about 3,700 tokens per support conversation, run on Claude Haiku 4.5, comes to roughly $37 per 10,000 tickets. Treat that as illustrative only. Your prompts, your knowledge base and your retry rate will move it.

All of these figures go stale fast. Anything you read about model names, prices, context windows or feature availability needs re-checking against the vendor's live page before you sign anything.

Build your own twenty-question benchmark

This is the part almost nobody does, and it takes an afternoon.

Pull 20 to 30 real items from your last month of work. Actual voicemails, actual emails, actual quote requests, including the ugly ones with typos and half a sentence. For each, write down what a good outcome looks like in one line. That's your test set, and it's worth more than every leaderboard combined because it matches your distribution.

Then run each candidate model against the same set, three times per item, because these systems are not deterministic. Have a person who knows the work score each response pass or fail. Log cost per item and time to first response while you're at it. Now you have three numbers that mean something: pass rate on your work, dollars per task, and seconds to answer.

You'll usually find the cheap model handles 80% of your volume and the expensive one earns its keep on the messy 20%. That's the actual answer to choosing an AI model for most Metro Vancouver businesses. It's rarely one model. It's a cheap default with an escalation path.

Where this doesn't apply

Honesty cuts both ways here, so here's the evidence against rushing.

A Statistics Canada study by Jiang Li and Huju Liu found that AI adopters showed 16.8% higher labour productivity than non-adopters under standard controls. Once pre-adoption productivity was accounted for, the premium fell to 10.2%. Once complementary capabilities such as R&D, cloud computing, data analytics and ICT training were included, it fell to 5.1% and became statistically insignificant (Li and Liu, 2026). The authors caution that gains "often take time to materialize and depend on organizational change and complementary investments." Read that as: the model is not the thing that pays off. The process change around it is.

Plenty of businesses are right to sit this out. In the second quarter of 2026, 40.0% of businesses stated that AI is not relevant to the goods they produce or the services they deliver (Statistics Canada, Canadian Survey on Business Conditions). Cybersecurity and privacy concerns were cited by 13.4% and cost by 10.6%.

If your work touches health information, hiring, lending or tenancy decisions, model choice is the smaller half of the problem. Canada's federal, provincial and territorial privacy regulators expect organizations to limit collection to "what is necessary for the purpose," to use anonymized or de-identified data where possible, and to be able to demonstrate compliance (OPC and provincial regulators, 2023). Those obligations don't change when a new model ships.

And benchmarks still have a job. They're a reasonable way to build a shortlist of three candidates. They're a poor way to pick one.

Sources

  • Anthropic. "Pricing," Claude platform documentation, accessed 30 August 2026. https://platform.claude.com/docs/en/about-claude/pricing
  • OpenAI. "Pricing," API documentation, accessed 30 August 2026. https://developers.openai.com/api/docs/pricing
  • National Institute of Standards and Technology. NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), January 2023. https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
  • Singh, S., Nan, Y., Wang, A., et al. "The Leaderboard Illusion," arXiv:2504.20879, 2025. https://arxiv.org/abs/2504.20879
  • Miller, E. "Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations," arXiv:2411.00640, 2024. https://arxiv.org/abs/2411.00640
  • Statistics Canada. "Analysis on artificial intelligence use by businesses in Canada, second quarter of 2026," released 11 June 2026. https://www150.statcan.gc.ca/n1/pub/11-621-m/11-621-m2026010-eng.htm
  • Statistics Canada. "Canadian Survey on Business Conditions, second quarter 2026," The Daily, released 27 May 2026. https://www150.statcan.gc.ca/n1/daily-quotidien/260527/dq260527a-eng.htm
  • Li, J. and Liu, H. "Artificial intelligence adoption and productivity in Canadian firms," Statistics Canada, released 22 April 2026. https://www150.statcan.gc.ca/n1/pub/36-28-0001/2026004/article/00002-eng.htm
  • Office of the Privacy Commissioner of Canada and provincial and territorial privacy regulators. "Principles for responsible, trustworthy and privacy-protective generative AI technologies," 7 December 2023. https://www.priv.gc.ca/en/privacy-topics/technology/artificial-intelligence/gd_principles_ai/

If you'd rather not spend an afternoon building a test set, that's the kind of thing we do for a living. Book a free call with Autana Solutions and we'll walk through your real workload, score a couple of models against it, and tell you honestly if the answer is "not yet."

AI modelsbenchmarksevaluationsmall businessAI strategy

Want an AI employee for your business?

We install a 24/7 AI worker for businesses in Vancouver, Burnaby, and beyond. Book a free Discovery Call.

Book a call

Keep reading