Playbook · AI evaluation

How we evaluateAI projects.

Most AI projects don't fail on the model. They fail on data nobody could reach, an error rate nobody priced, a workflow nobody watched, or a definition of working that nobody wrote down. These are the seven questions we answer first.

Coloured blocks stacked into an interlocking structure, each one load-bearing

The demo is the easy part. It has been the easy part since the models got good, and it is why so many organizations now have a folder of impressive prototypes and nothing in production. The gap between a convincing demo and a system people rely on is not model quality. It is everything around the model.

So we evaluate AI projects the way you would evaluate any other capital decision: what changes, what it is worth, what it costs at scale, what happens when it is wrong, and who is accountable. Seven questions. Two weeks. A scored answer in writing.

Three of those questions have ended projects for us before we wrote any code. That is not a failure of the process — it is the entire point of running it early, while stopping is still cheap and nobody has to defend a year of spend.

Seven questions, answered honestly

  1. 01 · Value

    What decision or action changes — and what is that worth?

    Not “it summarizes documents.” What happens differently on Tuesday because of it, who does it, and what does that hour or that error cost today? If the honest answer is that a person still reads everything afterwards, the value is a demo, not a number.

    How we test it. We test it by pricing the manual version. If paying a person to do this by hand would be worth it, the AI version has a business behind it. If it wouldn't, no model fixes that.

  2. 02 · Data

    Does the evidence exist, and can we actually reach it?

    Intelligence has to be grounded in something. The question is never whether your company has data — it always does — but whether the specific evidence this answer depends on exists in a retrievable form, and whether legal, security and the team who owns it will let you near it.

    How we test it. We ask for a real sample in week one. Access granted in week one is a good sign. Access still pending in week three is the answer.

  3. 03 · Error tolerance

    What happens the day it is confidently wrong?

    Every one of these systems is wrong sometimes, and the failures do not look like failures — they look fluent and assured. So we work out the blast radius up front. A wrong draft a human edits is an inconvenience. A wrong dosage, price, or eligibility decision is a different category of thing entirely.

    How we test it. We map the worst plausible output to a real consequence, then design where the human sits relative to it — before the architecture, not after.

  4. 04 · Workflow fit

    Whose job does this change, and have they agreed?

    The most common cause of an unused AI feature is not quality. It is that the tool lives beside the work instead of inside it, and the person it was built for was never asked. A perfect answer in the wrong window is a tab nobody opens twice.

    How we test it. We watch the work happen before we design anything, and we name the specific moment the system intervenes. If nobody can point to that moment, we stop there.

  5. 05 · Evaluation

    How will we score it, before we build it?

    Traditional software either works or throws an error. These systems degrade quietly and gradually, which means “it seems better” is the default review process unless you build something better first. That something is a graded set of real cases with known good answers.

    How we test it. We build the evaluation set before the feature. If we cannot write fifty realistic cases and agree on what a good answer looks like, we do not understand the problem well enough to solve it.

  6. 06 · Unit economics

    Does this survive success?

    Costs that look irrelevant in a pilot with twelve users become the whole conversation at twelve thousand. Same with latency: three seconds is charming in a demo and intolerable inside a workflow someone repeats ninety times a day.

    How we test it. We model cost and latency per interaction at ten times your expected volume. If the margin only works at pilot scale, we redesign it now rather than discover it at the renewal.

  7. 07 · Accountability

    Who answers for the output?

    Somebody in your organization owns what this system says to a customer, a regulator, or an employee. When that person is unnamed, the project stalls at the exact moment it starts to matter — usually in a security or compliance review that nobody scheduled.

    How we test it. We put the name in the document. Along with what gets logged, how long it is kept, what a customer can ask to see, and how a bad output gets reported and fixed.

Three answers that end a project

We hear all three regularly, usually said cheerfully. Each one means the project as described will not survive contact with production, and each one is far cheaper to hear now than in month nine.

“The data exists — we just can't get to it.”

This is almost never a technical problem, which is why it is so rarely solved by a technical team. It is procurement, an ownership dispute, a vendor contract, or a security posture that predates the project. We would rather spend two weeks finding out whether that door opens than six months building against a system we were never going to be allowed to read.

“We'll know it's working when people like it.”

Sentiment is not a measure, and it is the first thing to move when the novelty wears off. Without a graded evaluation set, every future change becomes a guess and every regression goes unnoticed until a customer finds it. If nobody will commit to a number now, the honest read is that nobody believes the number will be good.

“It has to be right every single time.”

Then it should not be a probabilistic system, or a person needs to sit between it and the consequence. Both are fine answers — the deterministic path is often the better product anyway. What does not work is demanding certainty from something that trades in likelihood, then acting surprised.

Two weeks, fixed scope

  1. Days one and two

    Price the outcome

    We work out what actually changes and what it is worth, in your numbers rather than a vendor case study. This is also where most ideas get narrowed — usually to a smaller, sharper version of themselves that is worth far more.

    • The decision or task the system touches, named precisely
    • The cost of doing it manually today, per unit
    • The one metric that would prove it worked
  2. Days three to five

    Test the data reality

    We ask for a real sample of the real thing, not a scrubbed export. What we learn about quality, coverage and access in these three days is usually the single biggest predictor of whether the project ships.

    • A genuine sample, with its gaps and inconsistencies intact
    • Retrieval tested against the questions people actually ask
    • Legal, security and data ownership brought in now, not at the end
  3. Week two

    Build the scorecard, then a thin slice

    We write the evaluation set with your experts — fifty or so real cases with agreed good answers — and then build the narrowest working version we can measure against it. Not a chat window bolted on. The actual moment of work, running.

    • An evaluation set your team owns and keeps using
    • A thin end-to-end slice, in the real workflow, on real data
    • Cost, latency and quality measured rather than estimated
  4. The end

    Go, narrow, or stop

    You get a recommendation with the numbers behind it, and we are equally willing to write any of the three. Stopping in week two is a good outcome. It is a fraction of the price of stopping in month nine, and nobody has to defend it in a board meeting.

    • A scored answer against all seven criteria, in writing
    • The evaluation harness and thin slice, handed over either way
    • If go: a sequenced build plan. If not: what would have to change

Larger programs — several use cases, multiple data domains, regulated environments — run longer. The order holds: value, data, evaluation, then a thin slice that proves it in the real workflow.

What you get

  • A scored answer against all seven criteria, in writing, with the reasoning shown
  • An evaluation set of real cases your team owns and keeps using
  • A thin working slice in the real workflow, on real data
  • Cost and latency modelled at ten times your expected volume
  • A named accountability model — logging, retention, escalation, human checkpoints
  • A go, narrow, or stop recommendation we're equally willing to write

Who it's for

  • Leaders holding a mandate to "do something with AI" and no way to rank the options
  • Teams with a promising prototype that hasn't survived a security or compliance review
  • Organizations weighing build against buy on a feature vendors are pitching hard
  • Product teams whose pilot got good feedback and no measurable outcome
  • Anyone about to sign a large AI budget on the strength of a demo

Questions we get

How long does the assessment take, and what does it cost?
Two weeks is the standard shape. It is a fixed-scope, fixed-price engagement, and the evaluation harness and thin slice are yours regardless of the recommendation — including when the recommendation is not to build it. We would rather be paid once for an honest no than three times for a slow yes.
We already picked a model and a vendor. Is this still useful?
Usually more useful, not less. Almost none of the seven criteria are about the model. Data access, error tolerance, workflow fit, evaluation, unit economics and accountability all sit outside the vendor decision, and they are where projects actually fail. If your model choice turns out to be wrong we will say so, but that is rarely the finding.
Does this apply to agents, or only to retrieval and chat features?
Both, and agents raise the stakes on every criterion. An agent taking actions rather than producing text turns error tolerance and accountability from paperwork into the core design problem — what it is allowed to do unsupervised, what requires a human, what gets logged, and how it is stopped. We evaluate agents against the same seven questions with a much harder line on the last two.
What if we fail the assessment?
There is no pass mark — there is a scored picture of where the risk actually sits. Most projects come out amber somewhere, and the useful output is knowing which two things to fix first. The three answers we listed above are the genuine stoppers, and in every case the honest move is to fix the underlying condition rather than build around it.
Can you just build it? We're confident about the case.
Yes — if you can answer the seven questions, we will happily skip the assessment and start building. We are not attached to selling a discovery phase. We are attached to not spending your budget on something the data was never going to support.

Sitting on an AI idea you can't quite price?

Tell us what you're hoping it does and where the data lives. Two weeks from now you'll have a scored answer with the numbers behind it — including if the answer is don't.

Book a Strategy Session