AI · 12 min

How to tell if a team actually delivers AI - or just sells the buzzword

Questions we'd ask before funding an 'AI-powered' roadmap. Job-to-be-done, failure modes, metrics, and who owns the thing next month.

If they can't explain the job without saying 'AI'...

Walk away from the buzzword and ask what gets faster, safer, or newly possible for a real person. Good teams talk about the workflow first. Model second.

Make them walk one scenario end to end: who triggers it, what context they need, what a good answer looks like, and what happens when it's wrong. If that story is fuzzy, the build will be too.

Ask about failure out loud

Every AI feature fails sometimes. The tell is whether they designed for that - or only demoed the happy path.

  • Wrong, slow, or expensive model responses - then what?
  • Is there a human review path for high-stakes output?
  • Fallback UX when AI is down: degraded, not dead
  • Who gets pinged when cost or quality spikes overnight?

Got a vendor pitch or an internal roadmap slide? We'll help you poke holes in it.

Pressure-test an AI idea →

Metrics that aren't 'we called GPT'

Token counts look busy and tell you almost nothing. Ask about task success, latency under load, cost per successful completion, and whether support tickets about AI answers are going up or down.

If the only dashboard is API usage, you don't have a product feature. You have a bill.

The data story is usually the hard part

Models are the easy bit. Permissions, freshness, and 'where does truth live?' are where projects stall for months.

Ask who owns cleaning the data. Whether answers need citations. What's allowed into retrieval (PII, contracts, chat logs). And how they'll evaluate answers on your data - not a public demo set that makes everyone look smart.

Ask to see a boring test, not another polished demo

A useful test uses twenty or fifty examples from your actual work, including incomplete inputs and awkward edge cases. The team should be able to show which answers passed, which failed, and why.

This is less exciting than a live chatbot demo. It is also much closer to the work required after launch. If a supplier resists measurable examples, they may not know how to move beyond prompt tweaking.

  • Include easy, ambiguous, and deliberately adversarial examples
  • Agree who decides whether an answer is acceptable
  • Save the test set so future model changes can be compared
  • Track both quality and cost; the best answer may be too slow or expensive

Check the human workflow around the AI

Most valuable AI features do not replace an entire job. They remove a slow step: draft a reply, classify a request, extract fields, or suggest the next action.

Ask what the person sees before accepting the result, whether they can correct it, and where that correction goes. A useful feedback loop beats a clever model hidden behind a spinner.

Who runs it next month?

Prompts drift. Models change. Partner APIs change. If nobody owns evals after launch, quality dies quietly while the website still says AI-powered.

We build AI inside products we also know how to ship and operate - mobile, web, the boring release stuff. If you want a second opinion on a proposal, send it. We'll be blunt.

A reasonable first AI release is deliberately narrow

Pick one repeated task, a known group of users, and a clear fallback. Run it long enough to learn where people disagree with the output before connecting more tools or giving the system broader permissions.

That approach may look modest in a roadmap meeting. In practice it gets a trustworthy feature into use faster than building a general assistant that tries to understand the whole company on day one.

Next step

Got a prototype, a messy MVP, or just a problem?

Send it over. We'll tell you what we'd keep, what we'd rewrite, and what can wait - without a 40-slide deck.