The buyer's diagnostic for evaluating AI implementation partners.
There is a narrative gaining traction that AI is no longer the hard part. The tools are accessible, the models are powerful, and implementation is straightforward.
That narrative is wrong.
87% of AI projects never make it out of pilot. The tools are accessible. Production AI that stays running, stays within budget, and actually improves business outcomes is genuinely difficult. The gap between a demo and a production system is where most initiatives die. It is also where most consulting engagements reveal whether the partner is serious or selling.
The partner question matters precisely because AI is hard. Five questions cut through the sales process and test whether a prospective partner has the production disciplines that determine outcomes. Each question maps to a methodology that either exists or does not. There is nothing to fake.
Here are the five questions.
What should you ask about workflow architecture before any AI build?
Question 1. “Show me a workflow you audited before you built anything.”
The workflow audit is the upstream diagnostic that determines what gets built and where. Most production AI builds operate on workflows that are 70-80% deterministic rules with 20-30% genuine judgment at the edges. Partners that skip the audit treat the entire workflow as an agent problem. The cost compounds at scale.
A partner that runs workflow audits can explain the encodable layer versus the judgment layer, name the three tests that separate them (determinism, variance, reversibility), and show a real example of a workflow they decomposed before any code shipped.
A partner without a workflow audit practice builds agents on everything and hopes the cost works out.
Watch for partners that jump straight to building without being able to name what should have been automation.
How do you test whether the partner understands the build-versus-automate decision?
Question 2. “Show me how you decide what's an agent and what's automation.”
This is the follow-up to the audit question. It tests whether the partner has internalized the 80/20 split or whether they default to agents because the tools make it easy.
The answer should include a methodology for coding workflow steps as rules-required or judgment-required. It should reference specific examples where the partner moved work from the agent layer to the rules layer after an audit revealed the step was deterministic. The savings should be quantifiable. At 10,000 records per month across seven rules-required steps, the agent path runs $14K-56K monthly for outcomes that cost a couple hundred dollars in automation infrastructure.
Watch for partners that describe every workflow as “AI-powered” without distinguishing what the AI is actually doing versus what a conditional rule would handle.
What does a production-grade AI eval framework look like?
Question 3. “Show me your eval framework.”
The eval framework is what separates a production system from a demo. A demo works under ideal conditions with curated inputs. A production system handles incomplete data, unexpected edge cases, and model drift over time.
A serious partner has a formal evaluation methodology. Binary scoring per axis. A defined gate for autonomous deployment (five consecutive passing runs across all axes). An explanation of how the eval was built from observed failure modes rather than assumptions. Domain expert calibration on the LLM-as-judge.
If the partner's eval process is “we review the outputs and they look good,” the system will surprise the team in production. Quality by inspection is not quality by discipline.
Watch for partners that cannot explain how they built their eval rubric. If the rubric existed before the agent was shipped, it grades what the team imagined would break rather than what actually breaks.
How do you test for economic discipline in AI operations?
Question 4. “Show me a cost-per-business-outcome number from a production deployment.”
This question tests whether the partner has cost telemetry in place. Cost per API call is the wrong abstraction for leadership. The unit that matters is cost per qualified lead, cost per resolved ticket, cost per generated email that lands in a sequence.
A partner with economic discipline can produce per-run cost logging, category breakdowns (input tokens, output tokens, tool calls, sub-agent invocations), and the rollup from run-level cost to business-outcome cost. They can explain their budget gates and what happens when cost thresholds are exceeded.
A partner without cost telemetry cannot answer the question leadership will eventually ask. The bill arrives unattributed, and the pullback follows.
Watch for partners that quote total monthly cost without being able to attribute it to specific workflows and outcomes.
How do you verify that a partner has actual production experience?
Question 5. “Show me what's running in production right now, and how you monitor it.”
The simplest filter. Many firms can demonstrate impressive prototypes. Fewer can show systems that have been running for months with continuous telemetry, drift detection, and on-call protocols.
The answer should include how many workflows are currently live, what observability surfaces exist (dashboards, alerting, aggregate failure pattern surfacing), and what happens when something breaks in production. The on-call protocol matters. The debug path should start at the specific axis that failed, not at “the agent did it.”
A partner with production experience talks about monitoring, maintenance, and failure management. A partner without it talks about launches.
Watch for partners whose case studies end at deployment. Production is what happens after the launch.
The pattern across all five questions
Each question tests a different production discipline. Architecture (the workflow audit). Economics (the build-versus-automate decision and cost telemetry). Quality (the eval framework). Operations (production monitoring and maintenance).
A partner that passes all five has the disciplines required to ship systems that stay running. A partner that fails one reveals the exact gap that will surface during the engagement.
These questions do not require technical expertise to ask. They require honest answers. The answers are either specific and verifiable or vague and aspirational. The difference between the two is the difference between a production partner and a demo shop.
Five questions. Ask them before you sign.
FAQ
How do you evaluate an AI implementation partner?
Evaluate an AI implementation partner by testing five production disciplines. Ask for a workflow audit example (architecture). Ask how they separate agents from automation (economics). Ask for their eval framework (quality). Ask for a cost-per-business-outcome number from a live deployment (cost telemetry). Ask what is running in production right now and how they monitor it (operations). Partners that cannot answer all five with specifics are likely to produce demos rather than production systems.
What is the difference between a production AI partner and a demo shop?
A production AI partner has formal disciplines for workflow architecture, evaluation, cost telemetry, and ongoing operations. They can show live systems, per-run cost numbers, and eval frameworks built from observed failure modes. A demo shop can show impressive prototypes under ideal conditions but lacks the infrastructure to keep systems running at scale. The gap between the two is where 87% of AI projects fail.
Why should you audit a workflow before building AI agents?
The workflow audit determines what work should be an agent and what should be automation. Most GTM workflows are 70-80% deterministic rules with 20-30% genuine judgment. Treating the entire workflow as an agent pays agent-level token cost for the encodable 70-80%. The audit catches this architectural misalignment before the build ships and the cost compounds.
What is an AI eval framework and why does it matter for choosing a partner?
An AI eval framework is the quality gate that determines whether an agent is production-ready. Binary scoring per axis, five consecutive passing runs before autonomous deployment, domain expert calibration on the LLM-as-judge. Partners without an eval framework are shipping agents that grade themselves. The eval is the difference between hoping the system works and knowing it does.
How do you measure AI cost per business outcome?
AI cost per business outcome rolls up per-run telemetry to the unit that leadership cares about. A qualification agent gets measured in cost per qualified lead. A support agent gets measured in cost per resolved ticket. The rollup requires per-run cost logging, attribution of each run to a business outcome, and aggregation across runs. Partners that cannot produce this number are running blind on the metric the executive sponsor will eventually demand.
What are red flags when evaluating AI consulting firms?
Three red flags. The partner leads with models and platforms rather than business problems. The partner cannot show a single workflow currently running in production with continuous monitoring. The partner cannot explain how their eval framework was built from observed failure modes rather than assumptions. Each red flag signals a gap in the production discipline that will surface during the engagement.