Nitin Kumar SinghSolutions Architect

Type to search. to move, Enter to open.

    move open esc close
    Deep DiveAI Strategy

    Agree Before You Build: Five Gates for Enterprise AI Use Cases

    A five-gate method for picking enterprise AI use cases and measuring AI ROI in a way finance will accept and delivery can instrument, with a claims example.

    Most enterprises cannot tell you what their AI programme has returned. Not because the technology failed, but because the question was asked after the fact.

    McKinsey’s 2026 global survey of 1,719 executives found that 80% of AI users report improved individual productivity, while only 37% of organisations attribute any impact to enterprise earnings, a share unchanged from the previous year. Only 6% qualify as high performers, meaning they attribute at least 5% of EBIT to AI and describe the impact as significant.12

    The gap between “people feel faster” and “finance can see it” is the entire problem. The sharpest evidence of that gap comes from a randomised controlled trial. METR gave experienced open-source developers real tasks from their own repositories and randomly allowed or disallowed AI tools. Developers forecast a 24% speedup. Afterwards they believed they had been about 20% faster. Measured completion time was 19% slower.3

    METR now labels that result as historical, and later tooling may well perform differently. The three-way divergence between forecast, perception, and measurement is the point. Self-reported productivity is not a metric.

    I wrote this for the CIO or head of AI who has to answer to two people at once: a CFO who wants a defensible number, and a VP of Engineering who needs to know what to instrument. The argument is simple. Selecting an AI use case and measuring its return are not two activities. They are the same decision.

    A use case is “right” if, and only if, you can answer three questions before you build it. Will finance accept the metric as return? Can delivery capture that metric from the systems involved? What stops the project, and when? If those questions cannot be answered in writing before the first sprint, that is the answer.

    Why traditional ROI breaks for AI

    Finance already knows how to evaluate a capital project. What makes AI different is not the arithmetic but three properties that undermine the assumptions behind it.

    Performance drifts. A model that works at launch degrades as data, prompts, and upstream systems change. A static business case built on launch-week accuracy is wrong within a quarter. Deloitte’s guidance for CFOs names model drift, attribution opacity, and cost volatility as the reasons conventional ROI methods fail for AI, and identifies missing pre-deployment baselines as the root cause of weak evidence.4

    Attribution is contested. When a claims handler closes a file faster, was it the AI summary, the new intake form shipped the same month, or the two experienced hires? Without a control, every stakeholder claims the win.

    Cost is volatile and back-loaded. Pilots are cheap. Production is not. Gartner expects at least half of generative AI initiatives to exceed their planned budgets by 2028, attributing overruns to poor architectural choices and a lack of operational expertise, and expects inference to account for at least 70% of a model’s lifetime cost.5 In customer service specifically, Gartner predicts that by 2030 the cost per resolution for generative AI will exceed $3, higher than many offshore human agents, as vendors move from subsidised growth to profitability and use cases consume more tokens.6

    There is also a governance vacuum. KPMG’s 2025 CFO and CIO Collaboration Survey found that 59% of CFOs and 61% of CIOs each say they hold primary responsibility for AI and technology investment decisions.7 When both own it, neither does. Agreeing the metric, the cost, and the exit before build is how that ownership becomes explicit.

    The five gates

    Each gate is a question. A use case that cannot answer it does not proceed. The gates are ordered so that the cheapest question is asked first.

    Gate 1 — does the outcome already happen, and in what form?Proposed AI use casebefore any build budget is committedMeasured12 weeks of historyUncountedscored sample or shadow modeAbsentmeasure what it leaves unfixedIt is the platformfund as R&D, capped and datedGate 1Comparison named and accepted by finance?Gate 2One value type and one primary metric?Gate 3Fully loaded cost and production run-rate known?Gate 4Can it be tested against a control?the only gate that does not stop the workGate 5Kill criterion and review date written down?Construct the comparison firstNarrow the use caseModel the cost at scaleProceed, labelled UNVERIFIEDWrite the exit conditionsPre-build agreementfinance and delivery, signed before codeBuild
    Every gate can stop the proposal, and stopping is the cheap outcome. Gate 1 first classifies what kind of baseline exists, because the answer decides what the other four gates can even be measured against. A case that fails Gate 4 is not killed — it proceeds labelled UNVERIFIED in the portfolio, which is a different thing from proven.

    Gate 1: What is the comparison, and does finance accept it?

    Most write-ups of this gate ask whether a baseline exists, and that question quietly assumes every AI use case is an existing process done faster. Plenty are not. The requirement is a comparison finance will accept. Twelve weeks of history is the cheapest one, not the only one.

    The trap is the word “new”. A solution being new tells you nothing about whether a baseline exists. What matters is whether the outcome already happens, and those two come apart more often than proposals admit.

    This gate reclassifies more proposals than any other, and it should. The most-cited failure statistic in enterprise AI is the MIT NANDA report’s claim that roughly 95% of generative AI pilots showed no measurable P&L impact, and it is frequently misread as a technology failure rate.8 The report is not peer-reviewed, its dataset has not been fully released, and its own limitations section lists inconsistent success metrics across organisations and a six-month observation window that may be too short.9

    Read carefully, its central finding is about measurement: a pilot without a documented pre-deployment baseline cannot show impact regardless of how well it works.10 That is not an argument against AI. It is an argument for naming the comparison before you build.

    Ask one question: does the outcome already happen?

    One question, four answers. Each answer points at a different place to find the comparison, and only one of them lets a use case past this gate without a number.

    AnswerWhat it meansWhere the comparison comes from
    MeasuredIt happens, and a system records itTwelve weeks of history from the system of record
    UncountedIt happens, but nothing records itConstruct one: a scored sample of recent cases, or shadow mode
    AbsentNobody does this today, in any formThe problem it leaves unaddressed, priced from a figure finance already reports
    PlatformIt has no business outcome of its own. It is what other use cases will be built onNothing. Funded as R&D, capped and dated, never counted as return

    Uncounted is bigger than it looks, and this is where most arguments end. If a person does the work today on paper, in a spreadsheet, or through a workaround nobody documented, then the outcome happens. Absence of a system of record is not absence of a process.

    Most proposals that arrive claiming no baseline are uncounted rather than absent. The difference decides whether you spend four weeks constructing a comparison or go looking for the problem instead. Only the third answer is genuinely different, and it is rarer than any pipeline suggests, because teams reach for novelty when it sounds like an exemption from measurement.

    A proposal that spans two answers gets split before it proceeds. Each half has its own comparison and usually its own owner.

    Measured and uncounted: you already have a comparison

    For measured work, the recipe is short. Name the process, name the metric, pull twelve weeks, lock it in writing before anyone writes code.

    For uncounted work there are two ways to build one, and the first is cheaper than teams expect:

    • Score a historical sample. Take 200 recent cases, agree a rubric with the business owner, and have experts score them. That is a baseline built from records you already hold, available in days, with no waiting for a pilot.
    • Run shadow mode. The model runs alongside the humans on live work, both are scored, and nothing it produces reaches a customer. Slower, but it also proves the integration before it matters. Regulators tend to want this sequence anyway.

    Quality work carries a prerequisite the schedule usually forgets: agreement between the people scoring. If two senior underwriters disagree about what a good decision looks like, there is no quality metric to improve, and no model can beat a target nobody can define. Measure agreement first. If your experts concur 60% of the time, that figure is the ceiling, and learning it in week two costs almost nothing.

    Then price the errors. “Quality improved 12%” means nothing to a CFO. A missed fraud signal and a mis-coded claim cost different amounts, so sort errors into classes and attach a figure to each. That is the step that turns a quality score into money.

    One shortcut before building anything: rework rate, reopened claims, appeals, complaint volume, and audit findings are quality signals the business already reports. Blunt, but free, and finance already trusts them.

    Absent: measure the problem, not the process

    When the process has no baseline, the problem does. That is the whole rule. If nobody reviews subrogation on small claims, the comparison is not some earlier review process, because none existed. It is the money leaking now, and leakage is already on a report. If nobody answers the phone at 2am, the comparison is what happens to those claims today: they wait until morning, and that wait is measured.

    Four methods, cheapest and most auditable first:

    • Backtest against closed cases. Run the model over 12 to 24 months of closed files and count what it finds that was missed. For fraud, subrogation, or missed coverage, every hit is checkable against a file with a known outcome. This produces a recoverable-dollars figure before a single live decision, and it is the strongest answer available to “we have no baseline”.
    • Test volume before value. Before asking what after-hours intake is worth, count how many claims arrive between 10pm and 6am. That number is already in the claims system. If it is four a week, the business case dies for free.
    • Price the alternative, honestly. What would this cost with people? The figure only counts if leadership would genuinely have approved that headcount. If they would never have staffed it, the avoided cost is fictional and finance will strike it.
    • Commit in tranches. Absent work carries more uncertainty than automation, so fund it in stages tied to evidence: backtest, then a limited population, then rollout, each with its own threshold.

    If neither the process nor the problem can be measured, you are not missing a baseline. You are missing evidence that the problem exists, and the proposal belongs in the platform column below or nowhere at all.

    Platform: how capability building turns into return

    Platform work has no return of its own, and the honest thing is to say so. Its return arrives through the measured, uncounted, and absent use cases it makes cheaper, faster, or possible. Finance already knows this shape: nobody asks the integration layer for its ROI. They ask what it costs, what it is allocated to, and whether the things built on it pay back.

    That makes it measurable indirectly, on four numbers:

    MetricWhat it shows
    Time from approved use case to productionWhether the platform actually shortens delivery
    Cost to ship the next use caseWhether shared components are reused, comparing use case one to use case five
    Reuse count per componentA component used once is a project cost, not a platform
    Graduation rateHow many platform bets produced a use case with a signed pre-build agreement

    Then allocate the cost rather than leaving it floating. Every graduated use case carries its share of platform spend in its Gate 3 run-rate, so a use case that only pays back by ignoring the platform it runs on has not paid back. The programme total, capability included, is the board-level view. That is what JPMorgan’s $2 billion in and $2 billion out actually is: a portfolio number, not a use case number.

    Some of what platform spend buys never converts into a figure, and should not be forced to. The organisation learns what agents can be trusted with, the identity team learns to govern non-human principals, the first regulator conversation goes well. Name that as the reason the budget exists, then keep it out of the return line. Pretending it has a dollar value is what gets a programme’s real numbers discounted.

    Gate 2: Can you name one value type and one primary metric?

    AI proposals tend to promise everything: faster, cheaper, better, safer. Finance will discount all four to zero if none is primary. Choose one:

    Value typeExample primary metricHow finance treats it
    Cost avoidanceCost per resolved case; hours redeployedAccepted if redeployment is real and tracked
    Revenue liftConversion rate; win rate; retentionAccepted with a control; heavily discounted without
    Risk reductionError rate; loss ratio; audit findingsAccepted if tied to a quantified exposure
    Cycle-time compressionDays from submission to decisionAccepted only when converted to a downstream dollar effect
    New outcome createdCases handled that previously were notAccepted only when tied to a downstream dollar metric the business already reports

    Two value types deserve honest treatment because they are the ones most often oversold. “Productivity” that never reaches a decision, a deliverable, or a headcount plan is not a metric. It is a feeling, and METR showed how unreliable that feeling is. “Capability building” or option value is real, but Gate 1 already sorted it into the platform column: an explicit budget, and no line in the return.

    The most-cited public number comes from JPMorgan, which reports roughly $2 billion of annual cost savings against roughly $2 billion of annual AI spend, with the savings spread across hundreds of use cases rather than one headline project.11

    Notice the value type: the bank reports cost savings, not “productivity”. A number stated in one value type, at the level of the use case, is one finance can interrogate. A blended enterprise productivity figure is not. The bank’s CEO also puts the portfolio in proportion: almost 1,000 use cases, of which “the really important ones are 50”.12

    Gate 3: What is the fully loaded cost, and what is the run-rate at scale?

    A pilot cost estimate that does not include a production run-rate is not an estimate. The cost lines that get left out are consistent:

    • Inference at production volume, not pilot volume, with a stated assumption about token growth as prompts and context windows expand.
    • Integration with systems of record, identity, and audit logging.
    • Data readiness: cleaning, lineage, access controls. Usually the largest and least visible line.
    • Evaluation infrastructure: the harness that tells you the model still works next quarter.
    • Governance and change management: review workflows, training, policy.
    • Drift maintenance: re-evaluation, prompt and model updates, regression testing.

    The accounting treatment belongs in the agreement too, because it changes the payback period a CFO sees. Under US GAAP, proprietary AI built for internal use under ASC 350-40 can create a capitalisable asset. Licensing API access to a foundation model is typically a period cost. Buying a platform may be a capitalised intangible, a SaaS expense, or a hybrid.13 Internally generated data costs are generally expensed as incurred.14

    The FASB amended this area in September 2025 with ASU 2025-06.15 Organisations reporting under IFRS work from IAS 38, where research expenditure is expensed and development expenditure that meets the standard’s criteria is capitalised, so the same judgment applies.16 Your controller will have a view; get it before the pilot, not at year-end.

    The run-rate also carries this use case’s allocated share of shared platform cost, per the platform rule at Gate 1. A payback calculated on marginal cost alone flatters every use case after the first.

    Gate 4: Can it be tested against a control?

    The reason a control is not optional is that the same AI tool produces opposite results depending on the task. A field experiment in a public-sector setting found generative AI improved quality by 17% and cut completion time by 34% on a document comprehension task. On a data task it reduced quality by 12% with no time saving.17

    A study of seven generative AI deployments at an online retailer found five delivered measurable gains. It measured revenue outcomes rather than input-side efficiency, with the control group reflecting standard practice before adoption.18 Field experiments at Microsoft and Accenture found developers with an AI coding assistant completed roughly 8–22% more pull requests per week, though low compliance made the estimates imprecise.19

    Without a control, you cannot distinguish the AI’s effect from the seasonal trend, the concurrent process change, or the Hawthorne effect of being watched.

    The design follows the Gate 1 answer. Measured work compares before and after against the historical baseline. Uncounted work needs shadow-mode data or a scored control sample, because the comparison is decision quality rather than elapsed time. Absent work has no before at all, so it compares populations: the cases that got the new capability against those that did not, over the same period.

    Acceptable designs, in order of strength: randomised assignment of cases or users; phased rollout by team or region with the later cohorts as controls; matched comparison against a similar unit that has not adopted. If none is feasible (some processes are single-team and cannot be split), the use case may proceed, but its ROI is labelled unverified in the portfolio view and is never aggregated into the enterprise number as if it were measured.

    Gate 5: What is the kill criterion, and when is the review?

    Every use case needs an exit. Before build, write the threshold below which the use case is stopped, the date on which that judgment is made, and the person who makes it. Kill criteria are what convert a portfolio of hopeful pilots into a managed investment.

    This is also where the portfolio rule lives. Leadership should judge the programme by the distribution of outcomes and by kill velocity (how quickly failing use cases are stopped) rather than by any single success story.

    McKinsey’s survey shows why: respondents commonly report cost reductions at the function level even where enterprise-wide EBIT is unchanged, and about one in five says operating costs are already limiting AI use.2 Local wins do not sum to enterprise return unless something is stopping the losers.

    Platform bets get a second question at the same review, alongside the threshold: has this capability graduated anything? A platform with no use cases through the pre-build agreement after two review cycles is not building optionality. It is cost with a story attached.

    The pre-build agreement

    Before the first sprint, three questions are answered in writing by whoever owns finance and delivery in your organisation. No template is required; a page in the project record is enough. What matters is that the answers exist before money is spent.

    1. Will finance accept this metric as return?

      The Gate 1 answer and the comparison method it requires, the primary metric from Gate 2, the target, and the value type. Finance signs off on the method, not only the number: a constructed or zero baseline agreed now is not reopened at the review. If the answer is “we’ll see,” the metric is wrong.

    2. Can delivery capture this metric from the systems involved?

      Engineering confirms the metric can be instrumented, names the source system, and confirms the control design from Gate 4 is feasible. If the metric lives in a spreadsheet nobody maintains, this question fails.

    3. What stops the project, and when?

      The kill criterion, review date, and decision owner from Gate 5, plus the production run-rate from Gate 3 that the review will be judged against.

    Worked example: claims intake at a P&C insurer

    Consider a mid-sized property and casualty insurer evaluating two generative AI use cases.

    Use case A: first-notice-of-loss summarisation. When a claim is reported by phone or web form, a model produces a structured summary and suggested triage category for the adjuster.

    • Gate 1: Measured. The claims system records time from FNOL to first adjuster action, and twelve weeks of history show a median of 2.1 business days.
    • Gate 2: Value type is cycle-time compression, converted to a dollar effect through a known relationship between early contact and indemnity cost. Primary metric: median FNOL-to-first-action time. Guardrail: triage reassignment rate must not rise.
    • Gate 3: Pilot cost covers one line of business. Production run-rate models all lines at annual claim volume, including integration with the claims system, an evaluation harness that scores summaries against adjuster corrections, and quarterly re-evaluation. The platform subscription is a period cost; the integration build is capitalised.
    • Gate 4: Claims are randomly assigned to summarised and non-summarised queues within the same team for eight weeks. Volume is sufficient to detect a half-day improvement.
    • Gate 5: If median time does not improve by at least 0.5 days, or reassignment rate rises by more than two points, the use case stops on the review date. Decision owner: VP Claims.

    Finance accepts the metric. Claims technology confirms it can be captured. Build proceeds.

    Use case B: an internal “ask anything” policy assistant for underwriters. A chat interface over underwriting guidelines.

    • Gate 1: The proposal arrives claiming there is nothing to compare against, with the stated benefit “underwriters will be more productive.” The question splits it in two. Underwriters do look up guidelines today, on paper and by asking each other, so faster lookup is uncounted: time a sample of 40 referrals over two weeks and the baseline exists. The other half is different. Answers to questions underwriters never asked, because asking took too long, is absent, and the problem it leaves behind is already reported: referral rate to senior underwriters is where people land when they give up.

    Split that way, the gate produces two comparisons and a path rather than a refusal. Capture starts now, on the current population, and the proposal returns in four weeks with numbers instead of an adjective. It also inherits the evaluation harness built for use case A, which is what the platform spend on that first project was for.

    The difference between this and “underwriters will be more productive” is not enthusiasm. It is that one of them can be wrong, on a date, in front of a named decision owner.

    A closing caution, and where to start

    Even a well-measured programme may buy parity rather than advantage. On JPMorgan’s July 2026 earnings call, Jamie Dimon told analysts “you don’t uniquely benefit from AI. The ultimate beneficiary of AI will be our customers”, because competitors reach the same capability and a bank cannot simply keep the margin.12

    That is not a reason to skip measurement. It is a reason to measure honestly: the return that survives a control group and a kill criterion is the only return worth reporting to a board.

    Start with what is already in flight. Take the three AI initiatives currently consuming the most budget and run each through the five gates this week, on one page each. At least one will fail Gate 1, and that page is the most useful thing your programme will produce this quarter.

    References

    1. The Register, “McKinsey says enterprise AI is finally ‘on the road to ROI’”, August 2026.

    2. McKinsey & Company, “The State of AI: Global Survey 2026”, August 2026. 2

    3. METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”, July 2025; update, February 2026.

    4. Deloitte Switzerland, “A CFO’s Guide to AI Value Realisation”, July 2026.

    5. Gartner, “10 Best Practices for Optimizing Generative and Agentic AI Costs,” as reported by Campus Technology, June 2026.

    6. Gartner press release, “Gartner Predicts GenAI Cost Per Resolution for Customer Service Will Exceed Offshore Human Agent Costs by 2030”, January 2026.

    7. KPMG, “CFOs and CIOs: Partnering for innovation”, KPMG 2025 CFO & CIO Collaboration Survey, March 2025.

    8. MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025,” July/August 2025.

    9. AI Operator, “Do 95% of AI Projects Fail? What the MIT Study Actually Measured”.

    10. Agentmode AI, “MIT 95% AI pilot failure: what the GenAI Divide measured”, May 2026.

    11. Investing.com, “Jamie Dimon: JPMorgan spends $2 billion yearly on AI, saves same amount”, October 2025.

    12. JPMorgan Chase Q2 2026 earnings call, 14 July 2026, transcript via Yahoo Finance; see also PYMNTS, “Dimon Says AI Will Make Banking Better and Tougher”, July 2026. 2

    13. Embark, “AI on the Books: Accounting for AI Costs”, June 2026.

    14. EisnerAmper, “Accounting for AI Data and Consumption Cost”, July 2026.

    15. Deloitte DART, “FASB Amends Guidance on the Accounting for and Disclosure of Software Costs”, September 2025.

    16. IFRS Foundation, “IAS 38 Intangible Assets”: “Research expenditure is recognised as an expense. Development expenditure that meets specified criteria is recognised as the cost of an intangible asset.”

    17. “Assessing Generative AI value in a public sector context: evidence from a field experiment”, arXiv, February 2025.

    18. “Generative AI and Firm Productivity: Field Experiments in Online Retail”, arXiv, 2025.

    19. “The Productivity Effects of Generative AI: Evidence from a Field Experiment with GitHub Copilot”, MIT GenAI, 2024.

    Comments

    Comments are GitHub discussions. Sign in with GitHub to post; reactions need no account.