Back to insights

AI Investment Decisions

Before You Approve an AI Pilot, Ask for This Decision Record

Share on LinkedIn
Before You Approve an AI Pilot, Ask for This Decision Record, beside a paper checklist for value, authority, evidence and ownership.

A pilot should buy evidence for a decision. Give the team a clear question, a spending boundary and permission to conclude that a simpler approach works better.

Executive Summary

AI Investment Decisions | Published September 23, 2026

An AI pilot deserves approval when it can resolve a useful business question within a boundary the organisation can sustain. Record the outcome, alternative, evidence, authority, owner, full operating cost and stop conditions before the trial starts. Then compare complete, accepted outcomes. The fictional example below shows why an AI trial can succeed as an experiment while losing the investment decision to ordinary automation.

Decision Approve a defined experiment with an expiry date, not an open-ended commitment to a product.
Economics Compare AI, a simpler alternative and the current process using the same outcome and full cost boundary.
Accountability Name who accepts the work, who can suspend the trial and what evidence permits the next step.

Key takeaways

  • Keep the approval record to one page. Link the evidence behind it rather than hiding assumptions in a long presentation.
  • Separate permission to test from permission to act on customers, money or production records.
  • Measure fallback, review and rework alongside model and platform charges. Keep implementation spending visible separately.
  • Define success and stop conditions before seeing the results. A pilot may justify a narrower use case, ordinary automation or no further investment.
  • Use the downloadable blank record with the worked example, then replace every assumption with evidence from your own workflow.

What are you actually approving?

The difficult moment in an AI proposal often arrives after the demonstration. People can see the potential. The vendor has answered the architecture questions. Someone asks for a modest pilot budget. Yet the room still has different ideas about what approval means.

The sponsor may think the team will test whether service improves. Technology may think it has permission to integrate a new platform. Operations may expect the existing team to review every result. Procurement may be committing to minimum usage. Each interpretation can sound reasonable while producing a different cost, exposure and workload.

My starting point would be a short decision record, written before the pilot begins. It should say what uncertainty the business is paying to resolve and which decision the evidence will support. For example: can supervised quotation preparation reduce total handling effort without changing who approves price or commits the business to delivery?

That question makes the experiment testable. It also allows a useful negative result. If a rules-based workflow produces the same benefit with less review, the pilot has provided valuable evidence. The mistake would be treating adoption of the AI product as the only acceptable outcome.

This guide builds on how to defend an AI agent system to engineers and executives. That article examines the evidence needed to justify a system. Here the focus is the earlier spending decision: a reusable approval record, a completed example and the arithmetic that should determine what receives the next cheque.

A well-run AI pilot can end with a decision to buy no AI at all. If the sponsor cannot accept that result, the budget is funding a preferred answer instead of an experiment.

Michel Junior Julien

The one-page AI pilot decision record

Use the record as the cover sheet for the review. Each field should contain a short answer and, where needed, a link to the underlying evidence. A blank or disputed answer is useful information. Do not convert uncertainty into an optimistic assumption simply to finish the page.

Open the printable one-page approval record or download the editable Markdown template. The framework is an original JM Digital working tool. It is not a certification or a substitute for the organisation’s existing approval obligations.

  • 1. Decision requested. State the permission and funding being requested now, its expiry and the decision that will follow. Separate discovery, supervised testing and production operation.
  • 2. Outcome and eligible scope. Name the business result, population and exclusions. Define a completed, accepted outcome in terms an operator can verify.
  • 3. Current process and alternative. Record today’s quality, time and cost using representative work. Include the simplest credible improvement that does not require AI.
  • 4. Evidence and readiness. Identify authoritative inputs, access approval, unresolved dependencies and the owner of each gap. Say which gaps prevent the test from starting.
  • 5. Authority. List permitted reads and actions, prohibited actions, approval points and where enforcement happens. A written promise is insufficient without matching permissions.
  • 6. Model and configuration. Identify the selected model or service, version policy, routing, tools and configuration record. State which changes require retesting.
  • 7. Evaluation. Define case selection, reference answers, critical failures and acceptance thresholds. Assign someone outside the build team to challenge the results.
  • 8. Ownership and capacity. Name the sponsor, operator, technical owner and exception owner. Confirm their available hours and who takes over when someone is absent.
  • 9. Cost boundary. Show the trial spending cap, implementation effort, recurring charges, human review and fallback costs. Record the unit used for comparison.
  • 10. Stop and recovery. Specify observable suspension triggers, who can invoke them, how existing work is reconciled and what permissions are removed.
  • 11. Outcome decision. Predefine when to expand, narrow, fix or stop. Explain which conditions are mandatory rather than averaged into a favourable score.
  • 12. Approval and review. Record the named approvers, date, conditions, evidence location and next review. Preserve earlier versions so later changes remain visible.

A single page is a discipline for the decision, not a limit on investigation. Security evidence, workflow maps and evaluation results may be substantial. The cover sheet should let the sponsor see whether those documents support the actual authority being requested.

Resolve the prerequisites before granting access

Readiness starts with a small number of concrete facts. Can the team obtain the necessary records lawfully and through approved access? Can it identify the authoritative policy? Is there someone available to judge the result? Can the existing process continue when the pilot is unavailable? If those answers are unknown, the first budget should fund discovery or remediation.

Classify every important input as verified, assumed or unresolved. A catalogue export with a timestamp is inspectable evidence. A statement that the ERP has everything we need is an assumption until the fields and permissions have been checked. This distinction helps a sponsor fund the next useful action without approving the entire proposed architecture.

The voluntary NIST AI Risk Management Framework connects understanding an intended use with go/no-go decisions and context-specific measurement. The record here applies that general discipline to one investment decision; completing it does not establish compliance with the framework.

Permissions should reflect the experiment. If the trial only needs to prepare an internal quotation brief, access to send customer messages or update prices adds responsibility without helping answer the question. Restrict the service identity and operations at the system boundary. Do not ask a prompt to compensate for an unnecessarily powerful connector.

OWASP’s excessive agency guidance recommends limiting functionality and permissions, enforcing downstream authorisation and requiring approval for consequential actions. In practice, test attempted boundary crossings as well as normal requests. Record the evidence that a disallowed action was blocked.

Human review also needs a design. The reviewer must be able to see the source facts, understand the proposed outcome and decline it without creating a separate investigation. Allocate the time in the staffing plan. An unnamed human somewhere in the loop is neither a capacity commitment nor a dependable control.

Compare complete outcomes from the first trial

There is no need to wait three years for a useful cost analysis. Start with a representative sample and the current process. Run the candidate workflow against the same acceptance criteria. Include a simpler alternative where one exists. A short comparison can expose an unattractive mechanism of value well before a large commitment.

Choose a denominator that the business recognises. For quotation preparation, that might be an accepted internal brief containing the correct items, quantities, relevant account restrictions and unresolved questions. A model response, a tool call and a drafted document are activities. None necessarily completes that outcome.

Then calculate operating cost per accepted outcome: relevant service and infrastructure charges, plus staff handling, review, fallback, rework and support, divided by accepted outcomes. Avoid double counting time recorded in more than one category. State how shared costs are allocated, and retain the total work volume so excluded or failed cases remain visible.

Keep implementation spending alongside this operating calculation. Hiding setup costs exaggerates the economics; loading every experimental cost into a tiny sample can obscure the steady-state mechanism. Show both views so the sponsor can distinguish what the pilot costs from what a repeatable service might cost.

Released staff time is initially capacity, not automatically cash savings. Identify what useful work the team would do with it, or which external cost would actually disappear. Faster quotations may have commercial value, but do not count speculative revenue without a way to isolate and measure that effect.

A completed example: preparing a distributor’s quotation brief

The organisation, quantities, costs and results in this example are entirely fictional. They illustrate a decision method, not a client engagement, market benchmark or expected return. All monetary amounts are Canadian dollars. The assumptions would need replacing with measured evidence before use in a real proposal.

Imagine a distributor whose sales coordinators turn incoming requests into internal quotation briefs. They identify products, resolve quantities and pack sizes, check account restrictions and flag missing information. An authorised employee sets the final price and sends the quotation through the existing process. The proposed AI would prepare the brief, not negotiate or commit the company.

Fields 1 to 3: the decision, outcome and comparator

Request: a supervised, 30-day offline comparison with a total cap of $12,000, including setup and trial operation. Approval expires after the review. The question is whether AI reduces the cost of an accepted brief more than a structured intake form and rules-based product lookup.

Scope: 200 previously handled, permitted requests from an agreed catalogue and account population. Exclude new products, negotiated pricing and unsupported document types. An accepted brief must match the verified product, quantity and account facts, and explicitly flag anything missing. No customer contact is authorised.

Baseline: a fictional eight minutes of coordinator work per request. Compare the current method, rules-only preparation and AI-assisted preparation on the same cases. Time spent validating the reference answers is evaluation work, tracked separately from routine operating effort.

Fields 4 to 6: readiness, authority and configuration

Evidence: the product owner confirms the dated catalogue, pack-size rules and account restrictions. Operations verifies the historical answers and removes records outside the approved data scope. Conflicting product identifiers must be resolved or represented as deliberate exception cases before testing.

Authority: read access to approved test copies, plus permission to write drafts only in the evaluation workspace. Price updates, customer messages and production changes are unavailable. Each draft goes to a coordinator. The technical owner demonstrates that the test identity cannot reach those prohibited operations.

Configuration: a fixed model configuration, one retrieval path and a recorded prompt version. Routing is disabled for this comparison so changes in behaviour can be investigated. Any model or tool change creates a new test version rather than being silently pooled with the original results.

Fields 7 to 9: evaluation, people and money

Evaluation: operations creates the reference set; a second reviewer resolves disputed labels. Rotate the method order to reduce familiarity effects. Include unclear quantities, conflicting identifiers, missing restrictions and requests that should be refused. Report routine and difficult cases separately, alongside the whole set.

Acceptance: at least 95% of eligible briefs must be accepted on first review, with no critical account-data exposure or unauthorised action. Correctly flagging a missing fact counts as success when the reference answer requires that flag. These are invented case-specific thresholds, not general standards.

Owners: the sales operations manager sponsors the trial and accepts the outcome. A technical lead owns configuration and access. Two coordinators receive allocated review time and cover each other. The operations manager owns the exception queue; finance reviews cost categories before the first run.

Budget: proposed setup allocations of $6,000 for the AI path and $2,500 for the rules path leave $3,500 for trial operation and evaluation within the $12,000 cap. Staff time uses a declared $60 per hour loaded rate. Both paths record paid usage, support and manual fallback.

Fields 10 to 12: suspension, next decision and sign-off

Stop: suspend immediately if restricted data reaches an unapproved destination, the tool boundary fails or a prohibited action becomes available. Pause if projected spending exceeds the cap or the reviewers cannot keep up. The technical lead revokes access, preserves permitted evidence and reconciles outstanding drafts.

Next decision: only consider an operational rollout if the quality threshold, control tests, capacity commitment and economic comparison all pass. An AI path must justify its incremental cost against rules-only preparation. A narrower follow-up experiment requires a fresh scope and budget.

Approval: named sponsor, technical owner and data owner sign the dated record; finance confirms the cost basis. This fictional example uses roles for illustration. A real record needs accountable people, an evidence location and a calendar review date before access is granted.

The results: AI saves time, but does it win?

Suppose the fictional trial produces the following results. The current process delivers all 200 accepted briefs with 1,600 minutes of handling. At the assumed $60 hourly rate, its operating cost for this comparison is $1,600, or $8 per accepted brief. Common costs that do not change across the methods are excluded consistently.

The AI path produces 170 briefs accepted on first review, taking three minutes each to inspect and complete. Thirty require a full eight-minute manual fallback. Additional attributable quality checks and rework consume 120 minutes. Paid model, platform and allocated service costs total $180. The categories are mutually exclusive; fallback time is not also counted as successful review time.

AI operating cost: (170 × $3) + (30 × $8) + $120 + $180 = $1,050. All 200 outcomes are eventually accepted after human handling, so the completed-outcome cost is $5.25. First-review acceptance is 85%, below the pre-agreed 95% gate. The final completion count cannot erase that failure.

The rules-only path produces 190 briefs accepted on first review at four minutes each, with ten eight-minute fallbacks. Additional quality checks consume 60 minutes and paid service costs total $100. That gives (190 × $4) + (10 × $8) + $60 + $100 = $1,000, or $5 per accepted brief, with 95% accepted on first review.

Both approaches improve on the fictional baseline. AI saves $550 of operating effort and fees across the sample; rules-only saves $600. The AI setup allocation is also $3,500 higher. On this evidence, approving the broad AI workflow would require overlooking both its quality miss and the better-performing alternative.

There may still be a useful AI opportunity inside the difficult requests. Investigate that subset before dismissing it or funding it. Did language interpretation solve a problem the form could not? Did it reduce a costly delay? The aggregate result provides no automatic answer. A narrower experiment should test that specific hypothesis with its own comparator.

Make the decision now, then test its sensitivity

The immediate recommendation in this example is to decline the broad AI rollout and consider a limited rules-only rollout, subject to its own operational approval. Preserve the reusable catalogue and evaluation work. If the sponsor wants further AI testing, require a bounded proposal for the ambiguous-request subset rather than extending the original pilot by default.

For planning only, multiplying the observed unit costs by 1,000 comparable requests gives $8,000 for the existing process, $5,250 for AI and $5,000 for rules-only. That is arithmetic under unchanged assumptions, not a forecast. The proportions of difficult cases, review burden, supplier terms and support effort may change with volume.

Before committing, vary those assumptions. What happens if manual fallback doubles? If usage charges increase? If the people receiving the time savings are already constrained by another step? Estimate a plausible range and identify which variable changes the recommendation. A decision that reverses after a small change deserves more evidence or a smaller commitment.

The record should also identify measurement limitations. Two hundred cases cannot establish the rate of a rare failure. A historical sample may miss new account rules or seasonal requests. A carefully supervised trial may understate everyday support needs. Those limitations determine the next boundary; they do not require postponing every commercial judgement.

Approval has an expiry date

A signed record represents evidence at a particular point in time. New data sources, a different model, broader permissions or a changed customer population can invalidate parts of it. Assign someone to decide whether a change needs targeted retesting, a revised operating procedure or renewed approval.

Distinguish stopping new work from repairing work already completed. Revoking a connector can prevent further drafts, but it does not locate every item awaiting review. For a workflow that eventually performs transactions, disabling the agent does not reverse a payment or a message. The recovery plan needs an inventory of affected work and an authorised owner for resolution.

Keep the decision record alongside configuration identifiers, evaluation results, incidents and review dates in a location the business can access. Restrict sensitive evidence appropriately. The sponsor should be able to understand why authority exists without finding the developer who ran the original demonstration.

The same discipline applies when a vendor runs much of the platform. Ask for evidence you can inspect and for the contractual and operational boundaries that support the proposed scope. A supplier can provide capable software while your organisation still lacks the data ownership or staff capacity needed to operate the workflow.

Bring the record to the next funding discussion

Ask the proposal owner to complete the record with operations and finance before the meeting. Circulate the unresolved fields as questions for decision, rather than presenting them as minor delivery details. Start the discussion with the requested scope, total cost boundary and evidence that would change the recommendation.

For teams exploring AI consulting in Toronto and the GTA, that may mean reconciling the source facts across commerce, service and operations before selecting a pilot. For an Ottawa AI consulting engagement, a useful starting point may be the authority and review burden behind a document or proposal workflow. The record adapts to the work; geography does not alter the evidence required.

An AI readiness assessment can turn unresolved fields into a practical recommendation: proceed within a defined boundary, fix a prerequisite, choose a simpler alternative or stop. The work earns its value when leadership can make that decision with a clear view of the consequences.

At the end of the review, someone should be able to read the approved sentence aloud: we are funding this experiment, for these cases, under these conditions, until this date. If the room agrees on that sentence and the evidence behind it, the pilot has a sound starting point.

Related reading

Continue the production-readiness path

These connected JM Digital Corp insights add architecture, data, workflow, and delivery context around the AI series.

Read next How to Defend an AI Agent System to Engineers and Executives Read next AI ROI Is Being Measured Wrong: Cost Per Token Is Not Cost Per Outcome Read next AI Consulting and Readiness Assessment

Make the next AI investment decision concrete.

Bring one proposed workflow, its current process and the questions you cannot yet answer. JM Digital can assess the evidence, operating requirements and alternatives, then help you define a decision your team can act on.

Discuss an AI assessment Use the decision record