The working kit

Audit one workflow. Account for every attempt.

Use these free templates to connect your AI bill to work your team accepts, then decide whether a change can reduce actual expense. Start with your evidence, not an assumed savings percentage.

This kit is for technical and finance teams at US companies spending $1M or more annually on AI. The downloads are available to everyone, without a form. They contain column headers only: your team supplies the data and calculations in its own spreadsheet.

Download the three templates

1. Set the comparison boundary

Choose one production workflow and name its technical owner and finance reviewer. Record the evaluation period, task set, model and configuration versions, currency and cost boundary. Your company’s total AI budget is not the addressable spend of this workflow.

Use the same representative tasks, acceptance rules, trial count and operating conditions for both setups. Document load, concurrency, caching and allowed tools. Include difficult cases and repeat trials where outcomes vary. In this kit, a task trial is one assigned task with its permitted retries; it can contribute at most one accepted result.

Agree which costs belong in the comparison: model usage, tools, allocated recurring services and human review. Apply the same allocation method to both setups. Record exclusions and missing evidence. Leave unknown amounts blank; a missing charge is not a zero-cost charge.

2. Define acceptance before testing

In the scorecard, use one row per criterion. Set criterion_description, measurement_method and required_threshold before running the comparison. Mark mandatory requirements in critical_gate. Include task correctness, unacceptable errors, required human review and end-to-end response time at the agreed concurrency.

Use evaluation_id, workflow and task_set_version to connect the files. Each criterion_id identifies a stable check. Enter both setup results, the gate_decision, reviewer and an internal evidence_reference. A blank or failed mandatory gate means the candidate is not approved, even if its unit cost is lower.

3. Fill the cost ledger

Use one row per setup and evaluated task cohort. The identity columns record the evaluation, workflow, period, configuration and boundary. Set setup to baseline or candidate. Keep raw traces and invoices separately and reference them in evidence_reference.

task_trials counts assigned trials; accepted_task_trials counts those meeting your acceptance rules. total_attempts includes retries and failed attempts under those trials. Include their charges, plus fallbacks, in the applicable cost columns. Do not discard unsuccessful work or count each retry as a new successful task.

Enter model usage, tool/service charges, allocated recurring costs and human review costs without overlap. Reconcile effective billed rates and provider-specific cache accounting. Explain allocations in allocation_notes; use limitations for missing information or differences. Record quality and latency gate results, concurrency and p95 end-to-end latency: the duration at or below which 95% of trials finish. Document how timeouts are handled.

4. Calculate the observed result

Once costs are complete, manually sum the four cost columns into total_applicable_cost. These CSVs do not contain formulas or calculate an answer.

Cost per accepted task = total applicable cost of all attempts ÷ accepted task trials.

Acceptance rate = accepted task trials ÷ assigned task trials.

If there are no accepted tasks, unit cost is undefined. If there are no assigned trials, acceptance rate is undefined. Report those outcomes explicitly. Show absolute costs, counts, quality and latency alongside every comparison.

For comparable results that pass the required gates, calculate percentage unit cost reduction as (baseline unit cost − candidate unit cost) ÷ baseline unit cost × 100, only when baseline unit cost is positive. A negative result means increased cost. This describes the evaluated cohort, not a company-wide forecast. See the complete assessment methodology.

5. Make the finance decision

Use one finance row per cost component for the same workflow, currency and period. Identify it with decision_id; label cost_type recurring or one-time and evidence_class measured or projected. Enter baseline and proposed amounts, including retained services, licensing, operations and support. Put transition expenses on separate one-time rows.

removable_baseline_amount is the existing charge finance confirms can disappear within that period. newly_added_amount is incremental spending caused by the change, including transition charges. Record the earliest change date, contractual constraints or allocation basis, evidence, finance owner and approval status. Leave unconfirmed removal blank.

Calculate the period’s potential net cash reduction as removable baseline charges minus newly added charges, counting each once. Reconcile this with the proposed budget, including commitments that remain payable. A projection needs explicit volume, workload mix, adoption and operating assumptions. Lower unit cost, avoided future spending and staff time released are distinct from cash savings; realized savings require later billing evidence.

Use the result to choose the next step

Keep the current setup, investigate missing evidence or scope a pilot according to the gates. Our private enterprise AI guide covers deployment questions, and the enterprise cost guide explains workflow selection. A comparison with Happiest Labs considers its inference engine, runtime and harness together.

Request an assessment and demo → A high-level workflow summary is enough for the first conversation. Keep completed sheets, invoices, prompts and confidential data out of the public request form; agree a suitable sharing channel afterward.

Method references

This kit applies business-outcome measurement and evaluation principles; it is not a benchmark result. Read the FinOps Foundation’s unit economics guidance, Anthropic’s agent evaluation guidance and AWS’s workload-specific inference cost guidance.