The assessment method
Measure the work. Then the savings.
An inference cost assessment compares the same workflow in your current setup and the Happiest Labs local AI platform. The result must explain cost, quality, performance, and which charges can actually change.
What does “up to 76% savings” mean?
Happiest Labs reports savings of up to 76% for existing customers. That reported result is context for a conversation, not an assumed reduction for your company. Your assessment starts with your own baseline and an agreed scope; it does not apply that percentage to your AI budget.
1. Establish the current cost baseline
Use a representative 30-day period and reconcile usage with invoices. Record the providers, model versions, effective rates after discounts, and the workflow associated with each charge. Separate input tokens, output tokens, cache reads and writes, tool charges, retries, and fallback requests where applicable. Avoid counting cached tokens twice when a provider includes them within another usage total.
Use request counts, task counts, concurrency, end-to-end latency, and error data to explain demand and performance. Keep subscription seats and contractual commitments separate from variable usage. For managed or self-hosted services, include the applicable service and operating costs attributable to the selected work using a documented allocation.
Check whether those 30 days represent normal demand. Record unusual events, seasonality, growth, and any missing data. This baseline describes observed costs; annual volume remains a separate assumption.
2. Agree the workflow and acceptance criteria
Select one workflow with a clear business owner and technical owner. Define the task inputs, required outputs, allowed tools, and what a successful result means. Agree a representative task set, including difficult cases, before comparing systems.
Set thresholds for task success, output quality, error rate, and performance under expected concurrency. Measure p95 end-to-end latency: the completion time at or below which 95% of tasks finish. Include tool steps, retries, and fallbacks in that duration. Agree any human review requirements as part of acceptance.
These are evaluation criteria, not a claim that every integration or workload is already supported. Confirm the platform capabilities and test scope with your team before the comparison.
3. Compare the same representative tasks
Run the agreed task set against the current setup and the local AI platform under comparable conditions. Record model and configuration versions, load, caching conditions, and any differences that affect the comparison. Use a held-out sample for the final check, and repeat trials where variability could change the decision.
Count failed attempts, retries, tool calls, fallback models, and human correction. Apply the same acceptance criteria to both systems. Report task success, p95 latency, errors, concurrency, and cost together; a cost reduction is not an acceptable result if required quality or performance fails.
This approach draws on Anthropic’s guidance on designing agent evaluations around tasks, grading, and repeated trials. Read the evaluation guidance.
4. Calculate cost per successful task
Cost per successful task = total applicable cost of all attempts ÷ tasks meeting the agreed acceptance criteria
Use the same cost boundary for both setups. Include failed attempts in the numerator and only accepted results in the denominator. If no tasks succeed, report that outcome instead of a unit cost. Document cost allocations and show the sample size, success rate, and test duration next to the figure.
Measured unit cost reduction = (baseline cost per successful task − platform cost per successful task) ÷ baseline cost per successful task
Calculate a percentage only when the baseline unit cost is positive and both setups meet the agreed quality and performance gates. Report the absolute costs alongside the percentage. A negative result means the platform cost more for that tested scope.
The FinOps Foundation recommends defining unit metrics around business outcomes and documenting their cost inputs. Cost per successful task is our application of that principle to this comparison. Read the unit economics guidance.
5. Reconcile projected and realizable savings
A measured test result describes that task set under those conditions. A production projection needs explicit assumptions about task mix, volume, peak concurrency, adoption, and operating cost. Do not multiply a small test result across the whole company budget without checking those assumptions.
Compare the full recurring baseline cost with the full proposed recurring cost for the same scope and period. Include Happiest Labs licensing, retained services, ongoing operations, support, and any change in review effort. Show one-time implementation and transition costs separately so finance can assess first-year savings and payback.
Identify which charges can fall, by how much, and when. Account for minimum commitments, renewal dates, retained subscriptions, and ramp-up. Keep a measured unit cost improvement, projected cost avoidance, and realized cash savings distinct. Actual savings may be lower than the test suggests, absent, or negative.
What to share first
Prepare the comparison with the blank audit worksheets. If the proposed change involves deployment location or data control, also complete the private AI deployment scorecard.
The request form asks for your company’s annual AI spend, primary inference setup, and optional high-level technical context. Do not paste invoices, logs, prompts, credentials, or confidential data into the public form. After the initial conversation, agree a channel and scope for any redacted billing and technical summaries needed for the assessment.