AFPM — AI Flow & Performance Monitoring
The New Performance Metric: Cost per Successful Outcome
Counting how much computing AI uses is not enough; the number that matters is how much it cost to get the right result.

Two taxis leave the same hotel for the same airport. The first driver takes the motorway and arrives in thirty minutes for 40 dollars. The second gets lost, circles the city twice and arrives an hour later for 140 dollars. Both passengers made their flight. Nobody would say the two trips performed equally well.
Yet this is roughly how many organizations measure AI today. They check whether the task was completed. They rarely check what it took to complete it.
From counting resources to counting results
Traditional cloud optimization often focuses on infrastructure consumption. Teams ask how much CPU they used, how much storage they are paying for and how many API calls their applications made. An API call is simply one piece of software asking another for something, and cloud providers often charge for each one.
These are sensible questions. They helped a generation of companies cut their cloud bills. But they measure effort, not results. They tell you how much fuel the taxi burned, not whether it took the right road.
AI introduces a far more important metric:
How much did it cost to achieve the correct outcome?
This question joins two things that are usually measured by different teams. The finance team sees the bill. The business team sees the results. Cost per successful outcome puts them in the same sentence.
Two agents, one task
Consider two AI agents performing the same task. Agent A consumes 70,000 tokens and completes the task successfully. Agent B consumes 9,000 tokens and reaches the same result.
Both are technically successful. Operationally, they are very different. Agent A may be reading documents it does not need, repeating its instructions to itself or thinking in circles before it answers. Agent B goes more or less straight to the point.
On a single task, the difference is small change. At scale, it is not. And the picture becomes more interesting once you add the fact that not every attempt succeeds.
The lesson is not that big agents are better. In another situation, Agent B might be improved until it fails rarely, and then it would win easily. The lesson is that you cannot know which approach is cheaper until you measure cost and success together.
The model choice problem
The same problem appears with model selection. A model is the AI engine that does the thinking. Providers offer a range of them, from large, powerful and expensive models to small, fast and cheap ones. The price gap between the top and the bottom of the range can be very large.
A powerful reasoning model may be necessary for one stage of a workflow but completely unnecessary for another. Using the most expensive model for every step can dramatically increase operational costs without improving the final result.
Think of a law firm. A senior partner’s time is valuable, and some decisions truly need it. But nobody asks the senior partner to photocopy documents or book meeting rooms. Many AI workflows do the equivalent: they send a simple sorting or formatting task to the most expensive model available, because that is the model the system was built with.
What to measure instead
This is why AI performance needs to be correlated with outcomes. Correlated simply means measured side by side, so that you can see how one changes when the other does. A few measurements make this possible. Each is simple on its own. Together they give a far clearer picture than a monthly bill.
- Tokens per completed task
- How much text the AI processed, on average, to finish one real piece of work.
- Cost per successful workflow
- The full cost, including failed attempts, divided by the number of results that were actually right.
- Model cost by process stage
- Which steps of a process use which models, and what each step costs.
- Retry cost
- What the organization pays for second and third attempts after something did not work.
- Failed reasoning cost
- Spending on AI work that ended without a usable result.
- Tool-call efficiency
- Whether the agent takes a direct path or makes many unnecessary lookups and actions.
- Human intervention frequency
- How often a person has to step in, and how much of their time that takes.
- Routing strategy comparison
- The cost difference between sending work to one model or another for the same result.
The objective is not simply to minimize token usage. A process that saves tokens by failing more often is not cheaper. It has moved the cost somewhere less visible.
The objective is to remove wasted intelligence.
Wasted intelligence is thinking that did not need to happen: the long document read for no reason, the expensive model used for a trivial step, the fifth retry of an approach that failed four times already.
What this means for your organization
Define success for one process. Pick an AI process with a clear result, such as a correctly categorized support ticket. Agree on what counts as success before you measure anything.
Put cost and outcome in the same report. Ask the team that owns the AI bill and the team that owns the results to produce one number together: total cost divided by successful results.
Include the human cost of failure. Every time a person has to fix or finish AI work, that time belongs in the calculation.
Look at model choice step by step. List the steps in the process and the model each one uses. Ask whether each step truly needs that model.
Compare before you cut. Before switching to a cheaper model or a shorter prompt, run both versions side by side on the same work and compare cost per successful outcome, not cost per attempt.
Tracston works on these questions in Observatory.



