AFPM — AI Flow & Performance Monitoring
Your AI Is Working. But Is the Process Working?
A system can look perfectly healthy while the work it does is slow, wasteful or wrong, so AI has to be judged by the whole process, not just by whether the machine is running.

Picture a busy restaurant kitchen on a Friday night. The ovens are hot, the fridges are cold, every cook is at their station and the ticket printer never jams. If you asked each piece of equipment how it was doing, the answer would be “fine”. Yet at table twelve a customer has waited forty minutes for a dish that was cooked three times and still arrived wrong. The kitchen was healthy. The meal was not.
That gap between a healthy kitchen and a good meal is the gap many organizations are about to discover in their use of AI.
The old question: is the system healthy?
For years, infrastructure monitoring has focused on a relatively simple question:
Is the system healthy?
To answer it, teams watch a handful of vital signs. CPU usage, memory, latency, database response time and application errors can tell us whether software is technically functioning. If the servers are not overloaded, the pages load quickly and nothing is throwing errors, the dashboard turns green and everyone moves on.
This way of thinking served us well for decades. When a website was slow, it was usually because a server was overloaded or a database was struggling. Fix the machine and you fixed the problem.
AI changes the question
AI changes the question. An AI-powered process may be completely healthy from an infrastructure perspective while still producing poor business results. The machine can be running perfectly and still do the wrong work, or the right work in a very wasteful way.
Three short examples show how this happens.
An agent can respond in 900 milliseconds and still make the wrong decision. Less than a second is an excellent response time by any technical standard. But if the agent approved a refund it should have refused, speed is not the point.
A customer-service flow can achieve 99.99% availability while sending users through unnecessary reasoning loops. The service never went down. Yet customers were asked the same question twice, or passed between steps that added nothing, before they got an answer.
A development agent can successfully execute hundreds of tool calls while consuming ten times more tokens than necessary. Every call worked. None of them failed. The job simply cost ten times what it should have.
None of these problems would show up on a traditional dashboard. This creates a new observability gap. Observability is the ability to understand what is happening inside a system from the outside, by looking at the signals it produces. Our current signals tell us about the machine. They say very little about the work.
What “the AI flow” actually means
Organizations must monitor not only infrastructure and application performance, but the AI flow itself. A flow is the whole journey of one piece of work, from the moment something starts it to the moment it is finished. Think of it as the full story of one order in the restaurant, not just the temperature of the ovens.
To follow that story, a team needs to be able to answer a set of plain questions about every piece of AI work:
- What initiated the interaction? A customer message, an employee request, a scheduled job or another agent. The starting point shapes everything that follows.
- Which model handled each stage? Many AI processes use more than one model. Knowing which one did what is the first step to knowing whether it was the right choice.
- What prompts were generated? A prompt is the instruction given to the AI. Much of it is often written automatically by software, not by a person.
- Which tools were called? Every lookup, calculation or action the AI took along the way.
- How much reasoning occurred? How many rounds of thinking the AI went through before it answered.
- What resources were consumed? Tokens, time and money.
- Where did retries happen? A retry is a second or third attempt after something did not work. A few are normal. Many are a warning sign.
- Was human intervention required? Did a person have to step in, approve, or fix something?
- Did the final result actually achieve the desired outcome? The most important question, and the one that is most often never asked.
Monitoring a model versus monitoring a process
This is the difference between monitoring an AI model and monitoring an AI-driven process.
Monitoring a model is like judging a relay race by timing one runner. It tells you something useful, but it cannot tell you whether the baton was dropped, whether the team took the wrong lane or whether they crossed the line at all. Real work in an organization is a relay. A request passes between software, several AI steps, data sources and people. Each of them can look fine on its own while the handovers between them go wrong.
As AI moves deeper into operational workflows, the flow becomes the real unit of performance. It is no longer enough to ask whether the model is fast and available. We need to ask whether the whole chain of work delivered the result it was meant to deliver, at a sensible cost, without unnecessary detours.
What this means for your organization
You do not need new technology to start thinking this way. You need a different habit of looking. Here are practical steps a team could take this week.
Pick one process that matters. Choose a single AI-powered process that customers or employees rely on, such as answering support questions or summarizing documents. One is enough to learn from.
Write down what success looks like. Not “the system responded”, but “the customer got a correct answer without being handed to a person”. If you cannot describe a good outcome in one sentence, you cannot measure it.
Follow ten real cases from start to finish. Take ten recent examples and trace each one by hand. Where did it start? How many steps did it take? Did anything repeat? Did a person step in? Was the result right? This small exercise usually reveals more than a month of dashboards.
Check what you can and cannot see. For each question in the list above, ask whether your current tools could answer it. The gaps are your observability gap.
Add one outcome measure next to the technical ones. Keep watching speed and availability. Add a single number that reflects the result, such as the share of requests resolved correctly on the first attempt.
Infrastructure tells us whether the machine is running. AI Flow Performance tells us whether the machine is doing useful work.
Tracston works on these questions in Observatory.



