Skip to content

Show what AI returns with business metrics

"How is AI improving our bottom line?" is hard to answer from an AI bill. The bill says what AI costs, not what it produced. A business metric counts the outcome and records how much of it AI handled without a person. It then puts the cost of the people and the AI beside that count.

This guide works through two examples. Use the one closest to your own work.

Before you start

How the pieces fit

Flowstate already knows what your employees, contractors and AI agents cost, and Hybrid workforce shows them on one bill. A business metric adds what that work produced. Together, each metric shows:

On screenWhat it tells you
AI share, or Without a person on the metric's pageHow much of the output needed no person
Cost per unitWhat one unit cost, with the people and the AI behind it
The cost per unit by AI and by a person, named after the metric's unit, such as Cost per contact, by AIWhat a unit costs when AI handles it alone, and when a person handles it
Volume over timeHow the split between With a person and Without a person has moved
SatisfactionWhether quality held while AI took on more of the work

You enter only what was produced. Flowstate supplies the cost. See How costs are calculated.

Example: level 1 support resolved without a person

A support team adds an AI voice assistant to its phone line. It takes the simplest calls itself, such as "my computer won't turn on", and walks the caller through the fix. When it can't help, it hands the call to an advisor. The head of support wants to know whether the assistant is paying for itself, and how much more volume the team can take without hiring.

What to count

Create a metric:

  • Name: Contacts resolved
  • Unit: contact. Plural: contacts
  • Direction: Higher is better
  • Reported: Monthly
  • Whose cost this is: the support team, including the advisors who take the hand-offs
  • Description: "A contact counts once it's closed. It's resolved without a person when the assistant closed it with no hand-off to an advisor."

Create a second metric for your customer satisfaction score, and choose it as the Satisfaction guardrail.

How to record which contacts AI handled

Each month's reading carries two figures:

  • How many: every contact resolved, by the assistant or by an advisor.
  • Of which, without a person: the contacts the assistant closed on its own.

If your contact-centre platform reports both, a developer can push them in through the API. Otherwise, a team lead records them by hand from the platform's monthly report. See Event sources.

Load the months from before the assistant went live too, with zero under Of which, without a person. Zero is right there: the assistant handled none. Leaving it empty would say nobody measured it.

For the assistant's own cost to count, it must be an AI agent that reports its own spend, on a project the support team works on. See Agents. A tool whose provider only sends a total bill shows in AI spend, but not in the metric's cost.

What the screens show

On Insights → Business metrics → Metrics, the Contacts resolved row shows the Volume, the AI share, a Trend line, the Cost per unit, and what the People and Agents cost. If it has the widest gap of any metric between AI and a person, the sentence at the top names it.

Open the metric:

  • The figures across the top show Without a person, then Cost per contact overall, by AI and by a person.
  • Volume over time splits each month into With a person and Without a person, so you can see the assistant take on more of the calls.
  • Cost per unit against AI share shows how cost per contact moves as that share changes.
  • Satisfaction shows the latest satisfaction score and how far it has moved.

How to read it

  • Cost per contact falls as the AI share rises. The assistant is taking work you'd otherwise pay people to do.
  • The AI share rises, but cost per contact stays flat or rises. This is common. The assistant took the calls, but the team is the same size and the assistant's bill came on top. The capacity is real, but it isn't a saving until something else changes, such as taking more volume without hiring.
  • Satisfaction falls as the AI share rises. The saving is costing you something. Look at it before you move more calls to the assistant.

To see what a decision would do, use Model a change at the top of the metric's page. Raise Resolved without a person and read Capacity released. Then see how People taken out changes Cash realised, Net annual benefit and Payback. Use it to plan how the team absorbs more volume, for example by not replacing leavers. Nothing is saved.

Example: insurance claims verified by AI

For example, an insurer might count claims verified. An AI agent checks each incoming claim against the policy and the evidence. It verifies straightforward claims on its own and sends anything unusual to a claims handler as an exception.

What to count

Create a metric:

  • Name: Claims verified
  • Unit: claim. Plural: claims
  • Direction: Higher is better
  • Reported: Weekly
  • Whose cost this is: the claims team, including the handlers who work the exceptions
  • Description: "A claim counts once it's verified. It's verified without a person when the agent verified it and no handler reviewed it."

For the guardrail, create a quality metric such as claims reopened after verification, with Lower is better, and choose it as the Satisfaction guardrail.

How to record which claims AI handled

Each week's reading carries two figures:

  • How many: every claim verified.
  • Of which, without a person: the claims the agent verified with no handler review.

A claim the agent passed to a handler counts in How many only.

If your claims system records who verified each claim, a developer can push both figures in through the API every week, and load past weeks the same way. See the Business metrics API.

If the agent is a workflow your own team runs, add it as an AI agent that reports its own spend, and put it on a claims team project. That way its cost counts towards the metric.

What the screens show

The screens are the same as in the support example:

  • Cost per claim, by AI sits beside Cost per claim, by a person.
  • Volume over time shows the exceptions as With a person.
  • Satisfaction shows the reopened-claims guardrail.

How to read it

  • People handle the hard claims. Exceptions take longer, so a claim handled by a person costs more than one verified by AI, even if nobody got slower. Read the gap as the cost of the work AI can't do yet. It doesn't mean every claim could cost what AI's do.
  • Handlers' time checking the agent's work is in the people cost. If handlers review much of what the agent does, verifying a claim without a person saves less than the gap suggests.
  • The AI share rises and reopened claims hold steady. The agent is taking more of the straightforward work without letting quality slip.
  • Reopened claims rise. The agent is verifying claims it shouldn't. Check before sending it more.

Mistakes to avoid

  • Expecting more AI to lower cost by itself. Moving work to AI releases capacity. It only saves money once people or spend actually change.
  • Treating empty as zero. Not measured means nobody recorded how much was handled without a person. Zero means AI handled none.
  • Reading a saving without its quality measure. Pair every outcome with a guardrail, so a rising AI share can be checked against satisfaction or errors.
  • Adding costs across metrics. Two metrics can name the same people, so their costs overlap.
  • Reading it as a verdict on a person. A metric prices the work of a group of people, not any one person.

Compare before and after AI

Not available yet

Flowstate doesn't compare a metric with a baseline from before AI, or work out savings against one. To compare, open the period control, choose Custom, and set From and To to a stretch before AI took on the work. Note the Cost per unit and AI share, then choose a stretch after it. Readings and Volume over time show each period side by side.

See the bigger picture

  • The whole bill. Hybrid workforce shows employees, contractors, and agents and copilots together.
  • Projects. A project's Value tab shows the business metric its team produces beside its drivers. See Project value.
  • AI spend by project. In Breakdown, aggregate by Projects to see a Per unit column. Switch Cost basis to AI + human to see people and AI cost together.
  • Questions. Ask Eddy "What does a resolved contact cost us, with and without a person?" or "Which teams produce the most without a person?".

If something's not right

Cost per unit, by AI shows a dash. No reading in the period records any units handled without a person.

Agents cost is zero or very low. Flowstate isn't recording AI use for the metric's people or agents. Check the agent reports its own spend and is on one of the team's projects, or that people's AI sessions are coming in. A provider bill that only arrives as a total isn't included.

AI share says Not measured. Readings in the period don't include Of which, without a person.

Satisfaction says no metric is linked. Edit the metric and choose a Satisfaction guardrail.

Flowstate Documentation