Methodology

How we count

Two kinds of number appear in this product. Money, which is counted from a field your call platform recorded. And labels — what the caller asked for, how the call ended, whether it was spam — which are a language model's reading of a transcript. They deserve different levels of trust, so both are documented here, along with the method planned for the demand and team labels that are still in build.

Read this before you trust a single figure in the product. If a rule below is wrong for how your business works, the number will be wrong for you, and you should know that on day one rather than at the end of a quarter.

Revenue rules

The revenue rules, in one table

Every money figure on every screen follows from these five decisions. The label rules are further down, and they are looser by nature.

Rule What we do What we do not do
Revenue basis Read the platform's revenue field on each call row. Substitute a different amount field, read a boolean flag as if it were money, or derive a figure ourselves.
Duplicates Exclude duplicate calls from revenue totals; keep them visible as their own count. Silently delete them, or let them inflate revenue and call counts.
Paid call A call is paid when revenue > 0. Infer paid status from call duration, an agent's own label, or a flag that can disagree with the money.
Timezone Compute every aggregate in the workspace's configured timezone; month boundaries follow it. Mix UTC storage with local display, or move your month because a server is in a different zone.
Disagreement Treat your call platform as the system of record and show the variance. Quietly adjust our figure toward yours so the two appear to agree.
The rules in full

Which field we treat as revenue

We read the revenue field on the call row, exactly as your call platform recorded it. We do not read any of the secondary amount fields a platform may also carry, and we do not read a boolean "this call earned something" flag as if it were an amount. Those fields drift apart in ordinary operation: a call can carry a value set by a platform default while the recognised revenue is different, and a flag can be true on a row whose revenue is zero. Picking one field and publishing which one is what keeps an aggregate arguable.

We do not compute revenue ourselves. We never multiply a duration by a rate, apply a default rate to a call, or reconstruct a figure a platform did not report. If the field is empty, the call contributes zero and stays visible as a call — it does not vanish from the denominator.

Duplicates are excluded from revenue

Duplicate calls are dropped from revenue totals. They remain in the data, counted and filterable, because the rate at which one source produces repeat callers is itself worth watching. What they never do is add revenue twice.

The window that defines a duplicate is set on your side, not by us — it is commonly 24 hours to 30 days, and platforms differ on where the clock starts. We apply the platform's own duplicate marking rather than re-deriving it, so a call we exclude is a call your platform already called a duplicate. Where that marking looks wrong, the individual rows are there to check rather than being resolved silently inside a total.

What counts as a paid call

A paid call is a call where revenue is greater than zero. Nothing else. Not a call that ran long, not one an agent marked as won, not a flag set somewhere upstream. Those are all useful signals and we surface them, but they sit upstream of the money and they can disagree with it. Revenue per call and paid rate are computed off this single definition, so the ratios stay consistent with the totals above them.

Every aggregate drills to raw rows

A number you cannot audit is a number you cannot argue with. Every figure in the product — revenue by source, calls by hour, month-to-date, paid rate by location — opens the underlying call rows: the individual calls, their timestamps, their revenue values, their duplicate status. There is no aggregate whose inputs you cannot see.

The same rule applies to labels, not just to money. A call's outcome, intent or requested service opens the sentence in the transcript that produced it. That is the whole answer to "your AI is wrong": you can go and look. The demand and team counts still in build are being built to the same rule — a count will open the calls behind it, and each of those will open the sentence behind the label.

Timezones and month boundaries

Aggregates are computed in the timezone configured on your workspace. Month boundaries follow that timezone, so a month-to-date total starts at midnight local, not midnight UTC. Daily and hour-of-day breakdowns use the same clock, which is what makes a quiet Tuesday afternoon comparable to your opening hours.

Underlying timestamps are stored with their offset and converted at read time — we do not overwrite the original. Changing the workspace timezone recomputes the aggregates rather than rewriting the calls. If your platform's reporting is set to a different timezone, a day-boundary difference is the most common reason a daily figure differs while the monthly figure matches. Check that first.

When our number disagrees with yours

Your call platform is the system of record. It bills, it settles, and it is what everyone else argues from. When our total and its total differ, we show the variance rather than hiding it, and we do not nudge our figure toward theirs to make the two agree. A tool that quietly reconciles itself into agreement is a tool that cannot tell you when something is actually wrong.

Variance is worth reading, not just resolving. The candidates to check, in the order that is quickest to rule out: a timezone difference at a day boundary, a sync that has not caught the last few minutes of calls, a credit applied on their side after our read, or a duplicate marked late. Each of those is visible in the call rows, which is why the drill-through matters more than the headline.

How the call labels are produced

A recording is transcribed, and a language model reads the transcript. From that reading come the labels: what the caller asked for, how ready they were to buy, roughly what the job looked like it was worth, how the call ended, and whether it was spam or a voicemail. There is no rules engine and no keyword list underneath — it is a model reading words, which is a different kind of number from a revenue field, and we would rather say so than blur the two.

Transcription itself sets the ceiling. A bad line, heavy background noise, two people talking over each other, or an accent the transcriber handles poorly all degrade everything downstream. Where the transcript is weak the labels will be weak, and the transcript is right there for you to see why.

How demand and team labels are produced

in build Demand aggregation, unanswered-question detection and per-person team feedback are being built now — none of them is something you can open today, on any tier. The method is settled, so here it is, written as the design rather than as a description of something running.

Demand figures will be counts of call-level labels, nothing more exotic. The model records what each caller asked for in their own terms, similar requests will be grouped, and the group counted over a date range. "Twelve people asked about gutter cleaning last week" will mean twelve calls carry that label — and you will be able to open all twelve. A service will appear on the unmet-demand list when callers asked for it and your workspace has no matching service configured, which means that list will only ever be as good as the services you have told us you offer.

Team labels are designed to work the same way. A question will be recorded when a caller asks one, and marked unanswered when the transcript shows no answer given, an explicitly uncertain answer, or a promise to find out and call back. Those are judgement calls made by a model reading a conversation, and reasonable humans will sometimes disagree with them. Per-person feedback will be an aggregate over that person's calls, meant as input to a conversation and never as an automatic score to act on.

Every label points at the transcript that produced it

Each label carries the span of transcript it came from. Open a call's requested service, its outcome, its spam flag or its value estimate, and you land on the sentence that produced it. This is a hard rule in the data model, not a feature we hope to add: a label with nothing to point at does not get shown. The demand and team labels still in build inherit that rule — they cannot ship without a span to point at either.

Labels can be wrong, and a human can overrule them

They will be wrong sometimes. Sarcasm, a caller who changes their mind halfway, an internal call that looks like a customer, a service name only your industry uses — all of these produce bad labels. So any label can be corrected by a person in your workspace, the correction is what the aggregates then use, and the change is recorded with who made it and when. Your correction wins over the model, permanently.

The right way to use this product is to spot-check. Take twenty calls you already know the answer on, read what we said about them, and correct what is wrong. If the labels do not survive that test on your calls, they will not survive a decision made on top of them either.

What we do not claim

We publish no accuracy percentage for our labels. We have not evaluated them against a human-labelled holdout set yet, so we have no figure that would mean anything — and inventing one on this page would make every other number here worthless. When we do run that evaluation, the results and the method will be published together, including whatever they say.

We also hold no security certification, and do not claim one. What we do commit to is on the security page, in plain terms.

How to check us

Reconcile one figure yourself

Do not take the rules above on trust. Reconciling one figure takes about ten minutes and it is the only reason to believe the rest.

Step 01

Pick any figure.

Revenue for a single day, or one source's total for a week. Narrow beats broad — a small window makes the mismatch obvious instead of averaging it away.

Step 02

Open the calls behind it.

Drill through to the raw rows. You should see each call, its timestamp, its revenue value, and whether it was marked duplicate.

Step 03

Export the same window from your call platform.

Same date range, same timezone setting, same filters. Sum the revenue column, excluding duplicates.

Step 04

Compare the two totals and the two call counts.

Both should agree. If the counts match and the revenue does not, the difference is a field or a duplicate rule. If the counts differ, it is almost always a day boundary or an incomplete sync.

Step 05

Then check the labels separately.

Take twenty calls you remember, read the outcome and the requested service we attached to each, and open the transcript span behind any you disagree with. Correct the wrong ones. Twenty calls is enough to tell you whether the labels hold up on your callers, which is the only sample that matters.

Step 06

Send us the mismatch.

If you can reconcile a difference to a rule on this page, that rule is documented and working as described. If you cannot, that is a bug in our counting and we want the export. A label you had to correct is worth sending too. Fixes to counting rules go in the changelog with the date.

The short version

Money is counted. Labels are read. We tell you which is which.

We read from your call platform and never write to it, so nothing here changes how your phones work or how you get paid. If a rule on this page does not match how your business runs, tell us before you sign up — it is a faster conversation than a refund.

Ask about a counting rule What the platform does