The argument every agency eventually has
Somewhere around month four of a successful campaign, a client says the leads are not converting. It is the most common conversation in agency life and it is almost always unwinnable, because both parties are arguing from things the other cannot see. The client sees their sales team's frustration. The agency sees rising lead volume at a falling cost. Neither has any evidence about what actually happened on the phone.
What usually follows is a negotiation dressed as an analysis. The agency offers to tighten targeting. The client agrees the leads might be fine but wants "better quality". Volume drops, cost per lead rises, and nothing about the underlying situation changes, because the underlying situation was never diagnosed.
When the calls are recorded and transcribed, the argument becomes an audit. Pull the last fifty calls, read the summaries, and the answer is generally obvious within twenty minutes. Sometimes the leads genuinely are poor and the agency has work to do. Far more often the picture is some combination of the same handful of findings: a meaningful share of calls never answered, a receptionist who treats price questions as an interruption, no callback on voicemails, or an out-of-hours volume nobody knew existed.
Those are all fixable, and none of them is fixable by the agency alone. Which is precisely why the evidence matters — it moves the conversation from "your leads are bad" to "here are eleven calls last month where nobody picked up, worth roughly this much in lost work, and here is what to do about it." That is a consultative relationship rather than a defensive one, and it is worth considerably more than the campaign itself.
The uncomfortable corollary is that sometimes the evidence goes the other way, and the transcripts show the agency is sending genuinely unqualified traffic. That is also useful. An agency that finds its own targeting problem in week six and fixes it keeps the account; one that finds out at renewal usually does not.
What happens to a call after it ends
The pipeline runs in the background against every tracked call on a client where it is enabled. Nobody has to request an analysis.
The recording is stored against the lead
Not in a separate telephony portal with its own login. The recording sits on the same lead record as the caller, the source, the campaign and the landing page, so listening to it never requires reconstructing which call this was.
It is transcribed
Full text, timestamped, searchable. A transcript is more useful than audio for most purposes precisely because you can search fifty of them in the time it takes to listen to one.
A summary is generated
A few sentences: what the caller wanted, what was agreed, what was left open. This is the artefact that actually gets read, because nobody reads a four-minute transcript when they are scanning a week of calls.
The call is classified and scored
Genuine new enquiry, existing customer, supplier, wrong number, silent, spam. This is the classification that fixes lead counts, because a client's existing customers ringing the tracked number are not new leads and counting them as such inflates everything downstream.
Anything notable is flagged
Missed, abandoned, very short, out of hours, or containing words you asked to be watched for. Flags are what turn a pile of calls into a list of things to do on Monday morning.
What you find in the first month, almost every time
These recur across accounts and industries with unusual consistency. If you have never looked at a client's call handling, expect at least three of them.
The unanswered fifth
Somewhere between ten and thirty per cent of calls not answered at all, concentrated at lunchtime and after five. The client believes it is near zero, because nobody counts the ones they missed.
The out-of-hours volume
A steady stream of calls in the evening and at weekends going straight to an answerphone nobody checks until Monday, by which time the caller has hired somebody else.
The price-question hang-up
Callers asking what something costs, being told the business does not quote over the phone, and ending the call. A scripting problem that looks exactly like a lead quality problem in the report.
The repeat customers in the lead count
Existing customers ringing the tracked number for support, counted as new leads, inflating volume and deflating cost per lead across the whole account.
The voicemail nobody returns
Messages left, no callback recorded, no follow-up. Usually the single highest-value fix available and it costs the client nothing but a habit.
The best-performing keyword nobody trusted
A term the client wanted paused, whose calls turn out to be the longest and most qualified on the account. Transcripts settle that in an afternoon.
Scoring calls without pretending the machine is certain
Automatic call scoring is genuinely useful and it is also the part of this category most oversold. A model reading a transcript can tell a new enquiry from a wrong number with high reliability, because those look nothing alike. It can tell a strong lead from a weak one with moderate reliability. It cannot tell you whether the job was won, and anything claiming to predict that from four minutes of audio is selling confidence rather than information.
So the scoring here is built to be checked rather than believed. Every score carries the transcript that produced it, one click away. Scores can be corrected by hand and the correction sticks. And the categories are deliberately coarse — the useful distinctions in a local services account are "new enquiry", "existing customer", "not a lead" and "needs a human to look", not a nine-point quality scale with two decimal places.
That coarseness is what makes it usable at volume. An agency running thirty accounts is not going to review every call; what it needs is a reliable way to reduce two thousand calls to the forty that somebody should actually look at. Getting the "not a lead" bucket right does more for that than any amount of sophistication in the middle of the scale.
Keyword spotting fills the gap where the model has no context and you do. Tell it to flag calls mentioning a competitor, a specific service line, "emergency", "insurance", or the name of the client's biggest product, and those calls surface as a list regardless of how they scored. This is the crude tool and it is frequently the most valuable one, because you know things about the client's business that no general model does.
The costs are worth stating plainly: transcription and analysis consume paid capacity, metered per client, with a ceiling you set. A client whose call volume triples should not be able to triple your bill without warning, so the ceiling is enforced rather than advisory and you are told before it is reached rather than after.
Consent, storage and the law
Call recording is regulated, the regulation differs by jurisdiction, and getting it wrong is the kind of mistake that ends an agency relationship. These are the controls; the legal advice has to come from somebody qualified to give it.
- Recording is per client, off unless enabled. It is never switched on across an account base by default. A client who has not agreed to record calls does not have their calls recorded.
- An announcement can be played first. Configurable per client, in the client's own wording, before the call connects — which is the standard mechanism for one-party and two-party consent regimes alike.
- Retention is set, not indefinite. Recordings and transcripts expire on the schedule you configure. Holding audio of members of the public forever because nobody chose a number is a liability, not a feature.
- Deletion is real. A recording can be deleted on request, individually or for an entire client, and it is removed rather than hidden behind a flag.
- Access follows tenancy. A client login reaches that client's recordings and nothing else, enforced server-side on every request rather than by hiding a button.
- Transcription can run without recording retention. Where policy allows analysis but not storage, the transcript and summary can be kept while the audio is discarded on a much shorter clock.
Which calls deserve a human, and which do not
The point of classification is triage. This is roughly how the buckets map to action in a working agency.
| Bucket | Typical share | Who should see it | Why |
|---|---|---|---|
| New enquiry, handled well | 30–50% | Nobody, routinely | This is the system working. Sample a few for coaching and move on. |
| New enquiry, handled badly | 5–15% | The client, with the recording | The highest-value list in the product. Each one is recoverable revenue and a coaching point. |
| Missed or abandoned | 10–30% | The client, urgently | A caller who was not answered is not a lead quality problem and not the agency's fault, but it is the agency who can see it. |
| Existing customer | 10–25% | Nobody, but exclude from lead counts | Counting these as new leads inflates volume and understates cost per lead across the whole account. |
| Not a lead — spam, wrong number, silent | 5–20% | Nobody | Filtered out so the reported number is enquiries. Kept and visible so the filter can be checked. |
What a missed call is actually worth
It is worth doing this arithmetic once with a client, out loud, because the number surprises people and surprise is what makes them act. Take the client's average job value — a boiler replacement, a set of veneers, an initial consultation — and their own estimate of how often a genuine enquiry becomes work. Multiply by the number of unanswered calls in the last month. That figure is not precise and it does not need to be; it is an order of magnitude, and the order of magnitude is usually four or five figures.
Then compare it to what they are paying you. On a great many local accounts, the revenue lost to unanswered phones in a month exceeds the entire monthly retainer, sometimes by several multiples. That is an uncomfortable slide and it is the single most persuasive one an agency can put in front of a client, because it reframes the relationship: you are not the cost line being scrutinised, you are the only party in the room who can see where the money is going.
The follow-up matters more than the finding. A missed-call list is only useful if somebody rings those people back, so the practical version of this is an alert the moment a call goes unanswered — by SMS, by email, into Slack — with the number and the source attached. A callback within five minutes recovers a large share of missed enquiries. A callback on Monday recovers almost none, because the caller rang three businesses and the second one answered.
There is a second-order effect worth watching for. Once a client starts answering the phone reliably, their conversion rate rises without a single change to the campaign, and cost per acquisition falls even though cost per lead has not moved. Agencies who have never measured call handling routinely attribute that improvement to their own optimisation work. It is more honest, and considerably more valuable, to attribute it correctly — because a client who knows that answering the phone made them money will keep doing it.
Running this across thirty accounts without drowning
The failure mode of call intelligence at agency scale is not technical. It is that somebody enables it everywhere, produces a vast quantity of transcripts, reviews them enthusiastically for two weeks, and then never opens the section again. The data keeps accumulating and stops being read, which is worse than not having it, because the agency now believes it is monitoring something it is not.
The workflow that survives is exception-based. Nobody reads good calls. The account manager's weekly job is to open one list — the calls flagged as mishandled, missed or keyword-matched — which on a typical account is between five and fifteen items rather than two hundred. Fifteen minutes per account per week is sustainable; two hours is not, and anything that requires two hours will be abandoned by March.
The second discipline is to enable it selectively. Transcription on every account from day one is expensive and mostly wasted. It earns its cost on accounts where the phone is the primary channel, where the client has questioned lead quality, where a new campaign is being validated, or where you suspect a handling problem. Turning it on for a month to answer a specific question and then turning it off again is a perfectly respectable pattern, and per-client metering exists precisely so that is a decision you can make rather than an all-or-nothing switch.
The third is to route the findings to whoever can act. An agency reading a missed-call list changes nothing; the client's office manager reading it changes everything. Where a client is willing, give them the exception list directly and let the weekly digest arrive in their inbox. Your job then shifts from reporting the problem every month to occasionally checking that it is still being handled — which is a much better use of an account manager and a much stickier relationship.
What this is not
It is not a substitute for the client's CRM. The transcript tells you what was said; whether the job was quoted, won or lost lives in the client's own systems, and the most this can do is push the lead there with its summary attached so somebody has a reason to update it.
It is not accurate on every call. Heavy accents, bad lines, two people talking over each other and trade jargon all degrade transcription, sometimes badly. The audio is always one click away for exactly this reason, and a summary that reads strangely is a prompt to listen rather than a fact to report.
It is not a performance management tool, and agencies should be careful about presenting it as one. Handing a client a league table of which receptionist scores worst is a fast way to have recording switched off entirely, and usually misdiagnoses a scripting or staffing problem as an individual one. The framing that survives is "here is what callers are asking for and where we are losing them", not "here is who to blame".
And it is not free. Transcription and analysis carry a real per-minute cost, which is why it is metered, capped and visible per client rather than bundled invisibly into a flat fee that quietly stops being profitable on your busiest account.
Where it sits
Call intelligence sits directly on top of call tracking: no tracked calls, nothing to analyse. It feeds two things downstream. The lead pipeline, because a classified call arrives already sorted into enquiry or not, which is the hardest part of keeping a pipeline honest. And cost per lead, because excluding existing customers and spam from the denominator is the difference between a real figure and a flattering one.
It also sharpens call routing. Once you can see that a third of missed calls happen between twelve and two, a routing rule that sends lunchtime calls to a second number stops being a guess and becomes a response to measured demand.
For the client, the most visible output is usually the smallest: a weekly list of calls that need attention, delivered by email, with a summary and a link to the recording. Agencies that send that list consistently report it as the single thing clients mention unprompted at renewal.
Common questions
Is every call recorded, or only some?
Only where recording is enabled, and it is off until you turn it on for a client. Once enabled, every tracked call on that client is recorded and processed rather than sampled — sampling defeats the purpose, since the calls you most need to see are the ones nobody chose to review.
Do callers have to be told the call is recorded?
That depends on jurisdiction, and it is a question for the client's own legal advice rather than for a software vendor. What the system provides is the mechanism: a configurable announcement, in the client's wording, played before the call connects, set per client rather than globally.
How accurate is the transcription?
Good on clear audio and noticeably worse on heavy accents, poor lines, crosstalk and trade-specific jargon. The recording stays one click from the transcript for that reason. Treat a summary that reads oddly as a prompt to listen, not as a fact to put in a client report.
What does the AI scoring actually decide?
Deliberately coarse categories: new enquiry, existing customer, not a lead, and whether a human should look. It does not predict whether a job will be won, and a coarse classification you can trust is far more useful at volume than a nine-point quality scale nobody believes.
Can I override a score the model got wrong?
Yes, and the correction sticks. Every score carries its transcript so you can see why it landed where it did. A classification you cannot inspect or override is a classification that quietly edits your lead counts, which is the opposite of what this is for.
Are missed calls reported?
Yes, as their own list — unanswered, abandoned before pickup, and out of hours, separated rather than merged. This is consistently the most valuable output for a client, because a missed call is recoverable revenue and neither the client nor the agency usually knows how many there are.
Can I get alerted when specific words are used?
Yes. Set the terms per client — a competitor name, a service line, "emergency", "insurance", a product — and matching calls surface as a list regardless of how they scored. This crude tool is often the most valuable one, because you know things about the client's business that no general model does.
How long are recordings kept?
For the retention period you set per client, after which they expire automatically. Individual recordings and whole clients can be deleted on request, and deletion removes the data rather than hiding it. Indefinite retention of the public's phone calls is a liability rather than a feature.
What does transcription cost?
It carries a real per-minute cost from the underlying provider, which is why usage is metered per client with a ceiling you set and warnings before it is reached. An account whose call volume triples should not be able to triple your bill without telling you first.
Can the client see the recordings and transcripts?
Yes, if you give them access — a client login reaches their own calls and nothing else, enforced on the server for every request. Many agencies instead send a weekly digest of the calls that need attention, which gets read far more reliably than a dashboard somebody has to remember to open.