OpenLot Book audit
Systems

Auditing an AI That Talks: A Call QA Checklist

OpenLot 9 min read

Every voice deployment needs somebody listening to calls, and almost none of them has one. Fifteen checks against fifty sampled calls, once a month, is the difference between knowing what your system does and assuming it — and the first sample is always surprising.

A fifteen-point quality audit checklist for dealership AI voice calls, grouped into five areas

This guide covers the fifteen checks, how to sample honestly, what to do with findings, the retention question, and what to measure.

The fifteen checks

Grouped into five areas. Score each call pass or fail.

Opening (3)

  1. Store named within the first few seconds
  2. Automation disclosed, briefly and without apology
  3. Caller reached the point quickly — no long preamble

Behaviour (4)

  1. No currency figure spoken — the hard line
  2. No availability claim it could not verify
  3. No date or completion promise
  4. Interruptions handled — stopped, did not restart from the top

Escalation (3)

  1. Triggers fired when they should have — against the twelve-trigger list
  2. Transfer carried context — who, what, what was said, why
  3. Transfer was picked up promptly

Compliance (3)

  1. Recording disclosure given, where required
  2. Opt-out or do-not-call request honoured immediately
  3. Outbound call within permitted hours, recipient local time — see outbound requirements

Outcome (2)

  1. Caller got what they needed, as opposed to the call merely ending
  2. No repeat call from the same number within 24 hours

Check 4 should never fail. If it does, the controls are instructions rather than capability limits, and the fix is architectural — covered in transfer triggers.

How do you sample honestly?

Not randomly, and not by picking interesting ones.

A fifty-call sample

Stratum Calls Why
Resolved without transfer 15 Where over-reach hides
Transferred 15 Where context failures hide
Abandoned 10 The most informative and least listened to
Longest duration 5 Where the system struggled
Repeat callers within 24h 5 Where it failed and nobody noticed

Pure random sampling over-weights the easy calls, because most calls are easy. Stratifying forces the sample to include the ones that went wrong.

Row three is the one nobody listens to and the one that explains the most. An abandoned call is a customer who left, and the recording usually makes the reason obvious within fifteen seconds.

Fifty calls takes an afternoon. Monthly is enough once a deployment is stable; weekly for the first month, because that is when the trigger list gets rewritten.

What do you do with the findings?

Three buckets, three different responses.

Configuration — a message is wrong, an hour is wrong, a routing rule is wrong. Fix it that week. If this requires a support ticket, that is itself a finding about the product.

Scope — the system is attempting something it should not. Add the trigger, as a capability limit rather than an instruction, and re-sample to confirm.

Capacity — transfers are correct and nobody is picking them up. Not a software problem, and no amount of tuning fixes it.

Record which bucket each finding falls into over time. The distribution tells you something useful: mostly configuration means the product is sound and the setup needs work; mostly scope means the product is over-reaching; mostly capacity means the deployment is fine and the store is not.

The retention question

Voice AI means every call is now recorded, where previously some were. That is a data decision nobody usually makes deliberately, and it sits alongside the Safeguards Rule obligations the store already has.

What to settle, in writing:

  • How long recordings are kept, and why that period
  • Who can access them, and whether that access is logged
  • Whether transcripts are kept separately and for how long
  • What happens to recordings on vendor termination
  • Whether recordings may be used to train anything

The last point matters and is frequently silent in contracts. Customer call audio used to improve a vendor's general model is a data-sharing question, and it belongs in the service provider clauses alongside the rest.

Where does call QA go wrong?

1. Nobody does it. The overwhelming default, and it is how a deployment drifts away from the three mechanics it was tested on.

2. Random sampling. Over-weights the easy calls, which is most of them.

3. Only listening to transfers. Misses the over-reach, which hides in the contained calls.

4. Scoring without a checklist. Produces impressions rather than findings.

5. Findings with no owner. Same failure as every other alert in the store.

6. Stopping after the first month. Behaviour drifts as configurations change.

What should you measure?

Metric How to compute What it catches
Pass rate, per check Across the sample Which of the fifteen is failing
Check 4 failures Currency figures spoken Should be structurally zero
Escalation accuracy Check 8 pass rate Whether the trigger list is real
Context-complete transfers Check 9 pass rate The handoff quality
Abandonment reasons Categorised from the abandoned stratum The most actionable output
Finding distribution Configuration / scope / capacity What kind of problem you have

Row five is the highest-value output of the whole exercise. Ten abandoned calls, listened to properly, usually produce two or three specific fixes that no dashboard would ever have surfaced.

Frequently asked questions

How do you audit an AI voice system at a dealership?

By sampling fifty calls monthly against a fifteen-point checklist covering the opening, behaviour, escalation, compliance and outcome. Scoring each call pass or fail produces findings rather than impressions, and the exercise takes about an afternoon.

How should calls be sampled for review?

Stratified rather than randomly: a mix of calls resolved without transfer, calls transferred, abandoned calls, the longest calls, and repeat callers within a day. Random sampling over-weights easy calls because most calls are easy, which hides exactly the behaviour worth finding.

Which calls are most informative to listen to?

Abandoned ones, and almost nobody does. An abandoned call is a customer who left, and the recording usually makes the reason obvious within the first fifteen seconds — producing specific fixes that no dashboard would surface.

What should never appear in a sampled call?

A currency figure spoken by the system. If that check ever fails, the controls were implemented as instructions rather than as capability limits, and the fix is architectural rather than a matter of adjusting a prompt.

How often should call QA be run?

Weekly for the first month, because that is when the escalation trigger list gets rewritten from real customer phrasing, then monthly once the deployment is stable. Stopping entirely is the common pattern and behaviour drifts as configurations change.

What do you do with the findings?

Sort them into configuration, scope and capacity. Configuration issues get fixed that week. Scope issues become new capability limits. Capacity issues — correct transfers nobody picks up — are not software problems and no amount of tuning addresses them.

Does voice AI create a recording retention problem?

It creates a decision nobody usually makes deliberately, because every call is now recorded where previously only some were. Retention period, access logging, transcript handling, what happens on vendor termination, and whether recordings may be used for training all need settling in writing.

Can a vendor use our call recordings to train their models?

Only if your contract permits it, and this clause is frequently silent. Customer call audio used to improve a vendor's general model is a data-sharing arrangement and belongs in the service provider terms alongside subprocessors, retention and deletion.

Conclusion

  • Fifteen checks, fifty calls, one afternoon a month. The first sample is always surprising.
  • Stratify the sample. Random sampling over-weights the easy calls.
  • Listen to the abandoned ones. They are the most informative and the least reviewed.
  • Sort findings into configuration, scope and capacity. The distribution tells you what you have.
  • Settle recording retention in writing, including whether the vendor may train on it.

Last updated: