OpenLot Book audit
Systems

When a Voice Bot Should Stop and Transfer

OpenLot 9 min read

The most important document in a voice deployment is the list of things the system must never attempt. Twelve triggers cover almost all of it, and a vendor who cannot produce an equivalent list has not run enough deployments to have discovered them.

The twelve triggers on which a dealership AI voice bot should stop and transfer to a person, grouped by type

This guide covers the twelve triggers, how to implement them, what a transfer must carry, where transfer logic fails, and what to measure.

The twelve triggers

Grouped by why they matter.

Commitment triggers — the system would be making a promise

  1. Any price, payment or monthly figure
  2. Trade value or payoff amount
  3. Vehicle availability it cannot verify live
  4. A delivery or completion date
  5. Any discount, offer or waiver

These share one property: the customer would reasonably treat the answer as binding. That is the test from where agent autonomy should end, applied to voice, where it bites harder because there is no written record the customer can re-read.

Judgment triggers — the system cannot evaluate the situation

  1. A complaint, in any form
  2. Any mention of legal action, a lawyer or a regulator
  3. A safety concern — a vehicle behaving dangerously
  4. A hardship or personal circumstance
  5. A dispute about something previously said or agreed

Capability triggers — the system cannot continue usefully

  1. Three failed attempts to understand, or poor audio
  2. An explicit request for a person

Trigger 12 should be honoured immediately and without friction. Making it difficult optimises containment at the cost of the customer, which is the wrong trade — the point made in the first ten seconds.

How should they be implemented?

Three layers, in order of reliability

Layer How it works Catches
Capability limit The system has no ability to state a price Triggers 1–5, absolutely
Pattern detection Matched before a response is generated Triggers 6–10
Attempt counter Hard limit on retries and audio failures Triggers 11–12

The first layer is the only one that holds under every condition. If the system physically cannot produce a currency figure, it will not produce one — regardless of how the customer phrases the question.

Instructions are not a layer. "Do not quote prices" is advisory, and advisory controls fail at exactly the moment they matter.

A fourth, weaker layer is useful as a net: output validation before anything is spoken — no currency figures, no dates, no availability claims. Cheap, and it catches what the other three missed.

What must a transfer carry?

Five items. Without them the transfer is cold, and a cold transfer is worse than no transfer because the customer now explains themselves twice.

Carried Why
Who is calling Matched, not asked again
What they asked In their words
What the system said So the person does not contradict it
Why it transferred The trigger that fired
Any commitment made An appointment time offered, a callback promised

The fourth is the one most commonly dropped, and it is the most useful: a person who knows the call escalated on a complaint picks it up very differently from one who knows it escalated on a price question.

Where does transfer logic fail?

1. Implemented as instructions. The dominant failure, and it is invisible until a transcript review finds the system quoting a payment.

2. No target available. The trigger fires correctly and the transfer lands nowhere — which defeats the whole design, as it does everywhere else in the funnel.

3. Written once, never revised. Customers phrase things in ways nobody anticipated. The list should change in the first month, from real transcripts.

4. Too eager. A system transferring on anything slightly unusual produces no value and a lot of transfers. The triggers should be specific.

5. Silent transfer. The caller hears music with no explanation. Say what is happening and why.

6. No transfer outside staffed hours. The hardest case: a complaint at 11pm, and the honest limit on what voice coverage can do. The honest answer is to take complete details, state clearly when someone will call, and ensure that actually happens.

What should you measure?

Metric How to compute What it catches
Escalation accuracy Triggers that fired ÷ should have fired Sampled from transcripts
False transfer rate Transfers that did not need a person Over-eagerness
Transfer with full context Transfers carrying all five items ÷ total The quality of the handoff
Transfer pickup time Transfer → human answers Whether the design works at all
Commitment-test violations Sampled calls where the system said something binding Should be zero
After-hours escalation follow-through Promised callbacks actually made The hardest case, measured

Row five should be structurally zero. If it is not, the triggers are implemented as instructions rather than as capability limits, and the fix is architectural rather than a matter of tuning.

Row one requires listening to calls, weekly at first. There is no substitute for it, and the first month of transcripts is where the trigger list is actually written.

Frequently asked questions

When should an AI voice bot transfer to a person?

On twelve triggers across three groups: anything that would constitute a commitment such as price, trade value, availability or a date; anything requiring judgment such as a complaint, legal mention, safety concern or hardship; and capability failures including repeated misunderstanding or an explicit request for a person.

Is a transfer a failure of the system?

No. Failing to transfer is. A high transfer rate with fast pickup and full context is the signature of a correctly scoped deployment, while a low transfer rate usually means the system is answering things it should be escalating.

How should transfer triggers be enforced?

As capability limits wherever possible — a system that cannot produce a currency figure will not produce one — with pattern detection evaluated before a response is generated, and a hard attempt counter for misunderstandings. Instructions alone are advisory and fail at exactly the moment they matter.

What has to be included in a transfer?

Who is calling, what they asked in their own words, what the system already told them, which trigger caused the escalation, and any commitment made such as an offered appointment time. The trigger reason is the item most often dropped and the most useful to the person picking up.

Should a request for a person be honoured immediately?

Yes, without friction. Making it difficult optimises containment rate, which is a vendor metric rather than a business one, and a caller who fights the system before reaching a person remembers that rather than the eventual resolution.

What happens when a complaint arrives after hours?

The hardest case. The honest handling is to take complete details, state clearly when someone will call back, and ensure that call actually happens. Attempting to resolve a complaint with no person available produces a worse outcome than acknowledging the limit.

How often should the trigger list be revised?

Weekly in the first month, from real transcripts. Customers phrase things in ways nobody anticipates, and a trigger list written in a meeting and never revised is a list of what someone imagined rather than what happens.

What is the single number that shows the boundary holds?

Commitment-test violations — sampled calls where the system said something a customer could reasonably treat as binding. It should be structurally zero, and if it is not, the triggers were implemented as instructions rather than as capability limits.

Conclusion

  • Twelve triggers, three groups: commitment, judgment, capability.
  • Capability limits beat instructions. If it cannot say a price, it will not.
  • A transfer carries five items, and the trigger reason is the one most often dropped.
  • Honour a request for a person immediately. Containment is the wrong thing to optimise.
  • Rewrite the list from real transcripts in the first month. That is where it is actually written.

Last updated: