The value of conversational AI is that the thread continues, and the thread breaks in five specific places. Four of the five are invisible in a demo — because a demo is one person, one channel, one session, which is the only configuration in which nothing breaks.
This guide covers the five breaks, the test that finds them, what continuity requires, where it fails, and what to measure.
The five breaks
| # | Break | Visible in a demo? |
|---|---|---|
| 1 | Session to session — they come back tomorrow | No |
| 2 | Channel to channel — web chat, then text | No |
| 3 | Device to device — phone, then desktop | No |
| 4 | Bot to human — the escalation | Sometimes |
| 5 | Shift to shift — a different person picks it up | No |
Only break 4 is routinely demonstrated, and it is the one most products handle best. The other four are where real customers actually go, and a demo is specifically constructed to avoid all of them.
1. Session to session
A visitor asks about a vehicle on Tuesday and returns on Thursday. If the system greets them as new, everything learned on Tuesday is gone and the customer repeats themselves.
This is partly an identification problem — cookies expire, devices change — and partly a design choice about how hard to try to match a returning visitor.
2. Channel to channel
The expensive one. A customer chats on the site, then texts the number from the confirmation. If those are two threads, the customer has to restate everything, and from their side the store simply does not remember.
Continuity across channels requires a messaging layer rather than a chat widget, which is a different product — the distinction in one thread, many apps.
3. Device to device
Browsing on a phone at lunch, returning on a desktop in the evening. Without identification, two sessions.
4. Bot to human
The escalation, and the one most products handle. The requirement is the same as everywhere: the person receives who, what they asked, what was already said, and why it escalated — the payload in the showroom handoff.
5. Shift to shift
The thread escalates at 8pm, a person responds, and a different person picks it up at 9am. Without a summary, the customer explains a third time.
The test that finds them
Five steps, ten minutes, on a live product
Step Do this Break tested 1 Start a chat on a vehicle page, ask two questions — 2 Close the tab. Return an hour later 1 — session 3 Text the number from the chat confirmation 2 — channel 4 Ask about price, see the escalation 4 — bot to human 5 Ask the person what you already told the system 4, verified Break 3 needs a second device and break 5 needs a shift change, so neither fits in ten minutes — but steps 2 and 3 are the ones that matter most and both are quick.
Step 3 is the one that separates products. A chat widget and a texting product from the same vendor are frequently two separate threads, and the only way to find out is to try it.
What does continuity actually require?
Four things, and they are infrastructure rather than conversation quality.
1. Identity resolution. Matching a returning visitor, a phone number and a web session to one person. Without it, nothing else is possible — the matching rules problem, applied in real time.
2. A single conversation store. One thread, written to by every channel, rather than per-channel logs.
3. Context passed on escalation. The payload, every time.
4. A summary at handoff. Not the full transcript — three lines a person can read in ten seconds.
Products that have the first two have continuity. Products that have strong conversation quality and neither have a good demo.
Where does it fail?
1. Separate products stitched together. Chat from one vendor, texting from another, and no shared identity.
2. Identity by cookie only. Expires, and does not survive a device change.
3. Per-channel conversation logs. The data exists and cannot be assembled.
4. Escalation without payload. The most expensive, because it happens at the moment of highest intent.
5. No shift summary. The overnight work is discarded at 9am.
6. Evaluating on conversation quality alone. Which is what a demo measures and is not what decides the outcome — and it does not distinguish rules-based from conversational either.
What should you measure?
| Metric | How to compute | Which break |
|---|---|---|
| Repeat rate | Conversations where the customer restates known information | All five, combined |
| Cross-channel continuity rate | Threads continuing across channels ÷ cross-channel customers | Break 2 |
| Returning visitor recognition | Recognised ÷ returning | Break 1 |
| Escalations with full payload | Complete ÷ escalations | Break 4 |
| Shift-change summaries present | Summarised ÷ overnight threads | Break 5 |
| Identity match rate | Matched to one person ÷ interactions | The foundation |
Row one is the summary measure and the only one a customer would recognise. Everything else on the list is a cause of it.
Row six is the foundation. A low identity match rate means the other five numbers cannot improve regardless of what is done to the conversation layer.
Frequently asked questions
Where does a conversational AI thread break?
Five places: between sessions when a customer returns another day, between channels when they move from web chat to text, between devices, at the escalation to a person, and at a shift change when someone else picks the conversation up. Only the escalation is routinely demonstrated.
Why are most breaks invisible in a demo?
Because a demo is one person, on one channel, in one session, with one representative — which is the only configuration in which nothing breaks. Real customers switch channels, leave and return, and are handled by different people.
Which break costs the most?
Channel to channel. A customer who chats on the site and then texts the number from their confirmation expects the store to remember, and when it does not, the experience is that the business simply does not keep track of its own conversations.
How do you test continuity before buying?
Five steps in about ten minutes: start a chat and ask two questions, close the tab and return an hour later, text the number from the confirmation, ask something that triggers an escalation, and then ask the person what you already told the system. Step three separates products most clearly.
What does continuity require technically?
Identity resolution matching a returning visitor, a phone number and a web session to one person; a single conversation store written to by every channel; context passed on escalation; and a short summary at shift handoff. The first two are the foundation and the rest depend on them.
Can separate chat and texting products provide continuity?
Rarely, even from the same vendor. They are frequently separate threads with separate identity, and the only reliable way to establish it is to start a conversation in one and continue it in the other during evaluation.
Why is identity match rate the foundation?
Because without matching a returning visitor and a phone number to one person, no amount of work on the conversation layer produces continuity. A low match rate caps every other continuity metric regardless of product quality.
What should be measured after deployment?
Repeat rate — the share of conversations in which a customer restates something they already told you. It is the only metric on the list a customer would recognise, and every other continuity measure is a cause of it.
Conclusion
- Five breaks, and four are invisible in a demo by construction.
- Channel to channel is the expensive one, and the test takes two minutes.
- Continuity is infrastructure, not conversation quality: identity and one conversation store.
- Separate chat and texting products are usually two threads, even from one vendor.
- Repeat rate is the customer-visible measure. Everything else is a cause of it.
Last updated: