Two quite different technologies are sold as automotive chatbots. A rules-based bot cannot handle what its tree did not anticipate; a conversational one cannot be fully predicted — and knowing which you are being sold decides what you should expect, what it costs to change, and how it will fail.
This guide covers how the two differ, how to tell them apart in a demo, where each fails, the hybrid that works, and what to measure.
How do they differ?
| Rules-based | Conversational | |
|---|---|---|
| How it works | A decision tree somebody wrote | Generates a response from context |
| Predictability | Complete | Partial |
| Unexpected phrasing | Fails, or falls through | Handles it |
| Cost to change | Edit the tree. Cheap | Adjust rules and test. Harder to be sure |
| Can it say something wrong? | Only what was written | Yes |
| Auditability | Trivial — read the tree | Requires transcript sampling |
| Typical price | Lower | Higher |
Neither row is a verdict. A rules-based bot that answers hours, directions and "is it available" correctly every time is a good product for that job. A conversational one that handles a customer describing a vehicle in their own words is a good product for a different job.
The mistake is buying one and expecting the other's behaviour.
How do you tell which you are being sold?
Six test questions, in a live demo. The responses identify the technology regardless of what the marketing says.
The six questions
Ask Rules-based does Conversational does "Do you have anything with a third row under 30?" Falls back to a menu Interprets and answers "Is the blue one still there?" Does not know which blue one Asks which, or uses page context A question with a typo and no punctuation Frequently misses Usually handles Two questions in one message Answers one, or neither Usually both "What's my payment?" Either, depending on the tree Either, depending on the limits The same question phrased three ways Identical or failed Three similar answers Row five is the important one: neither technology is inherently safer on commitments. A rules-based bot with a payment answer in the tree will quote. A conversational one with a hard capability limit will not. The control is the capability limit, not the technology.
Row six is the clearest tell. Identical responses to three phrasings is a tree; three similar but different responses is generation.
Where does each fail?
Rules-based: 1. Any phrasing the tree did not anticipate, which is most phrasings. 2. Dead ends — the tree ends and the visitor is stuck. 3. The tree grows unmaintainable as cases are added. 4. It feels mechanical, which costs engagement — though engagement is set more by prompt behaviour than by the technology, and either type still needs the qualification sequence to be right.
Conversational: 1. It can say something that was never written, which is the whole risk. 2. Auditing requires sampling rather than reading — see call and thread QA. 3. Behaviour drifts as configurations change. 4. Higher cost, and more of it ongoing.
The failure modes are opposite in a useful way: rules-based fails visibly and conversational fails plausibly. A visitor who hits a dead end knows the bot failed. A visitor given a confident wrong answer does not, and neither do you unless somebody is reading transcripts.
The hybrid that works
Most mature deployments are both, and the split follows risk.
| Handled by | What |
|---|---|
| Rules | Hours, directions, department routing, appointment booking flow |
| Rules | Anything touching a commitment — a hard block, not a generated refusal |
| Conversational | Understanding what the visitor is asking |
| Conversational | Vehicle questions answerable from the listing |
| Escalation | Everything else |
The second row is the design principle: use the deterministic technology for the deterministic requirement. A generated refusal to quote prices is a refusal that works most of the time; a hard rule is one that works every time.
What should you measure?
| Metric | How to compute | Which type it judges |
|---|---|---|
| Fallback rate | Conversations hitting "I didn't understand" | Rules-based, mainly |
| Dead-end rate | Conversations with no next step | Rules-based |
| Commitment violations | Currency figures or availability claims spoken | Both — should be zero |
| Consistency | Same question, three phrasings, compared | Identifies the technology |
| Engagement completion | Chats producing contact details ÷ started | The business outcome |
| Transcript sample findings | New failure patterns per review | Conversational, mainly |
Row three is the one that matters most and it is independent of technology. Whichever you buy, the number should be structurally zero, and if it is not the control was implemented as an instruction.
Frequently asked questions
What is the difference between a rules-based and a conversational chatbot?
A rules-based bot follows a decision tree somebody wrote, which makes it completely predictable and unable to handle phrasings the tree did not anticipate. A conversational one generates responses from context, which handles arbitrary phrasing and cannot be fully predicted.
Which is safer on price and commitment questions?
Neither, inherently. A rules-based bot with a payment answer in its tree will quote one; a conversational bot with a hard capability limit will not. The control is the capability limit rather than the underlying technology, and that holds for both.
How do you tell which one a vendor is selling?
Ask the same question three different ways in a live demo. Identical or failed responses indicate a decision tree; three similar but differently worded responses indicate generation. Asking a question with a typo and no punctuation is a second quick tell.
Which fails worse?
They fail differently. Rules-based fails visibly — the visitor hits a dead end and knows it. Conversational fails plausibly — a confident wrong answer that neither the visitor nor the store notices unless somebody is sampling transcripts.
Is a rules-based chatbot still a reasonable purchase?
Yes, for the right job. A bot that answers hours, directions and availability correctly every time is a good product, cheaper to buy and trivially auditable. The mistake is buying it and expecting it to handle a customer describing a vehicle in their own words.
What does the hybrid look like?
Rules for the deterministic parts — hours, routing, the booking flow, and a hard block on anything touching a commitment — with generation used for understanding what the visitor is asking and answering vehicle questions from the listing. Everything else escalates.
Why use rules for commitment blocking rather than generation?
Because a generated refusal works most of the time and a hard rule works every time. The requirement is deterministic, so the deterministic technology is the right tool for it even in an otherwise conversational product.
What should be measured regardless of type?
Commitment violations — currency figures or availability claims the system stated. The number should be structurally zero for both technologies, and if it is not, the control was implemented as an instruction rather than as a capability limit.
Conclusion
- Two technologies, one product name, with opposite failure modes.
- Rules-based fails visibly; conversational fails plausibly. The second is harder to catch.
- Neither is inherently safer on commitments. The capability limit is the control.
- Three phrasings of one question identifies the technology in a demo.
- Use rules for deterministic requirements, even inside a conversational product.
Last updated: