"AI sales assistant" describes two products that share almost nothing. One assists your salesperson and one replaces part of their job with a customer-facing system — and the risks, the evaluation and the person who should be in the room differ completely between them.
This guide covers the two products, how to tell which is in a proposal, how each is evaluated, where each fails, and what to measure.
The two products
| Copilot | Autonomous agent | |
|---|---|---|
| Who it talks to | Your salesperson | Your customer |
| Output | A draft, a summary, a suggestion | A sent message |
| Human checkpoint | By construction | Removed |
| Risk if wrong | Someone notices and edits | The customer received it |
| Evaluated on | Time saved, adoption | Escalation accuracy, show rate |
| Who should evaluate | The salespeople | The GM, and compliance |
The checkpoint row is the whole distinction, and it is the same one that separates an assistant from an agent: one drafts and a person sends, the other sends.
How do you tell which is in a proposal?
One question: does anything reach a customer without a person pressing send?
If no, it is a copilot. Evaluate it on whether your salespeople use it, because an unused copilot is worth exactly nothing regardless of how good its output is.
If yes, it is an agent. Evaluate it on escalation accuracy and on what it is structurally prevented from saying.
Many products are both, with a setting. Ask which mode ships enabled, because the default is what you will be running.
Evaluating a copilot
Adoption is the whole evaluation
A copilot that is not used produces nothing, and copilots go unused for predictable reasons.
Reason Fix The draft needs more editing than writing from scratch Better context, or abandon it It lives in a tool they do not already have open Put it where they work It is slower than typing A real problem at short messages It sounds nothing like them Tone matters more than people expect Nobody showed them twice Training, and a second session Measure adoption weekly for the first month. A copilot at low adoption in week four will be at zero by week twelve, and the honest decision is to cancel rather than to wait.
The useful copilot applications in a dealership are narrow and real: summarising a long thread before a call, drafting a first reply the rep then edits, and surfacing what a customer said three weeks ago. All three save minutes and none touches a customer directly.
Evaluating an agent
Different evaluation entirely, and it is the one covered across the rest of this cluster: escalation accuracy, capability limits rather than instructions, handoff payload, and show rate as the guardrail.
The question that matters most for an agent is not what it can do. It is what it is prevented from doing, and whether that prevention is structural.
Where does each fail?
Copilot: 1. Low adoption, which is the dominant failure. 2. Output requiring more editing than writing. 3. Sitting in a separate tool. 4. Generating confident summaries that are subtly wrong, which is worse than no summary.
Agent: 1. Over-reach on commitments. 2. Escalation into an unstaffed queue. 3. Deployed with the copilot's evaluation — judged on time saved rather than on what it said. 4. Mode enabled by default without anyone deciding.
Failure 3 on the agent side is the specific risk of the shared name: a store that evaluates a customer-facing agent the way it would evaluate a copilot has skipped escalation testing entirely — including the continuity checks in where the thread breaks and the boundary tested in rules-based versus conversational.
Who should be in the room?
Copilot: the salespeople. They decide whether it gets used, and they will tell you within ten minutes of a demo whether the drafts are usable.
Agent: the GM and whoever owns compliance. The decisions are about what the store may say and who is accountable — not about convenience.
Getting this wrong in either direction produces a bad purchase. Salespeople evaluating an agent will like the demo and not ask about escalation; a GM evaluating a copilot will approve something nobody uses.
What should you measure?
| Product | Metric | Target |
|---|---|---|
| Copilot | Weekly active users ÷ licensed | The whole evaluation |
| Copilot | Drafts sent unedited ÷ drafts generated | Whether the output is usable |
| Copilot | Time per response, before and after | The claimed benefit |
| Agent | Escalation accuracy | The boundary |
| Agent | Commitment violations | Structurally zero |
| Agent | Show rate on its appointments | The guardrail |
The split in this table is the article's point. There is no metric in common between the two, which is the clearest possible evidence that they are not one product.
Frequently asked questions
What is an AI sales assistant?
Two different products share the name: a copilot that drafts, summarises and suggests for your salesperson, and an autonomous agent that communicates directly with customers. They have different risks, different evaluations and different people who should be assessing them.
How do you tell which one a vendor is selling?
Ask whether anything reaches a customer without a person pressing send. If not, it is a copilot and adoption is the evaluation. If so, it is an agent and escalation accuracy is the evaluation. Many products are both with a setting, so also ask which mode ships enabled.
How should a copilot be evaluated?
On adoption, almost entirely. A copilot nobody uses produces nothing regardless of output quality, and the common causes of low adoption are drafts needing more editing than writing, living in a tool nobody has open, being slower than typing for short messages, and sounding nothing like the rep.
What are the genuinely useful copilot applications?
Summarising a long thread before a call, drafting a first reply the rep then edits, and surfacing what a customer said weeks ago. All three save minutes, none touches a customer directly, and all three sit inside the rep's existing workflow.
How should an agent be evaluated?
On escalation accuracy, on whether its limits are structural rather than instructional, on what a handoff carries, and on the show rate of the appointments it books. The important question is what it is prevented from doing rather than what it can do.
Who should evaluate each one?
Salespeople for a copilot, because they decide whether it gets used and will know within minutes of a demo. The GM and whoever owns compliance for an agent, because the decisions are about what the store may say and who is accountable.
What is the risk of the shared name?
That a customer-facing agent is evaluated the way a copilot would be — on time saved and convenience — which skips escalation testing entirely and deploys a system with no verified boundary in front of customers.
Do the two products share any metric?
None. A copilot is measured on adoption, unedited send rate and time saved. An agent is measured on escalation accuracy, commitment violations and show rate. The absence of any overlap is the clearest evidence they are different products.
Conclusion
- Two products, one name, and the checkpoint is the whole distinction.
- One question identifies it: does anything reach a customer without a person pressing send?
- Copilots are evaluated on adoption. Unused is worth nothing, however good the output.
- Agents are evaluated on what they are prevented from saying, and whether that is structural.
- No metric is shared between them. That is the evidence they are not one thing.
Last updated: