OpenLot Book audit
Systems

AI Sales Assistant: Assisting Whom, Exactly?

OpenLot 9 min read

"AI sales assistant" describes two products that share almost nothing. One assists your salesperson and one replaces part of their job with a customer-facing system — and the risks, the evaluation and the person who should be in the room differ completely between them.

Copilot and autonomous agent compared as two products sold as an AI sales assistant for car dealers

This guide covers the two products, how to tell which is in a proposal, how each is evaluated, where each fails, and what to measure.

The two products

Copilot Autonomous agent
Who it talks to Your salesperson Your customer
Output A draft, a summary, a suggestion A sent message
Human checkpoint By construction Removed
Risk if wrong Someone notices and edits The customer received it
Evaluated on Time saved, adoption Escalation accuracy, show rate
Who should evaluate The salespeople The GM, and compliance

The checkpoint row is the whole distinction, and it is the same one that separates an assistant from an agent: one drafts and a person sends, the other sends.

How do you tell which is in a proposal?

One question: does anything reach a customer without a person pressing send?

If no, it is a copilot. Evaluate it on whether your salespeople use it, because an unused copilot is worth exactly nothing regardless of how good its output is.

If yes, it is an agent. Evaluate it on escalation accuracy and on what it is structurally prevented from saying.

Many products are both, with a setting. Ask which mode ships enabled, because the default is what you will be running.

Evaluating a copilot

Adoption is the whole evaluation

A copilot that is not used produces nothing, and copilots go unused for predictable reasons.

Reason Fix
The draft needs more editing than writing from scratch Better context, or abandon it
It lives in a tool they do not already have open Put it where they work
It is slower than typing A real problem at short messages
It sounds nothing like them Tone matters more than people expect
Nobody showed them twice Training, and a second session

Measure adoption weekly for the first month. A copilot at low adoption in week four will be at zero by week twelve, and the honest decision is to cancel rather than to wait.

The useful copilot applications in a dealership are narrow and real: summarising a long thread before a call, drafting a first reply the rep then edits, and surfacing what a customer said three weeks ago. All three save minutes and none touches a customer directly.

Evaluating an agent

Different evaluation entirely, and it is the one covered across the rest of this cluster: escalation accuracy, capability limits rather than instructions, handoff payload, and show rate as the guardrail.

The question that matters most for an agent is not what it can do. It is what it is prevented from doing, and whether that prevention is structural.

Where does each fail?

Copilot: 1. Low adoption, which is the dominant failure. 2. Output requiring more editing than writing. 3. Sitting in a separate tool. 4. Generating confident summaries that are subtly wrong, which is worse than no summary.

Agent: 1. Over-reach on commitments. 2. Escalation into an unstaffed queue. 3. Deployed with the copilot's evaluation — judged on time saved rather than on what it said. 4. Mode enabled by default without anyone deciding.

Failure 3 on the agent side is the specific risk of the shared name: a store that evaluates a customer-facing agent the way it would evaluate a copilot has skipped escalation testing entirely — including the continuity checks in where the thread breaks and the boundary tested in rules-based versus conversational.

Who should be in the room?

Copilot: the salespeople. They decide whether it gets used, and they will tell you within ten minutes of a demo whether the drafts are usable.

Agent: the GM and whoever owns compliance. The decisions are about what the store may say and who is accountable — not about convenience.

Getting this wrong in either direction produces a bad purchase. Salespeople evaluating an agent will like the demo and not ask about escalation; a GM evaluating a copilot will approve something nobody uses.

What should you measure?

Product Metric Target
Copilot Weekly active users ÷ licensed The whole evaluation
Copilot Drafts sent unedited ÷ drafts generated Whether the output is usable
Copilot Time per response, before and after The claimed benefit
Agent Escalation accuracy The boundary
Agent Commitment violations Structurally zero
Agent Show rate on its appointments The guardrail

The split in this table is the article's point. There is no metric in common between the two, which is the clearest possible evidence that they are not one product.

Frequently asked questions

What is an AI sales assistant?

Two different products share the name: a copilot that drafts, summarises and suggests for your salesperson, and an autonomous agent that communicates directly with customers. They have different risks, different evaluations and different people who should be assessing them.

How do you tell which one a vendor is selling?

Ask whether anything reaches a customer without a person pressing send. If not, it is a copilot and adoption is the evaluation. If so, it is an agent and escalation accuracy is the evaluation. Many products are both with a setting, so also ask which mode ships enabled.

How should a copilot be evaluated?

On adoption, almost entirely. A copilot nobody uses produces nothing regardless of output quality, and the common causes of low adoption are drafts needing more editing than writing, living in a tool nobody has open, being slower than typing for short messages, and sounding nothing like the rep.

What are the genuinely useful copilot applications?

Summarising a long thread before a call, drafting a first reply the rep then edits, and surfacing what a customer said weeks ago. All three save minutes, none touches a customer directly, and all three sit inside the rep's existing workflow.

How should an agent be evaluated?

On escalation accuracy, on whether its limits are structural rather than instructional, on what a handoff carries, and on the show rate of the appointments it books. The important question is what it is prevented from doing rather than what it can do.

Who should evaluate each one?

Salespeople for a copilot, because they decide whether it gets used and will know within minutes of a demo. The GM and whoever owns compliance for an agent, because the decisions are about what the store may say and who is accountable.

What is the risk of the shared name?

That a customer-facing agent is evaluated the way a copilot would be — on time saved and convenience — which skips escalation testing entirely and deploys a system with no verified boundary in front of customers.

Do the two products share any metric?

None. A copilot is measured on adoption, unedited send rate and time saved. An agent is measured on escalation accuracy, commitment violations and show rate. The absence of any overlap is the clearest evidence they are different products.

Conclusion

  • Two products, one name, and the checkpoint is the whole distinction.
  • One question identifies it: does anything reach a customer without a person pressing send?
  • Copilots are evaluated on adoption. Unused is worth nothing, however good the output.
  • Agents are evaluated on what they are prevented from saying, and whether that is structural.
  • No metric is shared between them. That is the evidence they are not one thing.

Last updated: