A group rolling out AI makes one decision that determines everything afterwards: which store goes first. The instinct is to pick the best performer or the worst, and both are wrong — the first proves nothing transferable and the second fails for reasons that have nothing to do with the software.
This guide covers why the obvious pilot choices fail, the four selection criteria, what must be true before store two, how to avoid a per-store re-negotiation, and what to measure.
Why the obvious choices fail
The best-performing store. Already has the process discipline, the staffing and the manager who follows up. A tool deployed there shows a smaller gain than it would anywhere else, and the other stores correctly point out that the result came from the store rather than the software. You have proven nothing transferable.
The worst-performing store. Fails for reasons the tool does not address — a staffing gap, a manager problem, an inventory position. The deployment then owns the failure, and the group concludes the category does not work.
The one that volunteered. Better than both, and still risky if they volunteered because their numbers are unusual in either direction.
The pilot's job is not to produce the best result. It is to produce a transferable one.
The four criteria
| Criterion | What it means | Why |
|---|---|---|
| Representative | Median volume, median staffing, typical lead mix | The result has to generalise |
| Measurable | Clean CRM, exportable timestamps, logged outcomes | No baseline, no evidence |
| Willing | A GM who wants it, not one who was assigned it | Pilots die of indifference, not of software |
| Stable | No manager change, no renovation, no DMS migration in the window | Confounders destroy attribution |
Scoring your rooftops
Score every store 1–3 on each criterion and total. The highest total pilots.
Representative Measurable Willing Stable Total Store A (best performer) 1 3 2 3 9 Store B (median) 3 2 3 3 11 Store C (worst) 1 1 2 1 5 Store D (median, new GM) 3 2 2 1 8 Store B wins on a criterion nobody usually scores — being ordinary. That is exactly what makes the result portable to the other rooftops.
Note store D: median and willing enough, disqualified by a manager change inside the measurement window. Stability is not a tie-breaker, it is a gate.
What must be true before store two?
Four things, and skipping any of them is how a group ends up with five separate deployments instead of one rollout.
1. The baseline was captured and the result is attributable. The four numbers from the lead response baseline, measured before and after, with no confounding change in the window.
2. The escalation rules are written and have survived contact. A trigger list that has been revised after reading real transcripts is worth far more than one written in a meeting.
3. The handoff is staffed. If the pilot store's gains were absorbed by an unattended queue, replicating it just replicates the queue — the failure covered in the handoff.
4. The integration is documented, per system. What was connected, which fields, which direction, what broke. Store two should not rediscover this — and the evaluation answers given before the pilot should be scored against what actually happened.
How do you avoid re-negotiating at every store?
This is where group rollouts actually stall — not technically, but commercially and politically.
Settle per-rooftop pricing before the pilot. Negotiate the price for all stores at the start, contingent on the pilot succeeding. A vendor negotiating store by store has every incentive to let the rollout proceed slowly, and per-rooftop pricing compounds across a group.
Decide what is standard and what is local. Escalation rules, consent handling and audit requirements should be group standards. Message copy, hours and vehicle-specific language should be local. Writing that split down before store two prevents the same argument happening five times.
Name a group owner. One person accountable for the rollout across stores, with the pilot GM as the operational owner of the first deployment. Without this, each store treats it as a new decision.
Publish the pilot numbers internally. The before and after, including what did not improve. A group that sees an honest result accepts the next deployment; one that sees a vendor case study resists it.
Where do group rollouts stall?
1. Pilot at the flagship. Covered above. Proves nothing portable.
2. No group-wide pricing agreed first. Each store becomes a negotiation.
3. Everything made local. Five variants of the escalation rules, none of them auditable.
4. Everything made standard. A message written for one market deployed in five, reading as generic in four of them.
5. Rolling out all stores at once after a successful pilot. Store two should be a second test of transferability, not the start of a simultaneous deployment.
6. No group owner. The most common political failure, and the one that makes the rollout stop after store two.
What should you measure?
| Metric | Where | What it decides |
|---|---|---|
| The four baseline numbers | Pilot, before and after | Whether it worked at all |
| Same four at store two | Before and after | Whether it transfers |
| Variance between stores | Compare results | Whether the result is about the tool or the store |
| Integration effort, store two vs one | Hours | Whether the documentation was real |
| Time from pilot end to store two live | Days | The rollout's actual speed |
| Per-rooftop cost, all in | Licence + integration + internal hours | What full rollout actually costs |
Row three is the one that decides the rollout. If store two produces a similar gain, you have a programme. If it produces nothing, the pilot result belonged to the pilot store, and the honest response is to find out why before store three.
Frequently asked questions
Which store should pilot a dealer group AI rollout?
The most representative one — median volume, median staffing, typical lead mix — provided it also has clean enough data to measure, a GM who actually wants it, and no major change planned during the window. The pilot's job is a transferable result rather than the best one.
Why not pilot at the best-performing store?
Because it already has the process discipline the tool is partly substituting for, so the gain is smaller than it would be elsewhere and the other stores can reasonably attribute the result to the store rather than the software. Nothing portable has been proven.
Why not pilot at the worst-performing store?
Because it usually fails for reasons the tool does not address — staffing, management, inventory position — and the deployment then owns the failure. The group concludes the category does not work when the pilot never tested the category.
What has to be true before rolling out to store two?
A captured baseline with an attributable result, escalation rules that have been revised after reading real transcripts, a staffed handoff so the gains are not absorbed by a queue, and documented integration details so store two does not rediscover them.
How do you avoid negotiating pricing at every store?
Settle per-rooftop pricing for the whole group before the pilot, contingent on it succeeding. A vendor negotiating store by store has every incentive to let the rollout proceed slowly, and per-rooftop costs compound quickly across a group.
What should be standard across stores and what should be local?
Escalation rules, consent handling and audit requirements should be group standards, because they are compliance and quality matters. Message copy, hours and vehicle-specific language should be local, because a message written for one market reads as generic in the others.
Should all remaining stores go live at once after a successful pilot?
No. Store two should be a second test of whether the result transfers, not the beginning of a simultaneous deployment. If store two produces a similar gain you have a programme; if it produces nothing, the pilot result belonged to the pilot store.
What makes group rollouts stall politically?
The absence of a single group owner accountable across stores. Without one, each rooftop treats the deployment as a fresh decision, the pricing is renegotiated, the rules fork, and the rollout quietly stops after the second store.
Conclusion
- Pick the ordinary store. The pilot's job is a transferable result, not the best one.
- Four criteria: representative, measurable, willing, stable. Stability is a gate, not a tie-breaker.
- Settle group pricing before the pilot, or every store becomes a negotiation.
- Standardise escalation and compliance, localise copy. Write the split down before store two.
- Store two is a second test, and the variance between the two is what decides the rollout.
Last updated: