OpenLot Book audit
Systems

AI Implementation: A 30-Day Plan With an Abort Gate

OpenLot 9 min read

AI implementation at a dealership is a four-week sequence, not an install. Three of those weeks happen before the system talks to a customer, and the plan has an explicit abort gate on day 21 — because the cheapest failed deployment is the one you stop before go-live.

Four-week dealership AI implementation timeline showing data, specification, pilot and scale phases with an abort gate marked at day 21

This guide covers a four-week implementation sequence, what has to be true before go-live, the day-21 abort gate, where rollouts fail, and what to measure at each stage.

Why does implementation order matter more than the product?

Because most of what determines the outcome happens before the software is switched on, and it happens in a fixed order where each step depends on the one before it.

A system deployed against an unmeasured baseline cannot be evaluated. A system deployed against a dirty database amplifies the mess. A system deployed without written escalation rules will answer a payment question in week one. None of those are product problems, and none of them are fixable afterwards at the same price.

The sequence below is deliberately front-loaded. Week four is the easy week.

The four weeks

Illustrative. A group with multiple rooftops or a legacy DMS should expect six to eight.

Week Focus Exit condition
1 Baseline and data Current numbers captured; duplicate rate measured
2 Specification Escalation rules written; systems and fields named
3 Pilot on one source Running on real traffic, humans reviewing every thread
Day 21 Abort gate Go or stop, against criteria written in week 2
4 Scale and hand over Second source added; owner named; runbook delivered

Note where go-live sits. The system touches a real customer on day 15, not day 1, and you keep the option to stop six days later for the cost of three weeks rather than a year.

Week 1 — Baseline and data

Two jobs, neither of them technical.

Capture the baseline. Pull 90 days of lead timestamps and compute first-touch time bucketed by hour and weekday, contact rate, appointment rate and show rate. These are the only numbers that will ever let you say what the system did, and they are the same four that populate the ROI model. They cannot be reconstructed later — see the response time benchmarks for how to compute each one.

Measure the data, do not fix it yet. Count duplicate contacts, records with no valid phone, and records with no consent flag. You are establishing whether data decay is a cleanup task or the whole project. If the duplicate rate is high, the implementation stops here and becomes a data project — which is a cheaper discovery in week one than in week five.

Exit condition: the baseline numbers exist in a document, and you know your duplicate rate.

Week 2 — Specification

This is the week that decides whether the deployment works.

Name the systems and the fields. Which lead sources, which CRM, which fields are read and which are written. Latency as a number. If the sync is nightly batch, the fast-response promise cannot be kept and you need to know that now rather than in week three — see what integration work actually contains.

Write the escalation rules. The twelve or so triggers on which the system stops and hands to a person: price questions, trade values, out-the-door figures, payoff amounts, complaints, legal mentions, anything the system is not confident about. This document is the single most skipped artifact in dealership AI and the most expensive to add after customers have seen the gaps.

Write the abort criteria. Before the pilot starts, state what would make you stop on day 21. Specific and measurable: escalation accuracy below a threshold, any compliance incident, no measurable movement in first-touch time. Criteria written after the pilot are not criteria, they are rationalisations.

Exit condition: escalation rules and abort criteria exist in writing and are signed off by whoever owns the outcome.

Week 3 — Pilot on one source

One lead source. Real traffic. Every thread reviewed by a person.

One source, because you need to be able to attribute the result. Real traffic, because test leads do not behave like customers. Every thread reviewed, because week three is when you discover what the escalation rules missed, and a missed trigger found in week three costs a conversation while one found in month three costs a reputation.

Run it for seven days including a weekend. The weekend is the point — it is where the coverage gap actually lives, and a pilot that runs Monday to Friday has not tested the hours the system exists to cover.

Exit condition: seven days of real traffic with a reviewed transcript log.

Day 21 — The abort gate

Stop and decide, against the criteria written in week two.

Signal Go Stop
Escalation accuracy Triggers fired when they should have Missed price or trade questions
First-touch time, uncovered hours Collapsed to seconds Unchanged — integration is batch, or routing is wrong
Customer reaction Threads continue normally Complaints, or obvious-robot drop-offs
Compliance Consent and opt-out logged on every touch Any gap at all
Internal load Handoffs picked up promptly Engaged conversations queuing for hours

The bottom row is the one people ignore. If the system books appointments that nobody picks up, the bottleneck has moved rather than closed, and scaling makes the queue longer rather than the outcome better.

A stop at day 21 is a success condition for the plan. It cost three weeks.

Week 4 — Scale and hand over

Add the second source, not the remaining five. Each source has its own data shape and its own customer behaviour, and adding them one at a time keeps attribution intact.

Hand over properly. Credentials in your accounts, runbook written, monitoring and alerting live with a named recipient, and an internal owner with allocated hours. A deployment without a named owner degrades quietly — a vendor API changes, something stops, and nobody notices until lead flow goes quiet.

Exit condition: second source running, runbook delivered, owner named in writing.

Where do implementations fail?

1. Go-live on day one. Skipping the baseline makes the result unarguable in both directions, forever.

2. Escalation rules written after launch. The system answers a payment question in week one and the store spends the next month unwinding a commitment it did not make.

3. The pilot avoids the weekend. Testing coverage during covered hours tests nothing.

4. Scaling before the handoff is staffed. The front of the funnel gets faster and the middle stays the same width. Net effect is a longer queue.

5. No abort criteria. Without them, the day-21 decision becomes a negotiation about sunk cost, which has exactly one outcome.

What should you measure at each stage?

Stage Metric Target
Week 1 First-touch time by day-part Captured, not improved
Week 1 Duplicate contact rate Measured; above threshold means stop
Week 3 Escalation accuracy Every trigger that should have fired, did
Week 3 Handoff pickup time Minutes, not hours
Week 4 First-touch time, uncovered hours Collapsed to seconds
Week 4+ Appointment show rate Unchanged or better — not just more bookings

The last row guards against the most flattering failure mode: a system that books more appointments that fewer people attend.

Frequently asked questions

How long does AI implementation take at a dealership?

A narrowly scoped deployment against systems with real APIs runs about four weeks, of which three happen before the system contacts a customer. Groups with multiple rooftops, a legacy DMS or significant data remediation should plan for six to eight weeks, and the extra time is almost always consumed by data and integration access rather than by the AI itself.

What has to be done before go-live?

Three things: a captured baseline of first-touch time, contact rate, appointment rate and show rate; a measured duplicate rate in the CRM; and written escalation rules naming every trigger on which the system hands to a person. None of the three can be recovered afterwards at the same cost, and skipping the baseline makes the result permanently unarguable.

What is an abort gate and why day 21?

An abort gate is a scheduled decision point with criteria written before the pilot starts, so the go-or-stop call is made against a standard rather than against sunk cost. Day 21 sits one week after go-live, which is long enough to include a weekend of real traffic and short enough that stopping costs three weeks instead of a year.

Should the pilot run on all lead sources at once?

No. One source keeps attribution clean and limits the blast radius of anything the escalation rules missed. Add the second source in week four and the rest one at a time, because each source has its own data shape and its own customer behaviour.

Why does the pilot need to include a weekend?

Because the uncovered hours are the reason the system exists. A pilot running Monday to Friday during staffed hours tests the period a human was already handling and says nothing about nights, Sundays or holidays, which is where the measurable gain is concentrated.

Who should own the deployment inside the dealership?

A named person with allocated hours, usually the BDC manager or whoever owns the CRM, with operational authority sitting above them. Deployments without a named owner degrade quietly: an upstream API changes, a connection stops, and the gap is discovered weeks later from a dip in lead flow.

What if the data turns out to be too dirty?

Then the implementation stops and becomes a data project, and that is the correct outcome. Automation running against duplicate records contacts the same person from multiple threads, which turns an internal reporting problem into something customers experience directly. Finding this in week one is far cheaper than finding it in week five.

What is the most common reason a rollout fails?

Scaling the front of the funnel before the handoff is staffed. The system responds faster and books more, the engaged conversations queue for hours waiting for a person, and the net result is a longer queue rather than a better outcome. The fix is capacity at the handoff point, not more automation in front of it.

Conclusion

  • Three of the four weeks happen before a customer sees anything. That is the plan working, not the plan being slow.
  • The baseline cannot be recovered. Capture it in week one or accept that the result will never be defensible.
  • Escalation rules are the artifact that decides the deployment. Write them before the pilot, not after the first bad thread.
  • Day 21 is a real decision, with criteria written in week two. A stop is a success — it cost three weeks.
  • Scale one source at a time, and only after the handoff point is staffed to absorb what the system sends it.

Last updated: