OpenLot Book audit
Systems

Does Your CRM Need AI, or Clean Data First?

OpenLot 9 min read

Running automation against a dirty database does not surface the problem — it broadcasts it. A duplicate that was a reporting annoyance becomes a customer contacted three times from three threads, and four checks, each taking under an hour, tell you whether yours is bad enough to fix before anything is switched on.

Four data quality checks that decide whether a dealership should clean its CRM before deploying automation

This guide covers why automation amplifies, the four checks with thresholds, what cleaning involves, the order of operations, and what to measure.

Why does automation amplify rather than reveal?

Because a person working leads by hand silently corrects for bad data and a system does not.

A rep who opens two records for the same customer notices and works one. A sequence does not notice: it sends from both, and the customer receives duplicate messages from the same store on the same day. The data was equally wrong in both cases; only one of them was visible to the customer.

The same applies across the board. A dead phone number costs a rep thirty seconds and costs an automated cadence nothing — except that the cadence will keep calling it for three weeks, and the activity log will show eight attempts that never had a chance.

This is the practical version of CRM data decay: the decay was always there, and automation changes who sees it.

The four checks

Each under an hour, and together they decide the sequence

Check How Threshold to worry
Duplicate rate Records sharing a phone or email ÷ total Above a few percent
Contactability Records with a valid phone or email ÷ total Below roughly 85%
Consent coverage Records with a dated, channel-specific consent ÷ total Below roughly 70%
Source integrity Records with a usable lead source ÷ total Below roughly 80%

The exact thresholds matter less than the shape. A store failing one check can usually proceed while fixing it. A store failing two or more should clean first, because automation will make all of them visible to customers simultaneously.

The consent row is different in kind from the other three. The others degrade performance; that one is an obligation, and deploying outbound automation against records with no consent history is an exposure rather than an inefficiency.

What does cleaning actually involve?

Four workstreams, in increasing order of difficulty.

1. Deduplication. Merging records that are the same human. Mechanically easy and judgment-heavy at the edges — two people at one address, or one person with a work and a personal number. Decide the merge rules before running anything.

2. Contactability. Validating phones and emails, marking the dead ones rather than deleting them. Deleting loses history; marking preserves it and stops the contact attempts.

3. Consent reconstruction. The hardest, and partly impossible. Where a dated, channel-specific consent record does not exist, the honest position is to treat it as absent rather than to infer it — the standard in SMS compliance.

4. Source integrity. Fixing lead source attribution going forward, which matters more than fixing it historically — the point made in lead attribution.

The order of operations

Not "clean everything, then automate." That takes a year and nothing improves in the meantime.

Step 1 — Run the four checks. An afternoon.

Step 2 — Deduplicate, and fix consent, on the segment you are about to automate. Not the whole database. If the first deployment is website-form leads from the last 90 days, clean that segment.

Step 3 — Deploy on the clean segment. The staged rollout in the lead-source sequence makes this natural.

Step 4 — Clean the next segment before adding it.

Step 5 — Fix the intake. Stop new duplicates and missing consent at the point of entry, otherwise the cleaning is a permanent job rather than a project.

Step 5 is the one that determines whether this is finished or recurring. A store that cleans its database and does not fix intake will be back in the same position within eighteen months.

Where does this go wrong?

1. Deploying first and cleaning later. The customers find the problem before you do.

2. Cleaning the whole database before deploying anything. A year of no improvement, and the dirt returns at the intake.

3. Deleting rather than marking. History disappears, and the same bad record gets re-created on the next enquiry.

4. Inferring consent. Treating an old interaction as permission is the wrong call, and it is an obligation question rather than a data one.

5. Not fixing intake. Makes it a permanent job.

6. Treating it as an IT project. The merge rules are commercial decisions about what counts as the same customer, and they need someone from the floor. They also depend on which system is authoritative for each field.

What should you measure?

Metric How to compute Threshold
Duplicate rate Shared phone or email ÷ records Track monthly, not once
Contactability Valid phone or email ÷ records Should rise after cleaning and hold
Consent coverage Dated, channel-specific ÷ records The obligation metric
New duplicates per month Created ÷ new records The intake test
Customer-visible duplicate contacts Reported or sampled Should be zero after cleaning
Source integrity on new leads Usable source ÷ new Going forward matters most

Row four is the one that distinguishes a project from a treadmill. If new duplicates are being created at the same rate after the clean-up, nothing structural changed and the whole exercise repeats next year.

Frequently asked questions

Does automation fix bad CRM data?

No, it broadcasts it. A rep who opens two records for the same customer notices and works one; an automated sequence sends from both. The data was equally wrong in both cases, and only one of them was visible to the customer.

How do you tell whether the data is too dirty to automate?

Four checks, each under an hour: duplicate rate, contactability, consent coverage and source integrity. Failing one is usually workable while you fix it. Failing two or more means cleaning first, because automation makes all of them visible to customers at once.

Which data problem is most urgent?

Consent coverage, because it is an obligation rather than an inefficiency. The other three degrade performance; deploying outbound automation against records with no dated, channel-specific consent history is an exposure of a different kind.

Should the whole database be cleaned before deploying?

No. That takes a year, nothing improves meanwhile, and the dirt returns through the intake anyway. Clean the segment you are about to automate, deploy on it, then clean the next segment before adding it.

Should bad records be deleted?

No, marked. Deleting loses the history attached to them and the same bad record is usually re-created on the customer's next enquiry. Marking stops the contact attempts while preserving what the record knows.

Can consent be reconstructed from old interactions?

Only where a dated, channel-specific record genuinely exists. Inferring permission from an old interaction is the wrong call, and the honest position where no record exists is to treat consent as absent and exclude the record from outbound automation.

What makes this a project rather than a permanent job?

Fixing the intake. A store that cleans its database and does not stop new duplicates and missing consent at the point of entry is back in the same position within about eighteen months, which is why new duplicates per month is the metric that matters most.

Is this an IT task?

Only partly. The merge rules are commercial decisions about what counts as the same customer — two people at one address, one person with two numbers — and they need someone who works the floor rather than only someone who works the database.

Conclusion

  • Automation broadcasts bad data rather than revealing it. The rep was silently correcting.
  • Four checks, one afternoon. Failing two or more means clean first.
  • Consent is different in kind. The other three cost performance; that one is an obligation.
  • Clean by segment, in step with the rollout. Not the whole database, not later.
  • Fix the intake, or it is a treadmill. New duplicates per month is the test.

Last updated: