Your Data Is Dirty, and the Faster AI Runs, the More You Lose
A learn article arguing that AI amplifies dirty marketing data rather than fixing it. It explains why regulated industries keep data disciplined, how customers end up paying for bad data, and offers a five-question backward audit plus a people-process-tools framework to clean data foundations before adopting AI.

One Phone Call
A couple of days ago, I had an hour-long phone call with an old friend who has spent more than twenty years working with data.
We used to build a data warehouse together at Tempur-Pedic. Even before that, he worked on one of the world's first "customer-centric" marketing platforms, did stints at Digitas and WPP, and ran global data at MRM. Today, he runs business and operations for iceDQ, a data-reliability company.
Put simply, his entire career has been about judging whether data can be trusted.
When I hung up, a chill ran down my spine.
These days, so many people in marketing use AI. Copywriting, audience selection — all of it gets handed over. But there is a habit of AI's that worries me more the longer I think about it.
It answers fast, and it answers with total confidence. Right or wrong, it never changes its expression.
And when it works, it feels like the future has arrived. When it gets it wrong? Copywriting is a fair game — you can see when something is off. But the moment you hand over audience selection, no one catches it. Until a customer calls and asks why that set of numbers doesn't match.
That false accounting had gone bad a long time ago.
AI, won't fix your data. It just takes an old problem in the data, and hands it back to you as a new, much bigger problem.
A Vocabulary for Good Data
Let me start with an old saying.
What do we call a solid understanding of marketing data?
It's simple: it's the quality of data you feed in, that defines what you get back.
Anyone working in data marketing learns this in lesson one. In the old world, the systems were slow, and you could watch them run, pull a lost record back if needed. AI is different. It's fast, it's confident, and it hands you the answer straight away — you don't have the time to verify it.
So the more urgency you feel about AI, the more important it is first to check the foundation.
And exactly what my friend worries about: many companies don't take the foundation seriously.
Some are Watched, Some are Not
Here's a conclusion for you. The more regulated the industry, the more disciplined the data.
Because they are afraid of being fined.
Banks, insurance. When the numbers break, there's the FINRA auditor in the building.
It's the same in health care, HIPAA is a line you don't cross. They fine you heavily, so you behave.
What about consumer and retail?
No one is watching. No one is writing you a fine.
So, the standard claim from marketing people goes like this:
- When the campaign under-delivers, it's the creative's fault.
- When the audience is wrong, it's the algorithm's fault.
- And when the system outputs something weird, it's the prompt in an email.
We pin each fault on someone else, but we never look at the raw data under the dashboard.
But when the AI is on, the worse the raw material, the more bitter the taste. It's not hidden anymore; it's amplified and put on a plate for you.
Who Suffers? Not IT — the Customer
Dirty data, sounds like an engineer's problem. In fact, no.
The customer is the first one asked to pay.
Think about these scenarios.
A customer who has been with you for ten years gets a new-customer coupon.
The same person gets three copies of the same email.
Someone just bought a blender, and now you push the same blender to them again.
For a customer, this isn't a connection problem; it's "you don't know me." Once the vibe stuck, it is hard to beat.
This is the cost, and it's living in technical, not the executive floor. It needs to be on the CMO's desk.
One Word That Tells You Data Is Bad
How to spot a bad data? My old friend gave me a word.
"Unpredictable."
Data that today says one thing, the tomorrow something else. That's a red flag.
He told me about a client last year. A campaign that did not match. The team claimed to send it to 3 million people. Who received exactly which email — they couldn't say. Whether the send log matches the original audience file — they couldn't say. Segmentation and strategy were all right. They were stuck on one sentence: Do I really send what I was supposed to send?
That's the biggest red flag, because a model trained on such data to study the wrong pattern. And it replicates itself to everywhere.
A Water Pipe to Explain the Four Layers
He made a comparison, and I wrote it down. Not a database. A water plant.
First, the reservoir. Water fills it to the rim — this is the moment the data enters.
Then the water in the pipe network, purified and filtered. Here, you build audiences within clean data.
From here, the water flows into every household through the pipes. In each home, there's an employee, a tool. The water keeps changing along the way.
The last stop: the tap. Reports and dashboards. Most marketers only test the last one.
If you want clean water, think about it: do you only measure the one output you drink at?
No. You check the entire pipeline: every segment pressure, the flow. Instantly you see where it's clogged.
Twenty Years of a Question. Don't Blame Technology
"A Full Picture of the Customer" is what marketers have been waiting for for nearly twenty years.
Why hasn't the customer data platform (CDP) ever really materialized?
My friend said it in one word: like a cold bucket of water.
What's blocking, isn't technology.
It's "people, processes, and platforms". Most companies have not changed any of them.
A real CDP, there are three things: collect the complete dataset, define the standards for data collection, then implement applied engineering thinking so you can combine identity and tag data from various teams into a coherent view of your customer. If you miss one of the three, you can throw in the most expensive system, and it's still running in circles.
Two Methods to Use Today
The first method: work backward.
Even if you're not a data team, you can get to the bottom of matters within a morning.
Take a recent campaign. Trace back: start with what the customer saw, and end with where the data came from.
Five questions, answered in order:
- What exactly did we send?
- Does it match the audience data we sent to the platform?
- Does that file align with the plan? If you hesitate even a moment, the first break in the process reveals itself.
- From here to the source, what has happened to the data? Cleaning, dedup, engineering — each step needs a person who can answer it.
- Where did it come from, and how do we prove that it reached its destination?
Once you've done this, add three columns for each step: Who is watching the step day to day? Who's responsibility the step? Who's accountable when something goes wrong?
When you've completed the round, most teams discover the same thing: almost all the middle column is blank, and the third column is completely blank. This empty table is your action plan. Requires no budget, no permission, just one room and two hours.
The second method: people-first, processes-next, and tools-last, layer by layer.
Take people first. Many people miss not the measure, but the question: who is responsible for data quality.
A Chief Data Officer is of course ideal. But applications, development, data, and business, are often siloed in four separate walls. Somebody writes code, nobody checks the quality. Business assumes data is correct, yet the requirements were never written down.
Two paths beside: capturing requirements about the data that it's testable, and updating them when the data drifts. And get someone in to translate business to engineering, bridging what each other says.
Next, the processes. Here I'll take a page from the factory. The order can't change.
Test before your release. New tools, new platform, new task — nothing goes live without tests.
Run it, and watch continuously. Data is like a river; you expect the issues to appear on the dashboard in due course.
The results. That's a nice step, but if the first two are done, 99.9% of dirty data would never get that far.
Toyota is Toyota. that's how they folded all their Work done.
For tools, here's the thing: collecting many? And each vendor has a definition... in fact, the whole picture: many vendors, tangled mess.
Stop asking "Which tool fits best?" and ask: "For what goal am I responsible?" "Where does data break in the pipeline?" "Does this? It fills the exact gap?"
Either way, when you need to rebuild the whole, the same principle stands: It's not the tool.
Finally
After a whole night of listening, I found myself wanting to take notes.
He said, in the market two years from now, marketing teams that stand will all have done one thing differently: fixed the data first, and only then had the AI. Not the other way around.
Their models are trained on real data, not on the closest pretend version.
So I'll end with a reminder for you — and for me.
Don't jump to AI right away. First, look at the data underneath you. Is it clean? Fix the water at the beginning of the river.
Tonight, you can find every number that doesn't match, and go through it one by one.
I will.