Why Are the Books on AI So Hard to Balance?
A learn article on why AI marketing gains are hard to verify in finance reviews, covering three measurement traps: counting activity instead of change, work shifting between roles, and gains and costs landing in different departments. It argues for measuring whole-workflow change.
A while back, a friend of mine in consumer goods invited me out for coffee. He runs marketing at a company.
Before the coffee even arrived, he pushed his phone across the table: Look — AI has been live for six months. Content ships faster, ad creative has multiplied, and the team is putting in less overtime. But at last week's business review, the CFO asked me one question, and it stopped me cold.
What question?
Where is the money you saved?
He had no answer. The output had visibly multiplied — yet finance simply wouldn't recognize it in the books.

Later I got hold of his numbers and broke them down, line by line.
The problem isn't AI. It's how the math gets done.
Look at two sets of numbers, and you'll see just how tangled this is.
The CMO Survey projects that by 2029, AI will drive more than half of all marketing activity in the US. Now look at what has already happened: in 2025, sales productivity rose 14.1%, customer satisfaction rose 10.8%, and marketing overhead fell 14.6%. Green across the board.
Witness.AI surveyed a group of executives on whether their AI projects actually made money. The result: only 9% dared to claim that more than three-quarters of their company's AI projects had delivered real financial returns.
Nine percent. What does that mean? Out of ten executives, you won't find a single one ready to pound his chest and vouch for it.
And 68% admitted that over the past 12 months, their AI projects ran over budget at least once.
The CMO side stings more. In 2026, Conviva ran a global survey of CMOs: only 16% of CMOs could confidently produce evidence that their AI investments were working; nearly 70% could not measure results precisely; 21% had no stable measurement system at all.
They hadn't even built the ruler.
On one side, green across the board. On the other, nine in ten executives who won't sign off on the gains.
What's going on?
The reason is actually simple: AI's impact cuts across far too many steps. I've grouped the common traps into three.

Trap One: Measuring the Wrong Thing
What does measuring wrong look like? Plenty of companies do the AI math like this: an article used to take 8 hours to write; now, with AI's help, it's done in 2 — 6 hours saved. Multiply the saved hours by labor cost, and there's your "ROI."
Anything wrong with that? Yes. It counts how much work AI did, not what changed because of AI.
More assets produced, faster turnaround, fewer person-hours — that only proves "operational efficiency improved." And between efficiency improving and money actually saved, there is still a long stretch of road.
So what should you measure? Change. What the workflow looked like before AI touched it, and what it looks like after. Only the difference between the before and the after has anything to do with money.
Measure change, not activity.
There's a counterintuitive rule here, too: the farther your claimed benefit sits from the work AI directly did, the harder it is to prove that AI caused it.
But hold on — even once the change is tallied, another trap is buried underneath: data.
With AI today, the hard part is no longer "connecting the model to a few more data sources." It's "who should the model believe?" Every extra system you connect exposes another batch of problems. The same field, defined differently in two systems. The same customer record, stored in three copies, nobody sure which is current. When two sides' data contradict each other — which one do you listen to?
Humans can go on gut feel. An AI agent cannot. It has to know what a field means, when it was last updated, which of several sources is authoritative, and what it is allowed to do with the data once it has it. Which is why your data architecture is, in essence, your AI architecture.
Oh, and one more thing — when you tally the costs, don't forget the side accounts: system integration, data cleaning, governance, monitoring, employee training, human review. These are all real costs of AI. Leave them out, and the ROI inflates all by itself.
Trap Two: The Work Didn't Disappear — It Just Moved
Picture this scenario.
AI generates 300 content variants overnight. Fast? Fast. Time saved? Not so fast.
Those 300 won't clean themselves up. Someone has to go through them one by one: Are the facts right? On brand? Duplicate of existing content? Any legal risk? By the end of the review loop, the time that was saved has been poured right back in.
Same with an AI agent. It cuts your research time by a chunk — but you spend new time reviewing its edge cases and watching that its decisions don't drift. A task that is 70% faster sounds lovely. But if the work merely moved from role A to role B, the company as a whole hasn't saved a cent.
So the unit of measurement has to shift from the "task" to the entire workflow.
Choosing tools works the same way. Some tools dazzle the whole room in the demo; in real use, employees spend their days shuttling data between systems, patching in context, fixing outputs, waiting on approvals. The stronger the features, the more the wrangling. Meanwhile, a plainer tool — less flashy, but one that slots smoothly into the workflow — ends up saving more time and more money.
A tool isn't good or bad depending on how many features it has. It's good or bad depending on whether it fits your team's hands.
Trap Three: The Gains and the Costs Sit in Different Pockets
This is the most hidden of the three traps.
On marketing's AI ledger, the gains are crystal clear: output is faster, person-hours are saved, and it is all booked to marketing. But the costs? Compute and infrastructure, paid by IT. System integration, done by engineering. Compliance and security, carried by legal and risk. The person-hours of human review, scattered across every team.
The gains are theirs alone; the costs are everyone's.
Do the math that way, and of course marketing's AI projects look great. Because part of the cost is sitting on somebody else's books.
So what do you do? Pull the camera back. Measure the change across the entire workflow. Charge each cost to the department where it actually occurs. And as for credit: only results that can reasonably be attributed to AI changes go on AI's ledger.
Finally, Back to That Cup of Coffee
Later, my friend changed how he made his case. No longer "how many person-hours AI saved" — instead, he pulled in colleagues from IT and legal and recalculated the entire content workflow, from cost to benefit.
The new numbers looked worse than the old ones.
But this time, the CFO nodded.
An ugly ledger is a credible ledger.
AI was never out to deceive you. The way the math was done is what inflated the books.
And one more wish for you: the next time you report AI's value, may every number in the books survive the follow-up questions.