
006
The Data Trap
Woody Bendle
Founder, LiftConductor
The Data Trap
When I began building predictive models in the 1990s, it was easy to believe more data would lead to better decisions.
Every new dataset promised another piece of the puzzle. Every additional variable seemed like another opportunity to understand customers, choose better store locations, or make smarter decisions.
At Blockbuster and later at Payless ShoeSource, we invested heavily in data.
We licensed extensive household-level demographic data from Acxiom, along with vehicle ownership data from RL Polk. We purchased demographic appends, household characteristics, segmentation systems, and lifestyle indicators. We hired respected consulting firms to develop customer segments and needs-based classifications. We built dedicated analytics infrastructure because our models demanded more computing power than our production systems could comfortably support.
It wasn't unusual to work with hundreds of household-level variables across millions of households.
Leadership wasn't chasing technology for its own sake.
They were investing in understanding.
And we learned an extraordinary amount.
Just not always what we expected.
For a while, our philosophy was simple.
If some data is good...
More data must be better.
It sounds perfectly reasonable.
Until you've lived through it.
Most of the variables were interesting.
Only some of them proved genuinely useful.
That distinction became one of the most valuable lessons of my career.
Some variables quietly earned their place.
One demographic measure, known as the HOB Index, identified neighborhoods with higher concentrations of manufactured housing. It consistently improved our understanding of certain trade areas - not because anyone expected it to, but because the evidence repeatedly demonstrated its predictive value.
Another surprise came from RL Polk vehicle data.
One of our strongest predictive variables wasn't household income or wealth.
It was a measure we created around the number of engine cylinders owned by a household - cars, boats, motorcycles, jet skis, snowmobiles, and other recreational vehicles.
It wasn't measuring income directly.
It was acting as a proxy for lifestyle and discretionary spending.
In several of our models, it contributed more predictive value than reported household income alone.
We didn't keep those variables because they sounded clever.
We kept them because they consistently improved our predictions.
Later, when we built models to predict customer response, future activity, or churn, another pattern emerged.
The household-level demographic information still helped.
But only at the margin.
The real predictive horsepower came from something much simpler.
Recency.
Frequency.
Monetary value.
Classic RFM-like metrics and analysis consistently explained customer behavior better than hundreds of demographic attributes ever could.
The demographics became frosting on the cake.
Helpful.
Worth having.
Paid for itself.
But they didn’t carry the model.
The more we learned, the less interested we became in accumulating data.
Instead, we became interested in discovering which information actually made a difference and affected decisions.
Organizations often believe they have a data problem.
More often, they have a discernment problem.
Information is abundant. Companies are swimming in data.
But, evidence is selective.
And wisdom comes from knowing the difference.
Looking back, I don't remember most of the hundreds of variables we licensed and analyzed.
I remember the handful that consistently changed our decisions.
That's the difference between collecting information and building understanding.
Reflection
The most valuable information isn't the information you have.
It's the information that consistently improves your decisions.