In my role as data quality advisor, I get the chance to meet with organizations of all shapes, sizes, and purposes. One overwhelming, common challenge for non-profits and government agencies: good people who want to do more for their constituencies are prevented from doing so – simply because they do not have the budget. I feel for them. It’s hard to feel so hemmed in by such constraints. It is doubly hard when you know you could do more to educate, feed, protect, and help people.
Small Data Quality Problems Add Up to Massive Inefficiencies
As a data quality advisor, however, I see things a little differently. Through the data quality lens, most organizations, in both the public and private sectors, are enormously inefficient. Here’s why: As the Figure below depicts, every department (i.e., Department B) depends on data created by another (i.e., Department A) to complete its work. But Department A, lacking any visibility into how B uses the data, creates and passes on data that are simply not fit for use. So, Department B must spend time and effort making that data fit for use. It is hard, time-consuming work, often completed under enormous pressure. In some cases, Department B doesn’t get things right, and errors propagate to Department C.
I personally experienced an example just recently. I had applied on-line for some government program and was advised that such decisions usually took 30 days. After those 30 days had expired, I checked in on progress every couple of weeks. Still not complete.
After about 70 days, I called the Help Desk. The person who answered was nice and helpful. After a few minutes research, she advised that I had been mis-informed. I should not expect a decision for at least another month.
In some respects, this is a small matter—most data quality issues are. I spent about ten minutes each time I checked, fifteen minutes waiting in queue, and ten minutes with the agent. Maybe 45 minutes altogether for four touchpoints. The agent spent about ten minutes. Note that all this work is “non-value-added,” in the sense that neither myself nor the agency are any better off.
Identifying and Eliminating “Hidden Data Factories”
I call this non-value-added work the “hidden data factory.” There are of course data errors that take a very long time to deal with, but it is in these sorts of “ten minutes here,” and “an hour there” that make hidden data factories so large.
To be clear, I am NOT pointing the finger at this agency. I find hidden data factories all over. In the private sector, Salespeople spend lots of their time correcting data sent to them by Marketing, and Operations spends lots of time correcting data created by Sales. Often the effects make it to the Customer Service Department or its equivalent (e.g., the Help Desk). Based on a McKinsey Global Data Transformation Survey, use 30% as a “getting started” estimate of the fraction of time organizations spend on this non-value added work (e.g., in hidden data factories).
So far anyway, I’ve not seen any real difference in this pattern between government, non-profit and for-profit organizations. This 30% provides an obvious way to free up resources: To do more, free up funds you need by ruthlessly searching for and eliminating hidden data factories! This option is open to all, but it should be especially attractive for government agencies and non-profits, which are hemmed in by their fixed budgets. (For-profits have budgets, but appear to have many ways to work around them).
There is no real secret here: To reduce and eliminate hidden data factories, quit making so many errors upstream. I’ve explained how to do so in many places, so no need to repeat them here. It really does work!
Before closing, I want to make one further point: Coming to work each day only to spend a meaningful portion of your time cleaning up other people’s bad data is no fun. It’s far more enjoyable to attack the root causes of hidden data factories. And it frees up the resources needed to do more for your constituents!
Opportunity for Federal CDOs
For Federal Chief Data Officers (CDOs), data quality should be spelled O-P-P-O-R-T-U-N-I-T-Y. They can do their agencies a great service by focusing on data quality as their first priority by explicitly identify and eliminating hidden data factories. The resulting efficiencies will lead substantial cost savings, enabling the agency’s senior leaders to “do more.”






0 Comments