My latest book, The Cost of “Good Enough” Data: Why Modern Architectures Fail at Scale, examines a problem that has been building quietly across enterprise IT for years: we have become increasingly willing to accept data that is merely “good enough.” And that approach is becoming increasingly expensive.
Organizations today are investing heavily in AI, analytics, cloud computing, data lakes, data lakehouses, streaming platforms, and increasingly complicated data pipelines. Yet underneath all of this technology lies something much more fundamental: the data itself. If that data is incomplete, inconsistent, duplicated, poorly governed, stale, or inadequately understood, no amount of sophisticated technology layered on top of it can completely compensate. Indeed, AI has changed the economics of data quality because errors that once affected a report or dashboard can now propagate through automated recommendations, decisions, and business processes.
This is where the mainframe perspective becomes particularly interesting.
Those of us who have spent our careers working with Db2 for z/OS and other mainframe technologies grew up in an environment where data management was taken very seriously. Mission-critical systems could not routinely tolerate questionable integrity, undocumented changes, unreliable recovery, uncontrolled access, or sloppy transaction processing. The mainframe disciplines surrounding ACID transactions, database administration, security, recovery, change management, performance engineering, and operational control evolved because the workloads demanded them.
Of course, mainframe environments are not perfect. Bad designs, poor SQL, inadequate governance, and questionable data can exist anywhere. But there is something important to learn from the engineering culture that developed around systems of record.
Accuracy matters. Consistency matters. Recovery matters. Lineage matters. Governance matters. And somebody has to be accountable for the data.
As enterprises distributed their data across warehouses, lakes, cloud platforms, SaaS applications, analytical databases, streaming systems, and countless pipelines, some of that discipline was weakened or lost. Copies multiplied. Ownership became less clear. Transformations accumulated. Metadata became inconsistent. And organizations accumulated what I refer to as data debt — the growing cost and risk created by compromises that seemed reasonable when they were originally made. The book examines how that debt can undermine AI initiatives, increase operational costs, introduce compliance risk, and impair business decision-making.
AI is now exposing that debt.
An AI model does not magically know which customer record is authoritative, whether replicated Db2 data is current, whether a transformation changed the meaning of a field, or whether a piece of information satisfies the organization's governance requirements. It consumes what the data infrastructure provides.
And this brings us back to the mainframe.
For many large enterprises, some of their most valuable, accurate, and carefully controlled business data still resides in Db2 for z/OS and other IBM Z systems. That should not be viewed as an impediment to modernization. Quite the opposite. Trusted systems of record can become an enormously valuable foundation for analytics and AI — provided organizations build architectures that preserve the integrity, context, security, and governance of that data as it moves into newer environments.
That is really one of the central themes of The Cost of “Good Enough” Data. The book is not an argument against the cloud, or any of the newer distributed architectures. It is an argument against complacency. The technology has changed, the scale has changed, and the pace has certainly changed, but the fundamental principles of sound data management have not.
Modernization should not mean abandoning the disciplines we learned building reliable enterprise systems. It should mean applying those disciplines to modern architectures.
Because whether your data resides in Db2 for z/OS, PostgreSQL, a cloud database, a data lakehouse, or all of them at once, the same basic truth applies:
If the data cannot be trusted, neither can anything you build with it.
And in the age of AI, settling for “good enough” data may be more costly than ever.