It’s Not Just About the Data Mesh: The Importance of Data Mastering and What We’ve Learned From the Past

Editor’s Note: This post was originally published in November 2022. We’ve updated the content to reflect the latest information and best practices so you can stay up to date with the most relevant insights on the topic.
Data mesh was all the rage when we first published this post. Four years later, the hype has cooled, the practitioners have gotten better at it, and the hard part hasn’t changed. Even after more than 50 years of investment in systems for enterprise data quality, most sizable corporations struggle to break down data silos and publish clean, curated, comprehensive, and continuously updated data. One of the primary reasons companies still grapple with this issue is that they’re missing a critical component of the modern data ecosystem: data mastering.
The History of Data and the Evolution of its Ecosystem
Let’s get out our flux capacitor, hop into the DeLorean, and go back in time! We’re going to explore the 1990s and the early 2000s when most data professionals were focused on building enterprise data warehouses, an approach that was somewhat successful because a company’s data ecosystem existed as a large monolithic artifact within the organization, one that could be governed and contained.
Now let’s speed ahead a decade. Here you’ll witness the introduction of next-generation analytics tools such as Qlik, Tableau, and Domo geared to democratize the data ecosystem. The goal of these tools was to have analysts—rather than database administrators—dictating how data should be processed and consumed in a distributed manner. The assumption at this time was that data aggregation was ineffective.
Next, we’ll make a pit stop in the mid-2010s to witness how cloud infrastructure provided the ability to quickly scale storage and compute efficiently. Everyone wanted to aggregate their data in a data lake during this time—or at least move their data to the cloud first, then figure out how to use it.
One more stop before we head home: the 2020s. The data lake became the data lakehouse. Data mesh promised to hand ownership back to the domains. And then generative AI and AI agents showed up and started asking the data questions no dashboard had ever asked—at a volume no data steward could review by hand.
If Biff Can Be Tamed, a New Approach to the Data Ecosystem Can Be Embraced
We are back…no, not in 1985, and not in 2022, either. It’s 2026. The data ecosystem is still changing constantly, and data volume and variety are still exploding. Data is becoming more and more external and the best version of it often exists outside of the firewall, not in your organization’s ERP or CRM solution.
To deal with data silos in analytics use cases, there are four different strategies one can employ:
- Rationalization: Consolidating data from different systems into one
- Standardization: Creating consistent vocabularies and schemas and pushing them from one system to the rest
- Aggregation: Assembling all data into a central repository such as a data warehouse
- Federation: Storing and governing data in a distributed manner with interconnected data sources by domain
In any successful data project, all four strategies are necessary but not sufficient on their own. They require a centralized entity table and persistent universal ID linking data together. This is where data mastering comes in.
Zhamak Dehghani introduced the concept of data mesh while at Thoughtworks and now leads Nextdata, the company she founded to build it as a product. Dehghani originally defined data mesh as a new enterprise data architecture that embraces “the reality of ever-present, ubiquitous, and distributed nature of data.” And it has four aspirational principles:
- Data ownership by domain
- Data as a product
- Data available everywhere (self-serve)
- Data governed where it is
In this paradigm, the data is distributed and external. That’s why traditional, rules-based master data management (MDM) simply will not work. Instead, organizations need to start with an AI-native approach to MDM that uses advanced AI/ML models, select business rules, and agentic data curation. It is a cornerstone of a successful data mesh because it provides the centralized entity table and the persistent universal IDs that make distributed queries answerable.
Data Mastering Within the Context of Data Mesh Strategy
What’s inherent in data mesh is the belief that data is more distributed. Today, we need to think about data in terms of logical entities: customers, products, suppliers, providers—the list goes on. But I’ll let you in on a dirty little secret: Most companies actually have thousands of sources that provide data about these entities, making it difficult, if not impossible, to do a data mesh.
Companies embarking on a data mesh strategy will quickly realize that they need a consistent version of the best data across the organization. The only method of achieving this at scale is through AI-driven entity resolution and data mastering.
Four years on, that’s no longer a contrarian position. Thoughtworks’ own 2026 assessment concludes that “Data Mesh has evolved from industry hype into a mature socio-technical paradigm,” and that “clean, owned, product-based data with clear contracts are the essential foundation for success with trustworthy, production-grade AI and ML.” Clean, owned, contracted data is the point. Somebody still has to decide which customer is which.
Peanut Butter Sandwiches Taste WAY Better With Jelly!
Think about data mastering as a complement to data mesh. On their own, each produces a good result. But when combined, the results are spectacular. You can’t have peanut butter without jelly to make a truly perfect sandwich, and data mesh without data mastering is a dry offering—no jam to make it juicier and sweeter.
Further, in a modern data architecture, a well-built data mesh acts as the supply chain feeding domain-owned data—unified, cleaned, and organized using AI-native MDM—into the context layer so AI agents can use it to make recommendations or act autonomously.
When you apply AI-native data mastering, you clean up your internal and external data sources, engaging in a bi-directional cycle that allows you to cleanse and curate your data efficiently. You also effectively realize the promise of distributed data mesh. It’s a critical, continuous loop so that you can incorporate changes to your data—or your sources—over time.
That loop matters more now than it did in 2022. Gartner reported in February 2025 that 63% of organizations either do not have or are unsure if they have the right data management practices for AI, and predicted that through 2026, organizations would abandon 60% of AI projects unsupported by AI-ready data. And while a dashboard with duplicate customers is an annoyance, an AI agent with duplicate customers is a liability.
As you build your data mesh strategy—or whatever you’re calling it in 2026—remember that the key to success is starting with AI-native data mastering.
And should you require a refresher course to leverage learnings from past mistakes, don’t forget you need to be traveling at 88 miles per hour and generating a charge of 1.21 gigawatts. And look out for Doc—he can be found by the clock tower!
Get a free, no-obligation 30-minute demo of Tamr.
Discover how our AI-native MDM solution can help you master your data with ease!


