For private equity investors, the rise of AI has made one diligence question increasingly important: does a company’s proprietary data create a durable competitive advantage, or does it simply make the business look stronger than it really is?
A recent paper by Erwan Morellec and Francesca Zucchi, “Creative Destruction in a Data Economy”, offers a useful framework for answering that question. The paper examines how proprietary data, AI, and computing capacity interact to shape competition between incumbents and new entrants. Its core insight is particularly relevant for pre-deal diligence: data can simultaneously strengthen incumbents and increase the disruptive power of the few challengers that manage to overcome the entry barrier.
That distinction matters because many investment theses still treat data ownership as an almost automatic moat. In reality, the competitive outcome is more nuanced.

The authors model an economy in which incumbents accumulate proprietary data through normal business activity. More customers generate more interactions, transactions, workflows, and operational information. That data improves the productivity of R&D when paired with sufficient computing capacity, creating a reinforcing loop:
Scale → Data → Better AI and R&D Productivity → More Innovation → More Scale.
For an incumbent, this can be highly attractive. A company with a large installed customer base may continuously improve products, pricing, underwriting, recommendations, workflows, or automation using data that competitors cannot easily reproduce.
From a PE perspective, that is exactly the type of mechanism that can support margin expansion, higher retention, stronger pricing power, and multiple durability.
But the paper also highlights the other side of the equation.
Proprietary Data Raises the Barrier to Entry
If proprietary datasets are difficult or expensive to replicate, new entrants face a meaningful disadvantage. They may have strong technology, strong management teams, or better product design, but still lack the historical data needed to make their AI systems effective.
That can reduce the number of viable competitors entering the market.
For investors, this may initially look like confirmation of the moat thesis. Fewer entrants should mean less competition.
But Morellec and Zucchi show why that conclusion can be misleading.
The firms that do overcome the data barrier can become significantly more innovative. In other words, the number of entrants may decline while the innovation intensity of surviving entrants rises.
That creates an important paradox for diligence teams: a market can look stable based on startup formation while simultaneously becoming more vulnerable to technological disruption.
The relevant question is therefore not simply:
How many competitors are entering the market?
It is:
How capable are the strongest emerging competitors once they gain access to the required data and compute?
Why This Changes Pre-Deal Diligence
Traditional commercial diligence often focuses on market share, competitor count, pricing, customer retention, and historical barriers to entry. Those remain essential, but AI introduces another layer.
A company may face almost no credible competitors today because the data requirements for entering its market are unusually high. That can create apparent competitive insulation.
However, investors should distinguish between structural scarcity of data and temporary scarcity of access.
If the target’s advantage comes from data that cannot realistically be recreated, bought, licensed, scraped, generated synthetically, or obtained through partnerships, the moat may be durable.
If challengers can obtain functionally equivalent data through alternative channels, the barrier could erode much faster than historical market structure suggests.
This is especially important in vertical software, healthcare, financial services, insurance, logistics, marketplaces, and other sectors where proprietary datasets are often central to the investment thesis.
Five Questions PE Investors Should Ask
The paper suggests a stronger diligence framework around data-driven competitive advantage.
First, where does the data actually come from? Investors should understand whether the target’s dataset is generated organically through customer activity, purchased externally, manually collected, or derived from third-party systems.
Second, is the data genuinely unique? A large dataset is not necessarily differentiated. The important question is whether competitors can obtain economically equivalent information elsewhere.
Third, does more data materially improve the product? If an additional million observations barely improve forecasting, automation, or customer outcomes, the theoretical data moat may have limited commercial relevance.
Fourth, what happens when competitors gain sufficient scale? Diligence teams should model the threshold at which a challenger begins to achieve comparable AI or R&D productivity.
Finally, who is the most dangerous entrant—not the average entrant? This may be the biggest implication of the paper for PE investors. Competitive risk should not be measured solely by the number of startups in the market. One well-capitalized challenger with access to strong datasets and compute can matter more than dozens of underfunded entrants.
The Investment Takeaway
Morellec and Zucchi’s framework challenges a common assumption in technology-enabled investing: that proprietary data mechanically leads to winner-take-all outcomes.
It can create real barriers to entry. It can also strengthen incumbent R&D and reinforce scale advantages.
But those same barriers may produce a smaller group of highly capable challengers whose innovation intensity is significantly higher.
For PE investors, that means data diligence should move beyond questions of ownership and dataset size. The real focus should be on replicability, access, marginal value, and the innovation capacity of potential challengers.
The most attractive businesses may not simply be those with the most data.
They may be the businesses whose data advantage is hardest to reproduce—and whose management teams are converting that advantage into innovation faster than competitors can close the gap.
Source: Morellec, E., & Zucchi, F. (2026). Creative Destruction in a Data Economy. SSRN Working Paper. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7204859

