Data Assets - Underappreciated beneficiaries of artificial intelligence

The Quick Take

  • AI had disrupted information services businesses, and investors are thinking deeply about what makes a company AI-resilient
  • Companies that derive earnings from hard-to-replicate, embedded datasets look set to thrive in an AI-driven environment
  • These companies will likely benefit from improved cost efficiencies and broader distribution channels; there will be a demand for premium AI-ready data

Danie Pretorius is an investment analyst with 18 years of investment industry experience.

The meteoric rise of artificial intelligence (AI) has been the defining theme in global equity markets for the past few years, with investors agonising over how to separate winners from losers. Much of the value to date has accrued to the providers of "picks-and-shovels". For instance, while the shares of companies in the semiconductor supply chain have outperformed, other industries have been recast as potential losers. In fields as diverse as software, consulting services, and wealth management, fears that AI could significantly lower barriers to entry for new competitors, or render existing product offerings obsolete, have led to meaningful share price declines.

Similarly, businesses with heavy reliance on the sale of data have fallen victim to the narrative that AI will erode the value of these data assets. Information services businesses that provide data across a range of sectors, including financial data, credit bureaus, rating agencies, legal, accounting, insurance, academic publishing, and healthcare, have all been tarred with the same brush. As shown in Figure 1, an equal-weighted basket of 20 information services businesses has underperformed the broader market by 50% over the past year.

Fig 1_Information Services Share Price Underperformance_V3.png

NOT ALL DATASETS ARE EQUAL

The market's fear that some data assets could be disrupted is rational, but lacks nuance. Businesses that rely on the sale of non-proprietary or easily replicable information, and where the cost of substitution is low, are indeed at risk, in our view. In fact, some of these businesses are already feeling the impact.

For instance, Chegg Inc., a provider of education services such as study packs and homework assistance for high school pupils and students, reported a revenue decline of close to 50% in the first quarter of 2026 (Q1-26). Chegg's datasets, consisting of decades of human-generated question-and-answer banks, have been proven replicable by foundational AI models at a much lower cost, resulting in a permanent impairment to the value of the business.

While there are many different factors that determine the value of any specific dataset, we find it helpful to think about the moats around these assets along two dimensions, as shown in Figure 2. We believe that resilience to AI disruption is a function of both dimensions being in play: a dataset ideally needs to be difficult to replicate and hard to substitute away from.

Fig2_Data-Moats-Proprietariness-Embeddedness_v2.png

The first is the Degree of Proprietariness. We use this to describe the extent to which the data is truly unique or replicable.

Data that is internally generated (e.g., credit ratings), sourced through regulated contributory networks (e.g., credit bureaus), sourced through proprietary physical infrastructure (e.g., real-time pricing data from financial exchanges), or otherwise cannot be reconstructed from alternative sources (e.g., time-stamped tick history) are highly proprietary. On the other hand, datasets that merely repackage or redistribute publicly available information (such as financial filings), are not proprietary.

The second dimension is the Degree of Embeddedness, which describes the extent to which switching away from a dataset is costly or disruptive to the end user, regardless of how replicable the underlying data itself is.

Some datasets enjoy far higher switching costs than others. Brand value, network effects, regulation, and integration with existing business processes are important drivers of value for various incumbent data providers. For instance, in the public debt markets, obtaining a credit rating from Moody's or S&P is a de facto requirement (in some cases it is an actual requirement). Indices from MSCI, S&P Dow Jones, and FTSE Russell are effectively industry standards, benefitting from the deepest liquidity pools. Evaluated pricing data from Platts is written into contracts for delivery in the energy markets.

Datasets that are both highly proprietary and deeply embedded are structural winners that are, in our view, likely to get stronger, not weaker, in the age of AI.

The challenge for investors is twofold. Firstly, assessing where in the matrix a particular dataset falls is a subjective exercise. Moreover, very few companies fall neatly into only one quadrant. Rather, most businesses have portfolios of assets of varying quality.

WINNERS HAVE DEEP MOATS

We hold several such businesses in our Global Equity Fund. In each, the market is overly concerned about a relatively minor exposure to replaceable data, while overlooking assets that sit firmly in the top-right quadrant, and each trades well below our assessment of its intrinsic value.

For instance, S&P Global is a leading financial information services business. In our assessment, the group's earnings overwhelmingly derive from highly proprietary and deeply embedded businesses. For instance, flagship assets like S&P credit ratings, S&P Dow Jones indices, and Platts are industry standards with deep moats. Smaller divisions, like pricing and reference data for credit default swaps (CDS) serve narrower niches but enjoy similar market positions (hard-to-replicate contributory network and deeply embedded in regulated workflows). A small part of the business relies on distributing undifferentiated data (e.g., financial statements and transcripts delivered through its Capital IQ desktop terminal) that is, in our view, disproportionately visible to equity investors (who are often customers) but relatively immaterial to earnings.

Similarly, LSE Group (LSEG) has sold off over concerns that its Refinitiv Workspace desktop terminal could be at risk from AI. However, the bulk of LSEG's business is proprietary, deeply embedded, and highly regulated. More than half of LSEG's earnings derive from trading venues, clearing houses, and indices that are not at risk from AI disruption. Within its Data and Feeds business, the vast majority of revenue derives from real-time data (that requires physical network connectivity to nearly 600 trading venues around the world) and other exclusive datasets (e.g., deals league tables and Tradeweb fixed income pricing). Its desktop terminal is more than simply a distribution channel for data and is largely used within trading and execution workflows. While LSEG does distribute some non-proprietary data as well, such as financial filings, these are rarely sold in isolation and represent a small fraction of the business overall.

Finally, Equifax is a leading credit bureau. Its data is critical to lenders in evaluating the credit risk of consumers. In addition to credit data on more than 300 million consumers globally, Equifax also owns one of the most differentiated data assets in The Work Number (TWN). TWN maintains a database of employment and income information on 125 million Americans, sourced through a contributory network of three million employers, largely through exclusive agreements. Stubbornly high interest rates have been a headwind to lending activity (particularly in mortgages) in recent times, but in our view this is temporary.

AI MAKES THE DATA MORE VALUABLE

Far from being disrupted by AI, we think there is a good chance that incumbent data assets could become more valuable, for a few reasons.

Firstly, AI can lower the cost of product development (cleansing data, writing code etc.) and accelerate the pace of innovation for data owners. For instance, Equifax's share of revenue from products less than three years old grew to 17% in Q1-26 (from less than 10% five years ago), with all new models and scores now powered by an in-house AI engine (EFX.AI).

Similarly, AI has opened new distribution channels and customer segments for data owners. For example, LSEG is already distributing data through platforms like Snowflake, Databricks and Microsoft, and more than 150 clients are already consuming LSEG data through MCP (Model Context Protocol) servers in frontier AI models like Claude and ChatGPT.

Finally, AI increases the capacity for customer organisations to consume data. Already S&P Global has seen growth 30% higher from AI-forward customers than others and seen clients willing to sign renewals at 35%-40% higher rates in order to access data in an AI-ready format.

CONCLUSION

The debate over which businesses are likely to benefit from AI, and which businesses will be impaired is unlikely to be definitively settled soon. Not all data assets will prove durable, and we avoid those we judge easily replaceable. However, the owners of proprietary, deeply embedded data are, in our view, well placed to compound value through the transition, and over time we expect it to become clear that AI reinforces their moats rather than erodes them. In S&P Global, LSE Group and Equifax we own three such businesses, and we believe patient, long-term shareholders will be rewarded as their quality is recognised.


Insights Disclaimer

Danie Pretorius is an investment analyst with 18 years of investment industry experience.


Related articles

A short note about our global portfolios today

A global markets portfolio that balances real returns and the risk of loss over the long term

How Interactive Brokers, LPL Financial and Charles Schwab are capitalising on emerging trends