TL;DR
Thorsten Meyer AI’s 2026 Control Series analysis identifies proprietary data as a rising AI chokepoint as public web text approaches limits cited by Epoch AI. The report ties the shift to copyright settlements, licensing deals, expert data demand and national control over battlefield or real-world datasets.
A 2026 Thorsten Meyer AI analysis says the AI industry’s main chokepoint is shifting from rented computing power to proprietary data, citing projections that public web text could be largely used for frontier training between 2026 and 2032 and a licensing market shaped by copyright disputes. The development matters because companies, publishers, experts and governments that control hard-to-copy datasets may gain leverage over AI labs as chips and base models become easier to buy or rent.
The analysis cites Epoch AI’s estimate that the public internet contains about 300 trillion tokens of high-quality text and that the stock of public human text could be fully used for training between 2026 and 2032, with a median around 2028. The finding is a projection, not a measured endpoint. The report also cites Elon Musk’s early-2025 claim that the cumulative sum of human knowledge had been largely exhausted for training.
Thorsten Meyer AI says the scarcity has moved attention away from larger open-web crawls and toward paywalled archives, enterprise records, expert-authored work, autonomous-driving telemetry and battlefield data. Those datasets are not interchangeable, and the source argues that their value rises as compute rental prices fall and base models grow more available.
The report points to Anthropic’s $1.5 billion authors settlement as evidence that courts and markets are narrowing the older practice of collecting training data first and dealing with rights questions later. According to the source, the judge drew a line between training on legally acquired books, described as transformative fair use, and downloading pirated books from shadow libraries. The settlement covered past piracy claims, not future training or model outputs.
Data: The One Thing You Can’t Rent
The free part of “all human knowledge” is running out. As compute and models commoditize, the corpus you can’t replicate becomes the moat — so data is being fenced, priced, and, in places, treated as a national asset.
Data was supposed to be the abundant input. It’s the scarce one. It’s also the chokepoint you can actually own — so guard your proprietary data, and don’t hand it to a provider who can become your competitor (the lesson everyone fled Scale to learn). Nations: license it like Ukraine — keep the model, keep the leverage.
Proprietary Data Gains Leverage
For businesses, the analysis turns data ownership from a compliance issue into a competitive one. A company’s customer records, process logs, internal knowledge base or specialized research may be more defensible than access to a cloud GPU cluster. The report warns that handing unique data to an outside AI provider can weaken the company’s position if that provider later competes in the same market.
Larger AI labs may be better placed to pay for licensed archives, expert feedback and private datasets. That could make the licensing market beneficial for creators while also raising costs for smaller competitors. The report says Anthropic’s settlement and publisher licensing deals show that data once treated as a free input is becoming a priced asset.
The report also frames data as a governance issue for states. It cites Ukraine’s approach to wartime data as an example of sovereign control: access can be licensed, while the model and leverage remain with the country providing the data.
AI training data licensing platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Open Web Scraping Narrowed
For much of the generative AI boom, public web text helped train large models at scale. Thorsten Meyer AI says that phase is ending because the highest-value public text has already been heavily used, while legal pressure has made unlicensed collection riskier.
Part 2 of Thorsten Meyer AI’s series argued that compute is becoming easier to rent, with H100 rental rates down 60% to 75% from peak levels. This third installment says that pricing shift is one reason data now carries more weight: when machines and model techniques spread, the corpus underneath a model becomes a stronger differentiator.
The response to data scarcity has included synthetic data. The source cites Nvidia’s $320 million purchase of synthetic-data company Gretel and Microsoft’s use of hundreds of billions of synthetic tokens in model training. But the report says synthetic data has limits, especially in domains where correct answers are hard to verify and errors can compound across model generations.
“You can rent compute. You can lease power. You cannot rent data that no one else has.”
— Thorsten Meyer AI
expert-authored datasets for AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fair Use Lines Stay Unfinished
It is not yet clear how quickly frontier labs will exhaust usable public text. Epoch AI’s dates are projections, and the timeline could change if labs improve data efficiency, find new sources, or use more synthetic data.
The legal boundary remains unsettled. The report says the New York Times case against OpenAI is still in discovery, and Anthropic’s settlement did not decide future training practices or model-output disputes.
Customer reaction to data deals is another open point. The source cites Meta’s reported $14.3 billion purchase of a 49% stake in Scale AI and says it triggered an exodus, but the longer-term effect on enterprise willingness to share training data is still developing.
public web data archives for AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Licensing Deals Face New Tests
The next tests are likely to come from three places: courts, contracts and governments. Copyright cases will shape how far AI labs can go with training data, while publishers and expert communities will keep pricing access to material that cannot be scraped freely.
Enterprises are expected to tighten data-use terms in AI vendor agreements, especially where proprietary records could train systems later sold to competitors. For national datasets, the report points to licensing and model-control conditions as a path that lets governments share value without giving away strategic leverage.

Splash It!: 99 Customizable Press Release Tools, Texts & Layout Templates (Sovereign Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the actual development in this story?
Thorsten Meyer AI has published a 2026 Control Series analysis arguing that proprietary data is becoming the next major AI chokepoint as public web text becomes less available and more legally contested.
Is public AI training data already gone?
No. The source cites Epoch AI’s projection that public human text could be fully used between 2026 and 2032, with a median around 2028. That is a forecast, not a confirmed endpoint.
Does the Anthropic settlement decide all AI copyright issues?
No. According to the source, the settlement covered past piracy claims and did not resolve future training, model outputs or other pending cases.
Why should companies care about their own data?
The report says proprietary business data can become a competitive asset. If a company gives unique records to an AI provider without tight limits, that data could strengthen a system later used by rivals.
Can synthetic data solve the shortage?
Partly, according to the report. Synthetic data is already used in training, but the source says it carries risks in areas where answers are hard to verify, making fresh human-verified data more valuable.
Source: Thorsten Meyer AI