gor.bio wiki

How AI Training Data Is Sourced (and Why It's Contested)

Where the data behind large AI models actually comes from — web scraping, licensed corpora, and open datasets — and the legal and ethical fight over each source.

Category: Artificial Intelligence · Created: 2026-08-29 · Updated: 2026-08-29

Illustration: BalticServers data center
Illustration: BalticServers data center · Image: BalticServers.com, CC BY-SA 3.0, via Wikimedia Commons.

Every large AI model is a compression of its training data, which makes the question "where does the data come from?" the industry's most consequential question. The answer is three different supply chains, running on three different legal theories, with different futures.

The three supply chains

1. Web scraping. Common Crawl and similar crawls, filtered and deduplicated, powered the current generation of large language models. Legal status: contested everywhere. Fair use is the defense in the U.S.; the EU's DSM Directive and text-and-data-mining exceptions set different conditions; courts are mid-stream. The New York Times v. OpenAI case and the wave of author and publisher suits will define this for a decade.

2. Licensed and proprietary data. Deals with publishers, stock libraries, and platforms (Reddit–Google, News Corp–OpenAI), plus data collected inside a company's own products. This is the industry's pivot away from scrape-first: expensive, clean, and defensible.

3. Open datasets. Deliberately public corpora — Wikipedia, Project Gutenberg, government data, scientific datasets, community-curated sets. Smaller but qualitatively central: heavily represented in high-quality training mixes precisely because licensing is known and provenance is documented.

Why open datasets matter disproportionately

Open datasets carry the metadata the other two chains lack: known license, known source, known collection method. That makes them the research substrate — the benchmark sets, the teaching corpora, the reproducibility layer — and increasingly the compliance asset. A model trained with documented open data can answer "what's in you, and by what right?" A scraped model mostly cannot.

The contested middle

The fight is not "AI vs. no data." It is about consent and compensation: opt-outs that work or don't (robots.txt honored by some, ignored by others); the EU AI Act's training-data transparency requirements, which force disclosure summaries from 2025 onward; and the open question of whether individual creators get anything for their contribution to models that compete with them. Where this lands determines who can build frontier models — only incumbents with licensing budgets, or everyone with an internet connection.

Related reading: large language models for what the data becomes, and public key cryptography for a precedent of open infrastructure beating proprietary lockdown.

Going deeper. Open Data for AGI: Why Public, Free, and Open Datasets Matter by the author of this wiki makes the full engineering and ethical case for open datasets as the foundation of general AI — history, economics, and what a healthy data commons requires. Instant download at the author's bookstore.

Tags

artificial intelligence bookshop datasets training data

Related articles

This text may be freely copied, modified, and reused. See Content Reuse.