WhichAI / AI Debates / Copyright and training data

Copyright and training data

Can models train on the open web without paying? · updated 2026-07-30

Publishers, authors and artists say training on their work without license is mass infringement; labs argue fair use and point to transformation. Courts are deciding case by case, while licensing deals grow in parallel: the market may settle it before the law does.

Dozens of suitsactive US training-data cases, led by The New York Times v. OpenAI and Microsoft
2 reportsUS Copyright Office studies on AI, digital replicas and copyrightability

The concern

The concern: creators' work built these models and they see no compensation; opt-outs came late and are hard to verify; style imitation hits working artists hardest.

The counter-view

The counter-view: training is argued to be transformative like search indexing was; blocking it in one country just moves it elsewhere; and licensing markets (news, stock media, music) are already forming.

Where it stands (July 2026): no final precedent; a mixed pattern of settlements, licensing deals and ongoing suits. Expect years, not months.

Sources: NYT lawsuit coverage (NYT) · US Copyright Office, AI studies