AI, Copyright, and the Scale Problem
Learning from prior work is not new. Industrialising it is.
There is a growing argument that the use of copyrighted material to train artificial intelligence systems is theft, and that governments need to act before writers, musicians, artists and journalists are damaged beyond repair.
The basic claim is not hard to understand. Creative people own their work. If a company uses that work to build a commercial product, especially one that may later compete with the creator, then the creator should have some control over the transaction. They should be able to consent or refuse. They should be able to negotiate a price. They should not discover, after the fact, that their books, songs, images or articles have been swallowed into a training set and converted into value for a technology company.
There is a policy argument here as well. Copyright law was not written for a world in which vast amounts of creative material could be scraped, copied, analysed and used to train systems capable of producing text, images, music or code on demand. If the existing law protects creators, those protections need to mean something in practice. If the law is no longer adequate, it needs to be updated. That may mean licensing schemes, compensation mechanisms, lawsuits, or some new system that allows creators to be paid when their work contributes to AI training.
I have a lot of sympathy for that position.
Creative work is work. It is not a decorative by-product of society. Writers, illustrators, musicians, photographers and journalists all have to eat. Many already live with precarious incomes, weak bargaining power and platforms that extract more value than they return. It is entirely reasonable for them to look at generative AI and ask: if my work helped make this system valuable, why am I the only person not being paid?
That question deserves a serious answer.
But I am less comfortable with the word “theft” being used as though it settles the matter. It does not. In some cases it may be legally accurate. In others it may be emotionally satisfying. But as a description of the whole problem, it is too blunt. It turns a difficult issue into a slogan.
The difficulty is this: learning from existing work is not new.
Long before AI existed, creative people learned by exposure to the work of others. Writers learned by reading. Musicians learned by listening. Painters learned by copying the masters. Film-makers learned by studying scenes, cuts, lighting and structure. Programmers learned by reading other people’s code, pulling it apart, adapting techniques, and discovering why one solution was elegant and another was a mess.
Nobody creates in a vacuum. We absorb patterns. We imitate before we innovate. We borrow structures, rhythms, techniques and habits of thought. Sometimes we do it consciously. Sometimes we do it without realising. A novelist may carry the influence of half a dozen earlier writers without reproducing a single sentence. A guitarist may learn from years of listening before playing something recognisably their own. A software developer may write code shaped by decades of examples, documentation, libraries and other people’s mistakes.
That is not theft. That is culture.
It is also craft.
So when people say that an AI system has learned from existing work, I do not think the mere fact of learning is the problem. At a high level, there is an analogy with human creative development. Large language models analyse patterns in existing text. Image models analyse relationships between visual features and descriptions. Music models analyse structure, style and probability.
They are not learning as humans learn. They are not human. But they are extracting patterns from prior work and using those patterns to generate new output.
That is close enough to human learning to make the simple theft argument incomplete.
But it is not close enough to make the problem disappear.
The difference is scale.
A person reads hundreds or thousands of books over a lifetime. A model may be trained on millions. A musician spends years listening, practising, forgetting, failing and trying again. A company can process vast catalogues of recorded music as data. A student copies a painting to understand brushwork. A commercial image model may ingest the work of living artists and then produce images close enough to threaten their commissions.
Scale changes the moral character of the act.
It also changes the economics. A human being influenced by another artist still has to do the work. They have to develop taste, judgement, dexterity, discipline and intention. The influence of prior work is filtered through a life.
An AI system has no life. It has no apprenticeship in the human sense. It does not admire another artist. It does not know what it is borrowing. It does not decide that a line is dishonest, or that a phrase is too easy, or that a particular image matters because it carries a memory. It produces outputs from patterns. Often impressive outputs, but patterns nonetheless.
And behind those outputs are companies with capital, infrastructure, lawyers and shareholders.
That is where the analogy with human learning starts to break down. Not because machines cannot analyse prior work, but because the surrounding conditions are completely different. A writer reading another writer is one thing. A corporation copying vast quantities of material into a training pipeline, building a paid product on top of it, and then telling the original creators that nothing of value has been taken is something else.
The problem is not learning from prior art. The problem is industrialising that learning without consent, attribution or compensation.
There is another distinction that matters: influence versus substitution.
If I read a novelist and become a better writer, I have not replaced that novelist in the market. I may have been influenced by them, but I still have to write my own book, find my own subject, develop my own voice, and persuade readers that my work is worth their time. The earlier writer’s work remains intact. Their readers still have a reason to read them.
But if a system is trained on thousands of illustrators and then sold as a cheaper way to avoid hiring illustrators, the situation changes. If journalism is used to train systems that answer questions without sending readers back to the publications that paid for the reporting, the economics change. If books are used to create synthetic competitors, summaries, imitations or study guides that reduce the market for the original work, the grievance is no longer abstract.
That is the test I find most useful. Not “was the system influenced by prior work?” Everything is. The better question is: did the use of that work help create a substitute for the thing itself?
Not every case will have the same answer.
There is a difference between a researcher studying language patterns, a student using AI to understand a difficult chapter, a disabled reader using AI to make text more accessible, and a multinational company building a commercial model from material it did not license. Lumping all of that together under one moral category does not help. Nor does pretending that all of it is harmless innovation.
We need a more adult conversation than that.
A sensible approach would recognise several things at once.
Creators have legitimate rights and legitimate grievances. Their work should not be treated as free industrial feedstock simply because it was accessible online or available in digital form.
Learning from existing work is also part of the normal development of culture and craft. A legal or ethical framework that forgets this will quickly become absurd.
Scale matters. Automation matters. Commercial use matters. Substitution matters. Consent matters. Compensation matters.
And there is probably no single rule that will cover every case cleanly. We may need licensing schemes for some kinds of training, collective compensation mechanisms for others, opt-out or opt-in systems depending on the class of work, and stronger remedies where companies have clearly copied protected material for commercial advantage.
None of that will be simple. But difficulty is not an excuse for doing nothing.
I keep coming back to the distinction between influence and extraction. Influence is how culture breathes. Extraction is how powerful organisations turn other people’s labour into their own asset while insisting that the original contributor has no claim.
AI training sits uncomfortably between those two ideas. Sometimes it looks like learning. Sometimes it looks like copying. Sometimes it looks like a new form of cultural metabolism. Sometimes it looks like the oldest form of corporate behaviour: take what you can, move fast, and let the lawyers argue later.
That is why I do not find either extreme very convincing. “It is all theft” is too simple. “It is just learning” is too convenient.
The hard truth is that both sides are pointing at something real. Creative people do learn from existing work, and always have. Technology companies have also used enormous quantities of human creative labour to build systems from which they expect to make enormous sums of money.
Those two facts have to be held together.
The answer should not be to freeze culture in place, or to make every act of influence legally suspect. Nor should it be to give technology companies a free pass because the machine is impressive and the economics are complicated.
We need to protect the human ecosystem that made these systems possible in the first place. That means recognising the legitimacy of learning from prior work, while refusing to let industrial-scale extraction masquerade as nothing more than a student reading in a library.
That, to me, is where the argument belongs.
Not in pretending that AI invented the idea of learning from others.
And not in pretending that scale changes nothing.


