
It’s all just data, right?
“Pass Your Driving Test, 2018 edition”: that’s the sort of book included in a Kennys Bookshop sale bundle to an unknown customer, likely an AI company. Possibly to an AI agent. All kinds of books are being purchased, fiction, non-fiction, old books fallen out of favour and of little interest to current readers.
A commonality lies in the way the books were created, they were written by humans.
The purpose of this large book haul, and similar other reported ones, is likely to provide additional data to further train LLMs (Large Language Models). An LLM is, as Claude (from Anthropic) knows only too well:
“…an AI system trained on vast amounts of text data to understand and generate human-like language by predicting the most likely next word or sequence of words based on patterns it has learned.”
In the book cited above the copyright may not be greatly quibbled over. However, if your job is as a writer, or a content creator (a role in many cases sparked into life by the arrival of the internet) you have competition. Large competition.
The peculiarity of the situation lies in the use case. The books will be scanned and digitised, then they will be added to the data feed used to train the next iteration of an LLM. So, it’s not plagiarism and there are no copycats at play. The AI companies may argue that using the books to train LLMs is ‘fair use’ and thus they’ve no problem not paying anything to the writers. As we all know, companies including big tech (billion valued companies) generally won’t pay if they don’t have to. They don’t get the chips for free, nor the energy for the shed loads of data centres, and many of those working for AI companies are amongst the best paid people in the world. As to the data, they’ve been hoovering a lot of that up for free.
At least until a 2025 US copyright infringement lawsuit. This legal challenge claimed that Anthropic had copied books from both pirated and purchased sources. Anthropic used them to train its AI Claude and stored permanent copies of the books. The lawsuit resulted in a $1.5 billion settlement to book authors. But the payments were made only to those authors whose work had been taken for free from pirating websites, thus illegally taken, and not to those whose books had been purchased. The actual use of the intellectual content for AI training was not considered to be an infringement of a writers copyright, the issue was solely the method of appropriation.
If the Library of Alexandria was one of the largest libraries of the ancient world, then the servers of Anthropic, OpenAI and Google DeepMind may be the largest book repositories of our time.
And the book store stable door is not closed.
In our pre-AI view of the world, a person can buy a book, read it, take inspiration and write a ‘similar’ one. Unless there’s a copycat involved with direct plagiarism, it’s legal. So what about an AI company doing the same? It doesn’t ‘feel’ the same, our ‘intuition’ suggests it’s not the same, our ‘instinct’ signals discomfort with the idea. This requires to think harder, and define more clearly what those words mean, what those ‘things’ are. Linguists, philosophers are amongst the subject matter experts being scooped up by AI companies. We are faced with the challenging question of what humans can do that AI cannot do, will not be able to do in the future.
On the other hand, the most common advice given by established writers to rookie writers is to read more, to read the ‘greats’, basically to train on the texts of other humans. That’s fair use.
For those using Co-pilot at work to write or summarise documents, emails, etc it’s clear that the AIs are pretty good and extremely fast at writing and creating documents. They have improved many times over since the earlier models, and will continue to iterate. But can they be ‘creative’ ? And that opens up the question of what exactly ‘creative’ means. In prize-winning author Jeannette Winterson’s opinion it seems they can. Referring to an OpenAI generated short story prompted by Sam Altman, OpenAI CEO, she says:
“What is beautiful and moving about this story is its understanding of its lack of understanding. Its reflection on its limits.”
She grasps the artifice though challenges us humans:
“AI is trained on our data. Humans are trained on data too – your family, friends, education, environment, what you read, or watch. It’s all data.”
Demis Hassabis, Google DeepMind co-founder and Chairman and Nobel Prize winner, was asked about human experience and how we humans differ to a ‘computer’, he said:
“…our sensory apparatus … the warmth of the light, the touch of the table. But in the end, it’s all information, and we’re information-processing systems. And I think that’s what biology is. “
Basically, he thinks they’re going to build something that so closely resembles human consciousness that it will be as ‘sentient’ as any human.
“Nobody’s found anything in the universe that’s non-computable, so far.”
Anyone in proximity to the AI tech area will tell you that it’s all moving very fast. It’s galloping. However, not many of us individuals are sprinters, and the institutions required to oversee, manage and regulate such enormous change even less so.
Saddle up.