Book after scanning. What exactly do we lose?
Is a book still the same book after being scanned? In the age of artificial intelligence, its content can become a valuable data source, but the physical copy carries something that cannot be recorded in a file: materiality, history, provenance, and traces of successive owners.
Book after scanning. What exactly do we lose?
Artificial intelligence may cause us to start buying books en masse not to read them, but to scan them. And then – perhaps – to destroy them.
Sounds absurd?
Until recently, old textbooks from the communist era, outdated travel guides, technical publications, cookbooks, studies on local history, or old manuals could sit on antiquarian bookstore shelves for years. Their prices were often symbolic, and finding a buyer was sometimes harder than finding the book itself.
Today, some of them are starting to have a completely new value.
Not because readers suddenly returned to them.
Their value may come from the fact that they contain something increasingly needed by artificial intelligence systems – data that can be used to build and develop AI models.
![]()
Book digitization station at the University of Liège laboratory. The photo shows an open book placed on a professional book scanner. Author: Sam.Donvil. Source: Wikimedia Commons. License: CC BY-SA 4.0
This phenomenon was recently written about by "Dziennik Gazeta Prawna". Antiquarians in various countries began receiving unusual, bulk orders for books, often very specific and previously unattractive from the perspective of the traditional market. In Poland, particular interest was sparked by the offer of the Canadian company Zoom Books, which approached antiquarians with a proposal for long-term cooperation and informed about the demand for over 800,000 Polish titles.
I am an antiquarian myself and had the opportunity to comment on this topic in "Dziennik Gazeta Prawna". And precisely from the perspective of someone who has dealt with physical books for years, I have a certain problem with this phenomenon.
It's not that I am against digitization.
Quite the opposite.
If thanks to digitization old books can be searched by researchers from the other side of the world, if someone can find a specific quote, information about a historical event, or forgotten technical knowledge within seconds, it is hard to consider this a bad thing.
Access to cultural resources should be as broad as possible.
Digital libraries, archives, and databases can make knowledge that was previously available to few accessible to millions.
The problem begins when we start to perceive the book solely as a collection of information.
Because a book is more than that.
This is no longer just a theory
In mid-July, Marcin Gałązka, owner of the Warsaw Antiquarian Bookstore Zakładka, received a message from the Canadian company Zoom Books.
The offer was unusual.
It was not about a few or a dozen specific books. The company was interested in bulk purchases. The message included information that the list of sought Polish titles included over 800,000 items. The antiquarian was asked to send a catalog with prices and ISBN numbers, preferably in Excel format.
Marcin Gałązka did not decide to cooperate. Instead, he publicized the received offer on his social media, which caused a great stir among internet users and media interest.
His experience is particularly interesting to me because it shows that we are not just talking about a hypothetical future.
Marcin Gałązka also drew my attention to an important context of the whole matter:
"It was not about a few or a dozen specific books. The company was interested in bulk purchases. The message included information that the list of sought Polish titles included over 800,000 items. The antiquarian was asked to send a catalog with prices and ISBN numbers, preferably in Excel format. Gałązka left this offer without a positive response and publicized everything on his social media, which caused a great stir among internet users and media interest."
Marcin Gałązka, Photo by Mikołaj Starzyński
It is worth clarifying one thing here. The mere fact of purchasing or scanning books does not mean that their content will then be made available to internet users. Such an assumption would be too far-reaching.
Much more interesting is the question of why such huge amounts of books are needed at all and what happens to the content obtained from them.
One possible use is using scanned materials as data to develop artificial intelligence systems. In such a case, the user does not get access to a specific scanned book. Its content may be used as one of the elements of the process that results in a system capable of searching, analyzing, processing, or generating information.
And this is what worries me.
Not because digitization itself is bad. On the contrary — it can be a huge opportunity to preserve and share cultural heritage.
The problem begins when a physical book becomes merely a raw material for data extraction, and its material and historical value ceases to matter.
Zoom Books officially presents itself as a company engaged in mass resale and recycling of used books. The company strongly denies allegations that it digitizes or destroys books to train AI models.
Therefore, we should not present this as a proven fact.
But the phenomenon of massive interest in certain groups of books is already a fact. And with it comes a question that, in my opinion, is much more important than who exactly is behind a specific order:
what happens to the book when its most important value becomes the information contained in it?
A book is more than just text
For artificial intelligence, a book may primarily be a source of data. For an antiquarian, a book is not always just text written on paper.
But what exactly is in a book besides text?
When we talk about scanning books, it is very easy to start thinking of them as text located between the first and last page. After all, you just need to scan the pages, recognize the characters, and save them in digital form.
But a book has never been just text.
A book contains illustrations, maps, photographs, engravings, charts, and tables. There is the way the text is arranged on the page, the font type, font size, footnotes, initials, vignettes, decorations. Sometimes it even matters how the author or publisher decided to combine text with illustration.
In old publications, it is especially clear that a book is also an object designed by humans.
Paper has its texture and color. Printing has a certain quality. Pages may be folded by hand. The binding may be made by a specific bookbinder. Someone may have written, underlined, or corrected something in the margins. There may be an ex-libris on the endpaper and a stamp of an old library on the title page.
Even signs of damage can be significant.
A bent corner, worn binding, or a stain on the paper are part of the history of a particular copy. Of course, they can be photographed. They can be described. You can even create a very detailed digital model of the book.
But it will still be a record of information about the object, not the object itself.
And here comes a question that seems particularly important in the whole discussion about digitization:
Did we scan the book, or only everything we could read from it?
That is not the same.
We can preserve the content, illustrations, and appearance of pages, yet lose something that no scanning technology can fully transfer — the material presence of the book and its individual history.
For an algorithm, two copies of the same edition may be almost identical.
For an antiquarian and collector, they can be completely different objects.
One may be clean, unread, and free of any traces of previous owners. The other may have a handwritten dedication from the author, an ex-libris of a famous person, a stamp of a pre-war library, and a binding made by a renowned bookbinder.
![]()
Natural leather binding - protection of the book, but also part of its material history. Photo: Antykwariat Sobieski workshop
The text is the same. The books are not the same.
And perhaps this difference will become increasingly important in a world where the text itself becomes something that can be copied, searched, and analyzed within seconds.
The most worrying scenario
That is why I cannot imagine a situation where, for the purpose of data acquisition, we dismantle a seventeenth-century old print or destroy the first edition of an important literary work.
Such cases seem abstract today.
And I hope they remain abstract.
But they are not entirely fictional.
In the United States, cases of so-called "destructive scanning" have been revealed, where books were mechanically cut to speed up the scanning process and then destroyed. In one of the court cases concerning Anthropic, the issue of such use of legally purchased books became the subject of broad legal and environmental debate.
And here arises a problem we previously did not really have.
If a book is bought as a data carrier, then from an economic point of view its physical form may become an obstacle.
You have to buy the book.
You have to transport it.
You have to find a specific copy.
You have to open it.
You have to scan hundreds of pages.
And if someone wants to scan millions of books, the natural question becomes: how to do it faster?
And here comes destructive scanning.
From an industrial process point of view, cutting the book may be an efficient solution.
From a cultural point of view, it is something completely different.
What if the book is private property?
However, there is one more, much more difficult problem.
Imagine a client comes to an antiquarian and says:
"I buy everything. I pay according to your prices."
He is not interested in one book.
He wants to buy a hundred, five hundred, or a thousand.
And then what?
If the books are his property, what can we actually do?
Can we tell the owner that he is allowed to possess them, but not allowed to do with them what he intends?
In the case of ordinary books, the answer seems relatively simple.
In the case of objects with monument status, the situation is more complicated. The law provides special rules for the protection of certain cultural goods.
And rightly so.
But with the development of technology, it is probably worth asking whether current protection mechanisms always correspond to the actual value of the object.
Because the value of a book does not always come solely from its age or price.
Sometimes it comes from what the book represents.
The first edition of an important work may have great significance for a collector not because it is old, but because it is a material witness to the moment when the work first appeared in circulation.
Similarly with a copy that once belonged to a specific person.
If it bears traces of successive owners on its pages, the book begins to tell a story independent of the text printed in it.
The AI paradox
And here we come to a paradox.
The easier access to the content of books becomes, the more important their physical original may become.
![]()
Władysław Reymont, "Osądzona" - title page of the first edition.
If with a phone we can search millions of scanned books within seconds, the text itself will become an increasingly accessible good.
Information will become common.
But then the importance of what cannot be copied may increase.
The original copy.
Its paper.
The binding.
The dedication.
The ex-libris.
Signs of reading.
The history of successive owners.
Provenance.
In a sense, AI may thus lead to a very interesting reversal of values.
The text will become cheaper. The copy may become more valuable.
Of course, this does not mean that every old book will suddenly become a collectible item.
Quite the opposite.
The mass book market may become even more divided.
On one hand, there will be millions of publications whose content can be found immediately in the digital world.
On the other – there will be books whose value will come precisely from their material uniqueness.
It may be a first edition.
A rare edition.
A unique binding.
A copy with exceptional provenance.
A book that once belonged to a specific person.
A manuscript.
A dedication.
Or simply a beautiful object preserved through generations.
![]()
Marbled endpapers – a decorative element of traditional bookbinding.
Not only an AI issue
There remains the issue of intellectual property rights.
Here the matter is definitely more complicated than the simple question: "Is it allowed to scan a book?".
The question arises about what happens later with the acquired material.
How is it used?
What is it for?
Does it go to a public digital library, a private database, or a commercial system?
Is it used for information retrieval or for training a model?
Where is the boundary between using existing cultural heritage and using it in a way that may affect the rights of its creators?
I am not a lawyer and I would not want to pretend that there are simple answers to all these questions today.
But that is precisely why it is worth discussing them.
Before technology outpaces our reflections.
What can we do right now?
There is, however, something we can do right now.
Read.
And encourage others to read.
Because the role of an antiquarian should not be only to buy and sell books.
Of course, collecting has a financial dimension.
There is nothing wrong with that.
A good book can be an investment. It can have high market value. It can be a sought-after collectible item.
But if we limit a book solely to its price, we lose something much more important.
Our role is also to shape future readers.
Perhaps in a dozen years some of them will become our clients and start collecting first editions, old prints, or beautiful bindings.
But that is a consequence of something much more important:
interest in the book itself.
Because I would not want to live in a world where books survive only as collectible objects, valued by auctioneers and collectors, while fewer and fewer people simply read them.
Perhaps the greatest paradox of the AI era will be that the machine learns to read books faster than humans, and our task will be to remind people why it is worth reading them at all.
Because a book is not just information.
It is also an object, memory, history, the hands of successive owners, the trace of time, and part of our culture.
And that cannot be reduced to a file.
Author: Paweł Kołata, Antykwariat Sobieski
sources:
-
Dziennik Gazeta Prawna” - “Scan and grind. How AI feeds on old books” https://edgp.gazetaprawna.pl/magazyn/artykuly/11289740%2Czeskanuj-i-zmiel-jak-ai-zywi-sie-starymi-ksiazkami.html?utm_source=chatgpt.com
-
Wirtualne Media - “Foreign companies buy Polish books. They may be used to train AI” https://www.wirtualnemedia.pl/zagraniczne-firmy-skupuja-polskie-ksiazki-moga-posluzyc-do-trenowania-ai%2C7315847212050496a?utm_source=chatgpt.com
-
Publishers Lunch - Zoom Books position https://lunch.publishersmarketplace.com/2026/05/zoom-books-responds-to-accusations-of-selling-books-for-ai-scraping/?utm_source=chatgpt.com
-
The Guardian - “More than just objects” https://www.theguardian.com/technology/2026/aug/02/australian-book-sellers-alarm-destruction-rare-titles-ai-supply-chain?utm_source=chatgpt.com
-
Fortune – investigation into destructive scanning https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phishing-3000-book-copies-scan-destroy/?gsid=5470a5df-39a5-4396-938c-51f00efbb95c&utm_source=chatgpt.com
Polish