AI training using second-hand books

AI training using second-hand books: Why AI firms are buying up second-hand bookshops and the books are being shredded

AI training using second-hand books: Why AI firms are buying up second-hand bookshops and the books are being shredded

ein krieger mit axt steht in einem altertümlichen raum und schaut auf ein buch

DESCRIPTION: AI companies are buying second-hand books as training data, scanning them and destroying the copies. What cultural heritage is being lost in the process, what the court case against Anthropic reveals, and why Walter Benjamin anticipated the nature of the problem.

They’re cutting up the originals: AI firms are buying up second-hand bookshops

AI companies are buying up stock from second-hand bookshops, cutting off the spines, scanning the pages and disposing of the copies. This affects cookery books, travel guides, out-of-date specialist texts and regional literature. The article asks what disappears along with the physical copy, why Walter Benjamin anticipated the structure of this process, and what effect it has on us when knowledge is delivered without a visible source.

How AI firms are buying up second-hand bookshops for AI training

It begins at night. At 2.53 am on 30 April 2026, the Bielefeld-based bookseller Michael Ströter received his first order, then the next, then the next. It was always the same company: Zoom Books, a Canadian platform that describes itself as North America’s leading book recycling firm. Ströter, now retired, sells second-hand books as a hobby, usually no more than three a day. He cancelled the orders.

The *taz* investigated the case in June 2026; SWR spoke to the Tübingen-based antiquarian bookseller Roger Sonnewald; *Börsenblatt* and BR followed suit. Antiquarian booksellers from all over Germany are discussing the matter in trade and online forums; reports of the same large-scale orders are coming in from Bulgaria, the UK, Germany and New Zealand. Initially, deliveries were to be made to addresses in the US. Meanwhile, goods from German dealers are being sent to a warehouse in Kodersdorf, Saxony.

The process is industrial in scale and largely automated. The volumes are purchased, the bindings removed, the pages trimmed, the loose blocks fed through sheet-fed scanners, and the remains shredded or recycled. In an interview with SWR, Sonnewald made the crucial distinction: this is not theft; the works are paid for. Nevertheless, cultural heritage is being lost because such large quantities are involved.

Zoom Books itself denies that the process is for AI training. A manager at the company told the *taz* that they digitise the books but do not destroy them. They do not disclose any information about the buyers of the purchased books.

Anthropic, the Washington Post andder Gerichtsakt the court file

For the AI firms themselves, the practice is well documented. A court document from the US proceedings against Anthropic describes precisely this process: the company acquired printed works, removed the bindings, trimmed the pages, scanned them and saved them as searchable PDF files. The Washington Post described the process to in early 2026 under the headline ‘Inside an AI start-up’s plan to scan and dispose of millions of books’.

The background to this is a legal dispute over training data. Anthropic had previously downloaded millions of books from the internet, from so-called ‘shadow libraries’ – that is, illegal collections hosted on third-party servers. US authors filed a class-action lawsuit against this. The proceedings ended in a settlement: the company agreed to pay at least 1.5 billion US dollars – according to media reports, around 3,000 dollars per book.

Buying from second-hand bookshops is the response to this ruling. Items acquired legally can be defended differently from those originating from a ‘shadow library’.

From slow-moving stock to rare single copies: the criteria for purchasing

The search is on for second-hand books with ISBN numbers dating from the 1970s onwards: old non-fiction, specialist texts, novels, cookery books, travel literature, regional history, out-of-print titles and slow-moving stock. Often, this involves double-digit quantities of inexpensive volumes, where the postage costs exceed the value of the goods. The ISBN is no coincidence here, and its role extends further than it might initially appear. According to research by 404 Media, staff at the Amazon warehouse begin by scanning the barcode and ISBN of each volume. This suggests a process that works through the ISBN list rather than searching by subject: a comprehensive scan of everything published with a number since the 1970s.

The reason for this selection lies in aErschöpfung. shortage. The freely accessible texts on the internet have largely been exhausted as training material for AI models, and they are unsorted and of mixed quality. What is missing,ist Sprache are linguistic documents , that have been available  never before in digital formvorlag. The SRF report lists the areas being sought: regional history, linguistics, law and economics – in other words, texts featuring historical language varieties and stylistic peculiarities that are absent from today’s internet.

This also dispels the comforting assumption that it is merely a matter of disposable goods. On 18 August 2026, 404 Media published a report that frames the issue differently: in collaboration with a bookseller, the journalists placed an AirTag inside a rare book that was dispatched as part of a bulk order. The broadcaster reported from an Amazon warehouse in Las Vegas, where, according to the report, books are taken apart and scanned. Amazon did not comment on the investigation, but stated that it purchases books through commercial channels to improve its own products and services.

This reveals two things. Firstly, one of the buyers – about whom Zoom Books refused to provide any information – is now known. Secondly, rarityein Exemplar  apparently does not protect a book from destruction. The value that the antiquarian book trade attributes to a book and the value that a training dataset attributes to it have nothing to do with one another: what matters to the model is that the text is new, not that the copy is rare. A volume of which only twelve copies remain worldwide is worth just as much under this calculation as one of twelve thousand.

The slow-moving stock becomes a sought-after commodity simply because no one has ever uploaded it. Its value lies in its invisibility. For the trade, this turns an old equation on its head: what previously ended up as waste paper suddenly finds a buyer, precisely because nobody wants to read it. The warehouse’s logo reveals how it sees itself: a Tyrannosaurus rex about to tear a book to pieces. According to 404 Media, internal discussions amongst staff indicate that the site was at one point in danger of running out of stock, and the workforce was bracing for closure. New stock arrived.

Ströter describes this with mixed feelings: he is getting rid of his slow-moving stock, most recently a book from the 1970s on television education in children’s homes. Such material is no longer relevant today, nor does it reflect the current state of scientific knowledge. He would like to see a debate on what material is actually used to train artificial intelligence.

Fair Use under US law: Why the books are being destroyed

Why is the copy subsequently destroyed? The answer lies in US copyright law. Anyone who copies digital texts online risks claims for damages. The practice of physically purchasing a book and then destroying it after it has been scanned is regarded within the industry as an attempt to invoke the fair use principle. This allows the use of copyright-protected material without express permission, provided it serves the purposes of education and the stimulation of intellectual production.

Whether this approach holds water remains to be seen. Anyone who buys a book acquires that particular copy. The exploitation rights to the text remain with the author; copyright is not extinguished by the purchase. Until the courts have ruled on the matter, the process remains a grey area.

Destruction is not a side effect of digitisation. It is its legal purpose. A copy that no longer exists cannot be cited as an unauthorised reproduction. The shredder is part of the legal framework.

Loss of cultural heritage: what disappears during scanning

It is tempting to dramatise the situation, but this is misleading. A second Library of Alexandria is not burning here. Knowledge does not disappear; the text is preserved. However, it is preserved in a form that is detached from its physical medium, its authors, its readers and its history.

The book is bought without being read. It is bought as raw material.

With the physical copy, things may disappear that the scan does not capture: different editions, dedications, marginal notes, signs of ownership, library stamps, dog-ears at the very spots where someone once paused. And the simple possibility that someone might open this book again in twenty years’ time.

Benjamin and the Tinned Food Can

Walter Benjamin described the structure of this problem long before the advent of language models. In *A Short History of Photography*, he criticises technical images that imbue objects with a suggestive, seemingly self-evident meaning, whilst the contexts in which these objects exist remain invisible. Photography, he argues, can turn a tin can into something cosmic whilst concealing the human relationships attached to it.

This can be translated into a linguistic model. It provides linguistically plausible, often impressively formulated answers. Yet it fails to show from which books, works, conflicts, legal relationships and material traditions its linguistic competence has emerged. Knowledge appears to be readily available, anonymous and devoid of context.

The shredded book is the perfect illustration of this. The text lives on as statistical material; the physical copy, as an object, disappears.

It is not about the aura of uniqueness

Benjamin is ill-suited as a key witness for a glorification of the book. In *The Work of Art in the Age of Mechanical Reproduction*, he describes the withering away of the aura without lamenting it. For him, reproduction could open up new forms of collective experience and political appropriation. Anyone who casts him as the guardian of sacred uniqueness misreads him.

His question is different. It is: Who controls the technology, what is done with the cultural material, and which social relationships are made visible or invisible in the process?

That is precisely what determines the issue. Not the fact that scanning takes place, but the fact that, in the end, a database ends up in private hands, the origin of which has been obscured in the very process by which it was exploited.

What is lost with every destroyed copy

The psychologically interesting aspect concerns the dedication. A second-hand book bears traces of someone who was there before us. The handwriting on the flyleaf, the date, the name, the underlined paragraph: these are evidence that another person held this text in their hands and found it important enough to give it away or mark it up.

Donald Winnicott coined the term ‘transitional object’ for such things: the teddy bear, the blanket, the well-worn handkerchief maintain a connection to someone who is not there at the moment. Winnicott’s point was that these objects belong neither entirely to the external world nor entirely to the internal world; they occupy a third realm, in which culture first comes into being.

A second-hand book occupies the same realm. The text belongs to everyone; the copy was once owned by someone; and the underlining on page 112 is the trace of another person’s attention. When you buy from a second-hand bookshop, you buy both.

The scan captures the text. The traces of previous users end up in the bin.

The antiquarian bookseller, the antiquarian book market and the holding period

Sonnewald made a comment on SWR that is easy to overlook: some antiquarian books need a certain amount of time on the shelves before they become sought-after. What is a slow-seller today may be a source of value in thirty years’ time.

This statement describes a value that no sales statistic captures. Which items will one day be needed is determined by questions that nobody asks at the time of purchase. Instruction manuals from the 1970s became important for the history of technology, travel guides for tourism research, and cookery books for the history of nutrition. None of these disciplines would have given a reason to keep the volumes when they were first published.

Exactly the same thing is happening again, only in reverse. The books have become sought-after because a technology required them that did not yet exist when they were printed. The difference from any previous demand is that this time, the stock is destroyed upon use.

Google Books and the shredder: the difference in digitisation

The objection is obvious: books have been scanned on a massive scale before. Since 2004, Google Books has digitised millions of volumes, mostly from libraries, and has been embroiled in years of legal disputes over this.

Two differences are key. Firstly, the books remained in the libraries at the time. They were opened, photographed and returned to the shelves. Secondly, the aim was to create a searchable catalogue that linked back to the physical copy: anyone who saw a reference would be told which book it was in and could order it.

The current process reverses both of these. The physical copy disappears, and the end result is a model whose answers no longer cite the source. The 2004 scan made books discoverable. The 2026 scan transforms books into empty spaces.

Knowledge without a visible origin

The impact on the user is the real question. A model responds in smooth, coherent language. It does not cite a source unsolicited, it does not show its working, and it rarely admits where it has obtained information from.

This form of information resembles less a reference work than a statement from an authority that owes no explanation. Anyone who follows this blog will be familiar with the argument from the post on AI as a secular oracle: the empty throne rarely remains empty. Language without a traceable origin invites us to ascribe to it an authority that no one has granted.

The process in second-hand bookshops reveals what this language is made of. In the process, the origin is physically removed.

What this process does to reading

The sense of offence that many feel at this news has a clear reason. Anyone who reads assumes there is someone at the other end. A sentence is the trace of an effort: someone has wrestled with this phrasing, discarded something, and settled on a word. This assumption underpins reading, even where the author has long been dead.

A language model produces sentences without anyone having struggled to craft them. The effort lies in the training data – that is, in those who wrote the text, but whose names are not provided. What emerges takes the form of personal speech and has no author.

For the reader, this is an unfamiliar situation. One can take issue with a text because one assumes that a particular viewpoint lies behind it. In the case of a response without an author, there is no one to whom the objection applies. What remains is the choice between acceptance and rejection, and that is a poorer experience than any dialogue with a book.

The shredder illustrates this situation vividly. It removes precisely what would have indicated that someone had been here before us.

What follows from this

Digitisation preserves texts that would otherwise have disintegrated and makes them accessible to people who would never set foot in a second-hand bookshop in Tübingen. None of this implies a ban on scanning.

The relevant question is: who is allowed to know what a model consists of? There are already two answers to this, and they are worlds apart.

The EU AI Act requires providers to be transparent about their training data. However, they are not obliged to disclose this in detail; according to the wording, a sufficiently detailed summary— —is sufficient. The courts will have to clarify what this means. Until then, the obligation to provide information remains a mere formality.

The Swiss Apertus model demonstrates that things can be done differently. Its developers have disclosed the training data. Anyone wishing to know which texts are included in this model can look them up. With commercial providers, no one – not even a court – can do so without legal proceedings.

Just how reliable this information is can be seen from the way the Amazon warehouse came to light. It took a bookseller willing to place a tracking device in a parcel, and journalists who tracked the signal. At present, the public is finding out what is contained in a model via AirTags.

This shifts the focus from cultural heritage to availability. Whether a cookery book from 1978 was worth shredding is open to debate. What nobody can verify after the destruction is the correctness of the decision, once the cookery book has been destroyed.

The Library of Alexandria went up in flames. The library of the entire world is now being dismantled, scanned, destroyed and converted into proprietary AI models.

Key points in brief

  • Since 2026, AI companies have been buying up book collections in German-speaking countries for AI training; booksellers report large orders from the Canadian-American company Zoom Books.

  • The volumes are cut open, scanned and then disposed of or recycled.

  • We are looking for books with ISBNs dating from the 1970s onwards: specialist books, cookery books, travel literature and regional history. Their value lies in the fact that they have never been digitised.

  • Rarity offers no protection. In August 2026, 404 Media tracked a rare book via AirTag to an Amazon warehouse in Las Vegas, where books are dismantled and scanned. Staff there record the barcode and ISBN, suggesting that the ISBN list is being systematically worked through.

  • The destruction is part of an attempt to invoke the US ‘fair use’ principle. The shredder is part of the legal framework.

  • The text is preserved as statistical material. Lost, however, are different editions, dedications, marginal notes, signs of ownership and the possibility of reading the book at a later date.

  • Walter Benjamin described this structure: technology can imbue things with meaning whilst simultaneously rendering the relationships in which they exist invisible.

  • Benjamin was not defending any sacred aura. His concern was with the control of technology and the visibility of social relations.

  • Second-hand books bear the traces of previous readers: dedications, underlinings, ownership marks. The scan captures the text and leaves these traces in the container.

  • Google Books left the books in the libraries and provided links to their locations. The current process eliminates the physical copy and provides answers without citing sources.

  • Reading presupposes someone at the other end. A digital model provides the form of personal discourse without a sender whom one might contradict.

  • A court document from the case against Anthropic confirms the process: bindings removed, pages trimmed, scanned, and saved as searchable PDFs. The settlement with US authors cost at least 1.5 billion dollars.

  • The German shipments are now being sent to a warehouse in Kodersdorf, Saxony. Zoom Books denies digitising and destroying books, and does not name any recipients.

  • The crucial questions concern distribution: public or restricted access, destruction as a legal consequence, notification of authors, and the authority to decide what is dispensable.

Sources


Related

DESCRIPTION: AI companies are buying second-hand books as training data, scanning them and destroying the copies. What cultural heritage is being lost in the process, what the court case against Anthropic reveals, and why Walter Benjamin anticipated the nature of the problem.

They’re cutting up the originals: AI firms are buying up second-hand bookshops

AI companies are buying up stock from second-hand bookshops, cutting off the spines, scanning the pages and disposing of the copies. This affects cookery books, travel guides, out-of-date specialist texts and regional literature. The article asks what disappears along with the physical copy, why Walter Benjamin anticipated the structure of this process, and what effect it has on us when knowledge is delivered without a visible source.

How AI firms are buying up second-hand bookshops for AI training

It begins at night. At 2.53 am on 30 April 2026, the Bielefeld-based bookseller Michael Ströter received his first order, then the next, then the next. It was always the same company: Zoom Books, a Canadian platform that describes itself as North America’s leading book recycling firm. Ströter, now retired, sells second-hand books as a hobby, usually no more than three a day. He cancelled the orders.

The *taz* investigated the case in June 2026; SWR spoke to the Tübingen-based antiquarian bookseller Roger Sonnewald; *Börsenblatt* and BR followed suit. Antiquarian booksellers from all over Germany are discussing the matter in trade and online forums; reports of the same large-scale orders are coming in from Bulgaria, the UK, Germany and New Zealand. Initially, deliveries were to be made to addresses in the US. Meanwhile, goods from German dealers are being sent to a warehouse in Kodersdorf, Saxony.

The process is industrial in scale and largely automated. The volumes are purchased, the bindings removed, the pages trimmed, the loose blocks fed through sheet-fed scanners, and the remains shredded or recycled. In an interview with SWR, Sonnewald made the crucial distinction: this is not theft; the works are paid for. Nevertheless, cultural heritage is being lost because such large quantities are involved.

Zoom Books itself denies that the process is for AI training. A manager at the company told the *taz* that they digitise the books but do not destroy them. They do not disclose any information about the buyers of the purchased books.

Anthropic, the Washington Post andder Gerichtsakt the court file

For the AI firms themselves, the practice is well documented. A court document from the US proceedings against Anthropic describes precisely this process: the company acquired printed works, removed the bindings, trimmed the pages, scanned them and saved them as searchable PDF files. The Washington Post described the process to in early 2026 under the headline ‘Inside an AI start-up’s plan to scan and dispose of millions of books’.

The background to this is a legal dispute over training data. Anthropic had previously downloaded millions of books from the internet, from so-called ‘shadow libraries’ – that is, illegal collections hosted on third-party servers. US authors filed a class-action lawsuit against this. The proceedings ended in a settlement: the company agreed to pay at least 1.5 billion US dollars – according to media reports, around 3,000 dollars per book.

Buying from second-hand bookshops is the response to this ruling. Items acquired legally can be defended differently from those originating from a ‘shadow library’.

From slow-moving stock to rare single copies: the criteria for purchasing

The search is on for second-hand books with ISBN numbers dating from the 1970s onwards: old non-fiction, specialist texts, novels, cookery books, travel literature, regional history, out-of-print titles and slow-moving stock. Often, this involves double-digit quantities of inexpensive volumes, where the postage costs exceed the value of the goods. The ISBN is no coincidence here, and its role extends further than it might initially appear. According to research by 404 Media, staff at the Amazon warehouse begin by scanning the barcode and ISBN of each volume. This suggests a process that works through the ISBN list rather than searching by subject: a comprehensive scan of everything published with a number since the 1970s.

The reason for this selection lies in aErschöpfung. shortage. The freely accessible texts on the internet have largely been exhausted as training material for AI models, and they are unsorted and of mixed quality. What is missing,ist Sprache are linguistic documents , that have been available  never before in digital formvorlag. The SRF report lists the areas being sought: regional history, linguistics, law and economics – in other words, texts featuring historical language varieties and stylistic peculiarities that are absent from today’s internet.

This also dispels the comforting assumption that it is merely a matter of disposable goods. On 18 August 2026, 404 Media published a report that frames the issue differently: in collaboration with a bookseller, the journalists placed an AirTag inside a rare book that was dispatched as part of a bulk order. The broadcaster reported from an Amazon warehouse in Las Vegas, where, according to the report, books are taken apart and scanned. Amazon did not comment on the investigation, but stated that it purchases books through commercial channels to improve its own products and services.

This reveals two things. Firstly, one of the buyers – about whom Zoom Books refused to provide any information – is now known. Secondly, rarityein Exemplar  apparently does not protect a book from destruction. The value that the antiquarian book trade attributes to a book and the value that a training dataset attributes to it have nothing to do with one another: what matters to the model is that the text is new, not that the copy is rare. A volume of which only twelve copies remain worldwide is worth just as much under this calculation as one of twelve thousand.

The slow-moving stock becomes a sought-after commodity simply because no one has ever uploaded it. Its value lies in its invisibility. For the trade, this turns an old equation on its head: what previously ended up as waste paper suddenly finds a buyer, precisely because nobody wants to read it. The warehouse’s logo reveals how it sees itself: a Tyrannosaurus rex about to tear a book to pieces. According to 404 Media, internal discussions amongst staff indicate that the site was at one point in danger of running out of stock, and the workforce was bracing for closure. New stock arrived.

Ströter describes this with mixed feelings: he is getting rid of his slow-moving stock, most recently a book from the 1970s on television education in children’s homes. Such material is no longer relevant today, nor does it reflect the current state of scientific knowledge. He would like to see a debate on what material is actually used to train artificial intelligence.

Fair Use under US law: Why the books are being destroyed

Why is the copy subsequently destroyed? The answer lies in US copyright law. Anyone who copies digital texts online risks claims for damages. The practice of physically purchasing a book and then destroying it after it has been scanned is regarded within the industry as an attempt to invoke the fair use principle. This allows the use of copyright-protected material without express permission, provided it serves the purposes of education and the stimulation of intellectual production.

Whether this approach holds water remains to be seen. Anyone who buys a book acquires that particular copy. The exploitation rights to the text remain with the author; copyright is not extinguished by the purchase. Until the courts have ruled on the matter, the process remains a grey area.

Destruction is not a side effect of digitisation. It is its legal purpose. A copy that no longer exists cannot be cited as an unauthorised reproduction. The shredder is part of the legal framework.

Loss of cultural heritage: what disappears during scanning

It is tempting to dramatise the situation, but this is misleading. A second Library of Alexandria is not burning here. Knowledge does not disappear; the text is preserved. However, it is preserved in a form that is detached from its physical medium, its authors, its readers and its history.

The book is bought without being read. It is bought as raw material.

With the physical copy, things may disappear that the scan does not capture: different editions, dedications, marginal notes, signs of ownership, library stamps, dog-ears at the very spots where someone once paused. And the simple possibility that someone might open this book again in twenty years’ time.

Benjamin and the Tinned Food Can

Walter Benjamin described the structure of this problem long before the advent of language models. In *A Short History of Photography*, he criticises technical images that imbue objects with a suggestive, seemingly self-evident meaning, whilst the contexts in which these objects exist remain invisible. Photography, he argues, can turn a tin can into something cosmic whilst concealing the human relationships attached to it.

This can be translated into a linguistic model. It provides linguistically plausible, often impressively formulated answers. Yet it fails to show from which books, works, conflicts, legal relationships and material traditions its linguistic competence has emerged. Knowledge appears to be readily available, anonymous and devoid of context.

The shredded book is the perfect illustration of this. The text lives on as statistical material; the physical copy, as an object, disappears.

It is not about the aura of uniqueness

Benjamin is ill-suited as a key witness for a glorification of the book. In *The Work of Art in the Age of Mechanical Reproduction*, he describes the withering away of the aura without lamenting it. For him, reproduction could open up new forms of collective experience and political appropriation. Anyone who casts him as the guardian of sacred uniqueness misreads him.

His question is different. It is: Who controls the technology, what is done with the cultural material, and which social relationships are made visible or invisible in the process?

That is precisely what determines the issue. Not the fact that scanning takes place, but the fact that, in the end, a database ends up in private hands, the origin of which has been obscured in the very process by which it was exploited.

What is lost with every destroyed copy

The psychologically interesting aspect concerns the dedication. A second-hand book bears traces of someone who was there before us. The handwriting on the flyleaf, the date, the name, the underlined paragraph: these are evidence that another person held this text in their hands and found it important enough to give it away or mark it up.

Donald Winnicott coined the term ‘transitional object’ for such things: the teddy bear, the blanket, the well-worn handkerchief maintain a connection to someone who is not there at the moment. Winnicott’s point was that these objects belong neither entirely to the external world nor entirely to the internal world; they occupy a third realm, in which culture first comes into being.

A second-hand book occupies the same realm. The text belongs to everyone; the copy was once owned by someone; and the underlining on page 112 is the trace of another person’s attention. When you buy from a second-hand bookshop, you buy both.

The scan captures the text. The traces of previous users end up in the bin.

The antiquarian bookseller, the antiquarian book market and the holding period

Sonnewald made a comment on SWR that is easy to overlook: some antiquarian books need a certain amount of time on the shelves before they become sought-after. What is a slow-seller today may be a source of value in thirty years’ time.

This statement describes a value that no sales statistic captures. Which items will one day be needed is determined by questions that nobody asks at the time of purchase. Instruction manuals from the 1970s became important for the history of technology, travel guides for tourism research, and cookery books for the history of nutrition. None of these disciplines would have given a reason to keep the volumes when they were first published.

Exactly the same thing is happening again, only in reverse. The books have become sought-after because a technology required them that did not yet exist when they were printed. The difference from any previous demand is that this time, the stock is destroyed upon use.

Google Books and the shredder: the difference in digitisation

The objection is obvious: books have been scanned on a massive scale before. Since 2004, Google Books has digitised millions of volumes, mostly from libraries, and has been embroiled in years of legal disputes over this.

Two differences are key. Firstly, the books remained in the libraries at the time. They were opened, photographed and returned to the shelves. Secondly, the aim was to create a searchable catalogue that linked back to the physical copy: anyone who saw a reference would be told which book it was in and could order it.

The current process reverses both of these. The physical copy disappears, and the end result is a model whose answers no longer cite the source. The 2004 scan made books discoverable. The 2026 scan transforms books into empty spaces.

Knowledge without a visible origin

The impact on the user is the real question. A model responds in smooth, coherent language. It does not cite a source unsolicited, it does not show its working, and it rarely admits where it has obtained information from.

This form of information resembles less a reference work than a statement from an authority that owes no explanation. Anyone who follows this blog will be familiar with the argument from the post on AI as a secular oracle: the empty throne rarely remains empty. Language without a traceable origin invites us to ascribe to it an authority that no one has granted.

The process in second-hand bookshops reveals what this language is made of. In the process, the origin is physically removed.

What this process does to reading

The sense of offence that many feel at this news has a clear reason. Anyone who reads assumes there is someone at the other end. A sentence is the trace of an effort: someone has wrestled with this phrasing, discarded something, and settled on a word. This assumption underpins reading, even where the author has long been dead.

A language model produces sentences without anyone having struggled to craft them. The effort lies in the training data – that is, in those who wrote the text, but whose names are not provided. What emerges takes the form of personal speech and has no author.

For the reader, this is an unfamiliar situation. One can take issue with a text because one assumes that a particular viewpoint lies behind it. In the case of a response without an author, there is no one to whom the objection applies. What remains is the choice between acceptance and rejection, and that is a poorer experience than any dialogue with a book.

The shredder illustrates this situation vividly. It removes precisely what would have indicated that someone had been here before us.

What follows from this

Digitisation preserves texts that would otherwise have disintegrated and makes them accessible to people who would never set foot in a second-hand bookshop in Tübingen. None of this implies a ban on scanning.

The relevant question is: who is allowed to know what a model consists of? There are already two answers to this, and they are worlds apart.

The EU AI Act requires providers to be transparent about their training data. However, they are not obliged to disclose this in detail; according to the wording, a sufficiently detailed summary— —is sufficient. The courts will have to clarify what this means. Until then, the obligation to provide information remains a mere formality.

The Swiss Apertus model demonstrates that things can be done differently. Its developers have disclosed the training data. Anyone wishing to know which texts are included in this model can look them up. With commercial providers, no one – not even a court – can do so without legal proceedings.

Just how reliable this information is can be seen from the way the Amazon warehouse came to light. It took a bookseller willing to place a tracking device in a parcel, and journalists who tracked the signal. At present, the public is finding out what is contained in a model via AirTags.

This shifts the focus from cultural heritage to availability. Whether a cookery book from 1978 was worth shredding is open to debate. What nobody can verify after the destruction is the correctness of the decision, once the cookery book has been destroyed.

The Library of Alexandria went up in flames. The library of the entire world is now being dismantled, scanned, destroyed and converted into proprietary AI models.

Key points in brief

  • Since 2026, AI companies have been buying up book collections in German-speaking countries for AI training; booksellers report large orders from the Canadian-American company Zoom Books.

  • The volumes are cut open, scanned and then disposed of or recycled.

  • We are looking for books with ISBNs dating from the 1970s onwards: specialist books, cookery books, travel literature and regional history. Their value lies in the fact that they have never been digitised.

  • Rarity offers no protection. In August 2026, 404 Media tracked a rare book via AirTag to an Amazon warehouse in Las Vegas, where books are dismantled and scanned. Staff there record the barcode and ISBN, suggesting that the ISBN list is being systematically worked through.

  • The destruction is part of an attempt to invoke the US ‘fair use’ principle. The shredder is part of the legal framework.

  • The text is preserved as statistical material. Lost, however, are different editions, dedications, marginal notes, signs of ownership and the possibility of reading the book at a later date.

  • Walter Benjamin described this structure: technology can imbue things with meaning whilst simultaneously rendering the relationships in which they exist invisible.

  • Benjamin was not defending any sacred aura. His concern was with the control of technology and the visibility of social relations.

  • Second-hand books bear the traces of previous readers: dedications, underlinings, ownership marks. The scan captures the text and leaves these traces in the container.

  • Google Books left the books in the libraries and provided links to their locations. The current process eliminates the physical copy and provides answers without citing sources.

  • Reading presupposes someone at the other end. A digital model provides the form of personal discourse without a sender whom one might contradict.

  • A court document from the case against Anthropic confirms the process: bindings removed, pages trimmed, scanned, and saved as searchable PDFs. The settlement with US authors cost at least 1.5 billion dollars.

  • The German shipments are now being sent to a warehouse in Kodersdorf, Saxony. Zoom Books denies digitising and destroying books, and does not name any recipients.

  • The crucial questions concern distribution: public or restricted access, destruction as a legal consequence, notification of authors, and the authority to decide what is dispensable.

Sources


Related

Anfahrt & Öffnungszeiten

Close-up portrait of dr. stemper
Close-up portrait of a dog

Psychologie Berlin

c./o. AVATARAS Institut

Kalckreuthstr. 16 – 10777 Berlin

virtuelles Festnetz: +49 30 26323366

E-Mail: info@praxis-psychologie-berlin.de

Montag

11:00-19:00

Dienstag

11:00-19:00

Mittwoch

11:00-19:00

Donnerstag

11:00-19:00

Freitag

11:00-19:00

a colorful map, drawing

Google Maps-Karte laden:

Durch Klicken auf diesen Schutzschirm stimmen Sie dem Laden der Google Maps-Karte zu. Dabei werden Daten an Google übertragen und Cookies gesetzt. Google kann diese Informationen zur Personalisierung von Inhalten und Werbung nutzen.

Weitere Informationen finden Sie in unserer Datenschutzerklärung und in der Datenschutzerklärung von Google.

Klicken Sie hier, um die Karte zu laden und Ihre Zustimmung zu erteilen.

Dr. Stemper

©

2026

Dr. Dirk Stemper

Samstag, 22.8.2026

a green flower
an orange flower
a blue flower

Anfahrt & Öffnungszeiten

Close-up portrait of dr. stemper
Close-up portrait of a dog

Psychologie Berlin

c./o. AVATARAS Institut

Kalckreuthstr. 16 – 10777 Berlin

virtuelles Festnetz: +49 30 26323366

E-Mail: info@praxis-psychologie-berlin.de

Montag

11:00-19:00

Dienstag

11:00-19:00

Mittwoch

11:00-19:00

Donnerstag

11:00-19:00

Freitag

11:00-19:00

a colorful map, drawing

Google Maps-Karte laden:

Durch Klicken auf diesen Schutzschirm stimmen Sie dem Laden der Google Maps-Karte zu. Dabei werden Daten an Google übertragen und Cookies gesetzt. Google kann diese Informationen zur Personalisierung von Inhalten und Werbung nutzen.

Weitere Informationen finden Sie in unserer Datenschutzerklärung und in der Datenschutzerklärung von Google.

Klicken Sie hier, um die Karte zu laden und Ihre Zustimmung zu erteilen.

Dr. Stemper

©

2026

Dr. Dirk Stemper

Samstag, 22.8.2026

Webdesign & - Konzeption:

a green flower
an orange flower
a blue flower

Anfahrt & Öffnungszeiten

Close-up portrait of dr. stemper
Close-up portrait of a dog

Psychologie Berlin

c./o. AVATARAS Institut

Kalckreuthstr. 16 – 10777 Berlin

virtuelles Festnetz: +49 30 26323366

E-Mail: info@praxis-psychologie-berlin.de

Montag

11:00-19:00

Dienstag

11:00-19:00

Mittwoch

11:00-19:00

Donnerstag

11:00-19:00

Freitag

11:00-19:00

a colorful map, drawing

Google Maps-Karte laden:

Durch Klicken auf diesen Schutzschirm stimmen Sie dem Laden der Google Maps-Karte zu. Dabei werden Daten an Google übertragen und Cookies gesetzt. Google kann diese Informationen zur Personalisierung von Inhalten und Werbung nutzen.

Weitere Informationen finden Sie in unserer Datenschutzerklärung und in der Datenschutzerklärung von Google.

Klicken Sie hier, um die Karte zu laden und Ihre Zustimmung zu erteilen.

Dr. Stemper

©

2026

Dr. Dirk Stemper

Samstag, 22.8.2026

a green flower
an orange flower
a blue flower