• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI Companies Shred Rare Books for Training Data

    Firms buy pre-2022 rare books via anonymous services, scan them, and destroy the originals to create clean training datasets.

    SZ
    DR
    GM
    65 Sources, 66d ago, first seen 66d ago

    TLDR

    AI firms including Anthropic purchase rare pre-2022 books in bulk through services like ISBNdb under NDAs. High-speed scanners cut spines and destroy the physical copies after digitization. The goal is high-quality training data free of synthetic text. A federal judge has ruled the scanning qualifies as fair use. Critics describe the process as book burning that risks permanent loss of unique volumes once the companies dissolve. The practice continues as firms prioritize clean datasets over preservation of the originals.

    Combined views

    4.9M

    65 Sources, first seen 66d ago

    Combined views

    4.9M

    65 Sources, first seen 66d ago

    91.6K likes
    91.6K likes
    1.8K comments
    10.5K saves
    35.4K reposts
    AnthropicISBNdb

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1.8K comments
    10.5K saves
    35.4K reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    65 Sources

    @suchenzangyeah this kind of sucks but also it's not like you were ever going to read that rare book anyway. at least now it dies an honorable death in the service of creating Superintelligence, and the commies will make sure to distill and distribute its benefits to all!
    @sytelusThis bad because it is depriving general public and all futire AI trainers from depriving these books. At some point, these companies will go bust and their properties will be sold off in auction for pennies on dollar and aquirer may simply not understand the value of these treasures and could get lost forever. Furthermore, it is downright incorrect to assume that no one ever reads them. Powells in Portland has special section on these books and there are many people including scholars hang out there studying them.
    @beffjezosAnthropic are literally effectively burning books and want to be the exclusive source of truth in the future.
    @DanielleFongI think this is really bad and a stupid consequence of laws that we have not reconsidered. I don't think the justifications are sufficient. Why the hell are we doing this? The AI is being told I can't reproduce more than 15 words from these rare books. No one gets to read them? fuck this!
    @xuanalogueRT @bradrcarson: So AI labs are buying old books by the pallet, slicing them apart, scanning the pages, and pulping what's left. The orders…
    @BrianRoemmeleRT @BrianRoemmele: @HedgieMarkets Thank you for writing this. It is something I have tried to alert folks about, quite repeatedly for the l…
    @packyMIn 2007, @pmarca called Vernor Vinge "one of the best forecasters in the world" along with Arthur C. Clarke. Rainbows End had just come out, and Marc said it was "the clearest and most plausible extrapolation of modern technology trends forward to the year 2025 that you can imagine." Central to Rainbows End was the Librareome Project, a mass-digitization effort at UC San Diego's Geisel Library in which books are fed into machines that cut the bindings, shred and toss the pages into the air, and photograph them mid-flight from every angle, destroying the physical books to reconstruct them as digital data as fast as possible. I'll be damned, Andreessen called that Vinge called it.
    @jasoncrawfordSeems like there is a really simple solution: the AI companies should also contribute the scan to an online repository like the @internetarchive, so there is a digital copy available. AI gets to read the books, and people get better access. Win-win
    @yishan@DanielleFong This is even worse because the books are not actually preserved verbatim inside the LLM, they are chopped up into tokens distributed across billions of weights so their full reconstruction is unreliable even if allowed.
    @chamathAnd you’re worried about distillation? This is sad if true. A slowdown in tokenmaxxing would stop this madness.

    Related

    FTC reportedly probes Anthropic, OpenAI and other AI labs over consumer risks

    Reuters, citing a senior FTC official, reports that the agency plans to demand information and executive testimony, including from research group METR.

    Anthropic's IPO prospectus reportedly warns of 'existential risks to humanity'

    A user pairs that claimed warning with a claim that OpenAI scrapped a new model rollout over safety concerns, pushing back on criticism of the EU AI Act.

    Anthropic reportedly plans IPO warning on existential AI risks

    Reuters reports Anthropic plans to caution potential IPO investors that advanced AI could pose “catastrophic or existential risks to humanity.”

    65 Sources

    @suchenzangyeah this kind of sucks but also it's not like you were ever going to read that rare book anyway. at least now it dies an honorable death in the service of creating Superintelligence, and the commies will make sure to distill and distribute its benefits to all!
    @sytelusThis bad because it is depriving general public and all futire AI trainers from depriving these books. At some point, these companies will go bust and their properties will be sold off in auction for pennies on dollar and aquirer may simply not understand the value of these treasures and could get lost forever. Furthermore, it is downright incorrect to assume that no one ever reads them. Powells in Portland has special section on these books and there are many people including scholars hang out there studying them.
    @beffjezosAnthropic are literally effectively burning books and want to be the exclusive source of truth in the future.
    @DanielleFongI think this is really bad and a stupid consequence of laws that we have not reconsidered. I don't think the justifications are sufficient. Why the hell are we doing this? The AI is being told I can't reproduce more than 15 words from these rare books. No one gets to read them? fuck this!
    @xuanalogueRT @bradrcarson: So AI labs are buying old books by the pallet, slicing them apart, scanning the pages, and pulping what's left. The orders…
    @BrianRoemmeleRT @BrianRoemmele: @HedgieMarkets Thank you for writing this. It is something I have tried to alert folks about, quite repeatedly for the l…
    @packyMIn 2007, @pmarca called Vernor Vinge "one of the best forecasters in the world" along with Arthur C. Clarke. Rainbows End had just come out, and Marc said it was "the clearest and most plausible extrapolation of modern technology trends forward to the year 2025 that you can imagine." Central to Rainbows End was the Librareome Project, a mass-digitization effort at UC San Diego's Geisel Library in which books are fed into machines that cut the bindings, shred and toss the pages into the air, and photograph them mid-flight from every angle, destroying the physical books to reconstruct them as digital data as fast as possible. I'll be damned, Andreessen called that Vinge called it.
    @jasoncrawfordSeems like there is a really simple solution: the AI companies should also contribute the scan to an online repository like the @internetarchive, so there is a digital copy available. AI gets to read the books, and people get better access. Win-win
    @yishan@DanielleFong This is even worse because the books are not actually preserved verbatim inside the LLM, they are chopped up into tokens distributed across billions of weights so their full reconstruction is unreliable even if allowed.
    @chamathAnd you’re worried about distillation? This is sad if true. A slowdown in tokenmaxxing would stop this madness.