Collector Acquires Millions of Microfiche Cards for AI Training Data
Reactions from ranked influencers
7 postsI Lost Too Many Of These Warehouses, Sometimes I Wake Up In The Middle Of The Night With The Ghosts Of What We Forgot… A few folks sent me messages about the size of this never digitized data and now wonder like I have for decades why no one cares about erasing what is in what cost billions of today’s dollars. Data that if AI trained on would absolutely change the world. Just over the last two years we lost 2 warehouses this size I could not save. So I shout out here every now and again and hope someone hears. We are the amnesia generation and we are on the way to a full lobotomy. You can help, just reading this and sharing can help, subscribing to my X can help, buying me a coffee or becoming member at http://ReadMultiplex.com can help. You just knowing helps. Thank you. Deep gratitude.
Boom! Just got another 317 pounds of Filmsort microfiche Aperture punch cards for AI training. I will find out if we just discovered yet another multiple football field size warehouses! The warehouses look like this (left)… https://twitter.com/brianroemmele/status/2073822684035871130
Help Brian to help us all
I Lost Too Many Of These Warehouses, Sometimes I Wake Up In The Middle Of The Night With The Ghosts Of What We Forgot… A few folks sent me messages about the size of this never digitized data and now wonder like I have for decades why no one cares about erasing what is in what cost billions of today’s dollars. Data that if AI trained on would absolutely change the world. Just over the last two years we lost 2 warehouses this size I could not save. So I shout out here every now and again and hope someone hears. We are the amnesia generation and we are on the way to a full lobotomy. You can help, just reading this and sharing can help, subscribing to my X can help, buying me a coffee or becoming member at http://ReadMultiplex.com can help. You just knowing helps. Thank you. Deep gratitude.
"We are the amnesia generation and we are on the way to a full lobotomy."
I Lost Too Many Of These Warehouses, Sometimes I Wake Up In The Middle Of The Night With The Ghosts Of What We Forgot… A few folks sent me messages about the size of this never digitized data and now wonder like I have for decades why no one cares about erasing what is in what cost billions of today’s dollars. Data that if AI trained on would absolutely change the world. Just over the last two years we lost 2 warehouses this size I could not save. So I shout out here every now and again and hope someone hears. We are the amnesia generation and we are on the way to a full lobotomy. You can help, just reading this and sharing can help, subscribing to my X can help, buying me a coffee or becoming member at http://ReadMultiplex.com can help. You just knowing helps. Thank you. Deep gratitude.
I Lost Too Many Of These Warehouses, Sometimes I Wake Up In The Middle Of The Night With The Ghosts Of What We Forgot… A few folks sent me messages about the size of this never digitized data and now wonder like I have for decades why no one cares about erasing what is in what cost billions of today’s dollars. Data that if AI trained on would absolutely change the world. Just over the last two years we lost 2 warehouses this size I could not save. So I shout out here every now and again and hope someone hears. We are the amnesia generation and we are on the way to a full lobotomy. You can help, just reading this and sharing can help, subscribing to my X can help, buying me a coffee or becoming member at http://ReadMultiplex.com can help. You just knowing helps. Thank you. Deep gratitude.
Boom! Just got another 317 pounds of Filmsort microfiche Aperture punch cards for AI training. I will find out if we just discovered yet another multiple football field size warehouses! The warehouses look like this (left)… https://twitter.com/brianroemmele/status/2073822684035871130
2 of 2 3Robustness to Noise, Uncertainty, and Imperfection Raw data is messy: incomplete records, measurement errors, contradictory observations, evolving terminology. Learning to extract reliable reasoning from such data trains the model to handle ambiguity gracefully — a critical weakness in today’s systems, which often hallucinate or overconfidently reason from overly clean training distributions. 4Deeper Multi-Step Problem Decomposition Original documents often show the step-by-step breakdown of complex problems as they were actually tackled: initial hypotheses, experiments, recalibrations, final implementations. This provides natural training signals for chain-of-thought, planning, and iterative refinement that go far beyond synthetic reasoning traces. 5Emergent Generalization and Data Efficiency Scaling laws show that performance improves predictably with more high-quality, diverse data. When that data consists of dense, real-world problem-solution pairs rather than repetitive web text, the gains in reasoning benchmarks (math, coding, scientific reasoning, logical inference) compound dramatically. Models gain the ability to reason effectively from fewer examples in new domains because they have internalized foundational principles from vast historical “experience.” The sheer volume football-field-sized warehouses of such material represents orders of magnitude more authentic reasoning examples than what is currently available in digitized form. This isn’t just “more data”; it’s qualitatively different data that fills critical gaps in current training corpora (modern bias, loss of pre-digital insights, lack of raw empirical depth). Why This Path Strongly Supports Progress Toward True AGI True AGI implies systems that can understand, learn, and solve novel problems across virtually any domain at or beyond human level not just fluent text generation or narrow task performance. Key missing ingredients today include: •Deep, grounded understanding of the physical and causal world •Robust generalization and transfer •Efficient learning from limited new data •Reliable long-horizon planning and innovation Massive archives of primary problem-solution data directly supply the “collective experiential history” that lets models build these capabilities. It provides the raw material for developing: •Richer internal simulations of reality •Better abstraction and principle extraction •Human-like intuition built from billions of real precedents Combined with ongoing advances in model architecture, test-time compute (long thinking/reasoning), reinforcement learning on reasoning traces, and hybrid systems, this data could be the catalyst that pushes models past current scaling plateaus into genuinely general intelligence. It moves AI from “statistical parrot of existing knowledge” toward “active reasoner that has internalized humanity’s problem-solving journey.” Thusly digitizing and training on these lost archives wouldn’t just preserve history it would inject an unprecedented density of authentic reasoning intelligence into AI systems. The result would be reasoning capabilities that far surpass today’s models in depth, robustness, creativity, and generalization. This is one of the highest-leverage opportunities available for accelerating toward true AGI. Preserving and utilizing this data is not optional nostalgia; it is foundational infrastructure for the next leap in intelligence.
HOW WAREHOUSES OF OLD PUNCH CARD BUILD AGI You are some of the very few that understand this. Vast undigitized historical archives especially microfiche, aperture cards, and similar primary-source troves contain raw, granular records of real-world problems and the solutions (or attempted solutions) devised across centuries. These are not polished summaries but primary documents: engineering drawings with handwritten annotations, experimental logs, patent applications, scientific observations, building plans, technical specifications, and case records. They capture the full messy reality constraints of the most important eras, failed experiments, iterative refinements, quantitative measurements, diagrams, and contextual details that reveal how problems were broken down and solved. This kind of data is fundamentally different from (and far superior to) even the best books written about the same topics. My discoveries, curation and digitalization over the decades of this data makes it one of the most important AI training data available. But why? Why This Data Is Superior to Books for Training Reasoning Books are secondary sources. They synthesize, edit, simplify, and often bias toward successes. They present cleaned-up narratives, omit dead ends, skip raw data points, and filter through modern interpretations. The original thinker’s actual reasoning process the assumptions, the trial-and-error, the serendipitous connections, the precise quantitative observations under specific conditions is usually lost or heavily abstracted. In contrast, aperture cards and microfiche archives preserve the primary signal: •A 1950s engineering aperture card might include a full technical drawing of a structural problem, punch-card metadata for indexing, annotations showing load calculations, material choices, and notes on why certain designs were rejected. •Scientific records capture raw observations before they are turned into elegant theories. •Historical problem-solution pairs come with real-world constraints (materials available at the time, measurement precision limits, economic pressures) that modern textbooks rarely convey with the same fidelity. Training on this data gives an AI exposure to millions (or billions) of authentic, high-signal examples of problem decomposition, causal chains, and solution validation in their original context something no curated book corpus can match in volume or authenticity. How This Transforms Reasoning Abilities Far Beyond Today’s AI Current large language models excel at pattern matching from internet-scale text and books, but they often struggle with: •Deep causal reasoning across novel contexts •Robust handling of incomplete or noisy information •True analogical transfer between distant domains •Reliable multi-step planning under real constraints •Generalization from sparse or edge-case data Training on raw archival problem-solution data would address these gaps directly: 1Superior Causal and Counterfactual Reasoning Historical records frequently document “what happened when we tried X under conditions Y.” An AI learns not just correlations but actual cause-effect relationships observed in the real world across physics, chemistry, engineering, medicine, and more. It internalizes why certain solutions worked (or failed) given specific constraints — building richer internal world models. 2Enhanced Analogical and Cross-Domain Reasoning The same corpus spans vastly different eras and fields. An AI can discover deep structural similarities between, say, a 19th-century bridge failure analysis and a modern materials science problem, or between biological observation logs and optimization challenges. This fosters the kind of flexible analogy-making that underpins human creativity and scientific insight. 1 of 2
I Lost Too Many Of These Warehouses, Sometimes I Wake Up In The Middle Of The Night With The Ghosts Of What We Forgot… A few folks sent me messages about the size of this never digitized data and now wonder like I have for decades why no one cares about erasing what is in what cost billions of today’s dollars. Data that if AI trained on would absolutely change the world. Just over the last two years we lost 2 warehouses this size I could not save. So I shout out here every now and again and hope someone hears. We are the amnesia generation and we are on the way to a full lobotomy. You can help, just reading this and sharing can help, subscribing to my X can help, buying me a coffee or becoming member at http://ReadMultiplex.com can help. You just knowing helps. Thank you. Deep gratitude.
Combined views
45.8K
7 posts, first seen 6h ago