• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    The case that PDFs aren't built for machines

    A reshared post argues that PDF text appears as glyphs with coordinates, while tables appear as line segments rather than table structures.

    JL
    CJ
    2 Sources, ,

    TLDR

    The reshared post argues that PDFs aren't built for machines, pointing to how they represent text as glyphs with coordinates and tables as line segments.

    Combined views

    819

    2 Sources, first seen 7h ago

    4 likes

    Combined views

    819

    2 Sources, first seen 7h ago

    4 likes
    7h ago
    first seen 7h ago
    3 comments
    3 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    3 comments
    3 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @CoreyGallonPDFs aren't built for machines: text shows up as glyphs with coordinates, tables as line segments instead of table structures. @jerryjliu0, CEO of LlamaIndex, digs into what that means for AI agents in "Building the Document Context Layer for AI Agents," posted by @aiDotEngineer on YouTube. The talk lays out what it actually takes to turn the messy stack of enterprise documents into context an agent can use, from parsing through search to structured extraction. - 10 trillion-plus pages, locked up. That's how much human-native knowledge Jerry pegs as sitting in PDFs, PowerPoints, Word docs, and Excel sheets today. - Parsing is hard because the format fights you. PDFs are built for printing, so text comes in as individual glyphs with coordinates and tables render as line segments, not table structures. Word and PowerPoint aren't much better: their XML is bespoke and full of fluff an agent has to see through. - Pipeline parsers vs. vision models, and why LlamaIndex uses both. Classic tools like PyPDF and PyMuPDF handle structure well but miss visual context; VLM-based parsing reads the page but can hallucinate on text-only content and gets expensive fast. Jerry describes combining both, with auto-routing between cheaper specialized models and frontier ones. - ParseBench. A public benchmark of 2,000 human-verified pages testing document parsing across roughly 50 models on tables, charts, content faithfulness, and semantic formatting, live at http://parsebench.ai and mirrored on Hugging Face and Kaggle. - LightParse. A free, Rust-based, MIT/Apache-licensed open source markdown parser built for a fast first pass over a big batch of documents, before handing the harder pages to a VLM-based tool. - Structured extraction at scale. Pulling data out of invoices, receipts, claims, and contracts with confidence scores and citations back to the source document. I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!
    @jerryjliu0RT @CoreyGallon: PDFs aren't built for machines: text shows up as glyphs with coordinates, tables as line segments instead of table structu…

    2 Sources

    @CoreyGallonPDFs aren't built for machines: text shows up as glyphs with coordinates, tables as line segments instead of table structures. @jerryjliu0, CEO of LlamaIndex, digs into what that means for AI agents in "Building the Document Context Layer for AI Agents," posted by @aiDotEngineer on YouTube. The talk lays out what it actually takes to turn the messy stack of enterprise documents into context an agent can use, from parsing through search to structured extraction. - 10 trillion-plus pages, locked up. That's how much human-native knowledge Jerry pegs as sitting in PDFs, PowerPoints, Word docs, and Excel sheets today. - Parsing is hard because the format fights you. PDFs are built for printing, so text comes in as individual glyphs with coordinates and tables render as line segments, not table structures. Word and PowerPoint aren't much better: their XML is bespoke and full of fluff an agent has to see through. - Pipeline parsers vs. vision models, and why LlamaIndex uses both. Classic tools like PyPDF and PyMuPDF handle structure well but miss visual context; VLM-based parsing reads the page but can hallucinate on text-only content and gets expensive fast. Jerry describes combining both, with auto-routing between cheaper specialized models and frontier ones. - ParseBench. A public benchmark of 2,000 human-verified pages testing document parsing across roughly 50 models on tables, charts, content faithfulness, and semantic formatting, live at http://parsebench.ai and mirrored on Hugging Face and Kaggle. - LightParse. A free, Rust-based, MIT/Apache-licensed open source markdown parser built for a fast first pass over a big batch of documents, before handing the harder pages to a VLM-based tool. - Structured extraction at scale. Pulling data out of invoices, receipts, claims, and contracts with confidence scores and citations back to the source document. I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!
    @jerryjliu0RT @CoreyGallon: PDFs aren't built for machines: text shows up as glyphs with coordinates, tables as line segments instead of table structu…