• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    LlamaIndex uses Markdown by default for parsing, switching to HTML for tables with merged headers

    LlamaIndex says Markdown keeps headings, lists and tables intact and stays readable when debugging bad answers.

    Jerry LiuJL
    LlamaIndex 🦙L🦙
    2 Sources, 2h ago, first seen 2h ago

    TLDR

    LlamaIndex says Markdown is its default parsing output because it preserves headings, lists and tables in a readable format. A parser can capture every word but lose which column a number belongs to, it says, leaving a model to guess. For tables with merged headers, LlamaIndex switches to HTML.

    Combined views

    6.1K

    2 Sources, first seen 2h ago

    Combined views

    6.1K

    2 Sources, first seen 2h ago

    59 likes
    59 likes
    7 comments
    56 saves
    4 reposts
    Featured Source
    7 comments
    56 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    LlamaIndex 🦙@llama_indexMarkdown is all you need. (Mostly.) A parser can get every word on the page right and still lose which column a number belongs to. Then your model has to guess. Markdown keeps headings, lists, and tables intact, and it stays readable when you're debugging a bad answer. For tables with merged headers, we switch to HTML. Read our breakdown on why it's our default output for parsing below! ⬇️2h
    Jerry Liu@jerryjliu0RT @llama_index: Markdown is all you need. (Mostly.) A parser can get every word on the page right and still lose which column a number…2h

    2 Sources

    LlamaIndex 🦙@llama_indexMarkdown is all you need. (Mostly.) A parser can get every word on the page right and still lose which column a number belongs to. Then your model has to guess. Markdown keeps headings, lists, and tables intact, and it stays readable when you're debugging a bad answer. For tables with merged headers, we switch to HTML. Read our breakdown on why it's our default output for parsing below! ⬇️2h
    Jerry Liu@jerryjliu0RT @llama_index: Markdown is all you need. (Mostly.) A parser can get every word on the page right and still lose which column a number…2h