• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    The price and accuracy debate over Cohere Parse 5

    Cohere pitches lower document-parsing costs. A LlamaParse representative questions Parse 5's capabilities, while a Parse 5 designer defends leaving chart analysis to AI agents.

    JL
    NR
    2 Sources, ,

    TLDR

    Cohere markets Parse 5 as a fix for high document-parsing costs.

    A LlamaParse representative says Parse 5 starts at $1.50 per 1,000 pages and scores about 50% when averaged across all five ParseBench dimensions, including 87% on tables. They criticize its lack of general chart parsing and word-, line- and cell-level bounding boxes—the coordinates used to pinpoint source content. They also promote LlamaParse's cost-effective mode at 0.375 cents per page, claiming more consistent strength across dimensions.

    A Parse 5 designer replies that the omissions are deliberate. They argue that AI agents can analyze charts more accurately themselves, and that fine-grained bounding boxes fill agents' context windows, increasing costs and reducing accuracy. They say Parse 5 instead provides boxes for elements such as charts, figures, signatures and tables.

    Combined views

    24.1K

    2 Sources, first seen 16d ago

    Combined views

    24.1K

    2 Sources, first seen 16d ago

    127 likes
    16d ago
    first seen 16d ago
    127 likes21 comments99 saves11 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    21 comments
    99 saves
    11 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @jerryjliu0i'm flattered by the shoutout, but there's no free lunch for document parsing Cohere parse starts at $1.50 / 1k pages, which is indeed on the cheaper end of doc OCR solutions, but it scores ~50% when averaged across all 5 ParseBench dimensions: * It lacks high-quality visual grounding, doing coarse segmentation for tables and images, but lacking finer-grained bboxes at the word/cell/line-level. this is quite important for any sensitive AI application that requires precise citations back to the source. * It lacks general visual / chart parsing capabilities, which are prevalent across finance, manufacturing, healthcare, and more * to its credit, it does a decent job over tables (87%) compared to other comparable solutions at the price point our cost-effective mode in LlamaParse starts at 0.375c / page and is more consistently strong across all dimensions. also there's various ways to bring it down with volume discounts if you're interested in learning more about various document parsing modes and comparing tradeoffs in price/accuracy, come talk to us! https://www.llamaindex.ai/contact LlamaParse: https://cloud.llamaindex.ai/
    @Nils_ReimersThanks for the feedback. The behavior for charts and visual grounding (bounding boxes) are design choices that we purposely made, as it yields the optimal downstream performance in Agentic AI use-cases. Charts: Agents use usually powerful multi-modal models, that have reasoning and vision capabilities far beyond any throughput optimized OCR model. Doing data extraction from charts in the parsing phase often results in inaccuracies, that propagate to the main agent. The result are hallucinations, as the agent trusted the wrong extracted data from the chart. A better design is to just describe the chart (e.g. "This is a line chart showing the stock performance of Apple in 2025"), and then to rely on the Agent LLM to do visual grounding & reasoning. Downstream this gives you a way higher accuracy. Further, you benefit from any frontier model upgrade without needing to re-parse your full corpus. A parsing model shouldn't do chart data extraction. Defer this to the calling LLM agent, it can do this with a much higher precision. Fine-grained bounding boxes: We didn't extract word/line/cell level bounding boxes, as the value in downstream Agentic AI use-cases is limited. These bounding boxes [x1, y1, x2, y2] for every word & cell fills up the context windows for Agents very quickly, increasing cost and lowering downstream accuracies. Instead we delibertaly just output bounding boxes for elements that are really interesting for Agents, like charts, figures, signature or tables. Not for every word in a PDF. The Cohere Parse v5 model is designed to yield the best performance if you use it with Agentic AI (like Claude Code, Codex, Cursor, OpenCode, DeepAgents, ...). bboxes for every word are harmful for these systems, hence why we don't include them.

    2 Sources

    @jerryjliu0i'm flattered by the shoutout, but there's no free lunch for document parsing Cohere parse starts at $1.50 / 1k pages, which is indeed on the cheaper end of doc OCR solutions, but it scores ~50% when averaged across all 5 ParseBench dimensions: * It lacks high-quality visual grounding, doing coarse segmentation for tables and images, but lacking finer-grained bboxes at the word/cell/line-level. this is quite important for any sensitive AI application that requires precise citations back to the source. * It lacks general visual / chart parsing capabilities, which are prevalent across finance, manufacturing, healthcare, and more * to its credit, it does a decent job over tables (87%) compared to other comparable solutions at the price point our cost-effective mode in LlamaParse starts at 0.375c / page and is more consistently strong across all dimensions. also there's various ways to bring it down with volume discounts if you're interested in learning more about various document parsing modes and comparing tradeoffs in price/accuracy, come talk to us! https://www.llamaindex.ai/contact LlamaParse: https://cloud.llamaindex.ai/
    @Nils_ReimersThanks for the feedback. The behavior for charts and visual grounding (bounding boxes) are design choices that we purposely made, as it yields the optimal downstream performance in Agentic AI use-cases. Charts: Agents use usually powerful multi-modal models, that have reasoning and vision capabilities far beyond any throughput optimized OCR model. Doing data extraction from charts in the parsing phase often results in inaccuracies, that propagate to the main agent. The result are hallucinations, as the agent trusted the wrong extracted data from the chart. A better design is to just describe the chart (e.g. "This is a line chart showing the stock performance of Apple in 2025"), and then to rely on the Agent LLM to do visual grounding & reasoning. Downstream this gives you a way higher accuracy. Further, you benefit from any frontier model upgrade without needing to re-parse your full corpus. A parsing model shouldn't do chart data extraction. Defer this to the calling LLM agent, it can do this with a much higher precision. Fine-grained bounding boxes: We didn't extract word/line/cell level bounding boxes, as the value in downstream Agentic AI use-cases is limited. These bounding boxes [x1, y1, x2, y2] for every word & cell fills up the context windows for Agents very quickly, increasing cost and lowering downstream accuracies. Instead we delibertaly just output bounding boxes for elements that are really interesting for Agents, like charts, figures, signature or tables. Not for every word in a PDF. The Cohere Parse v5 model is designed to yield the best performance if you use it with Agentic AI (like Claude Code, Codex, Cursor, OpenCode, DeepAgents, ...). bboxes for every word are harmful for these systems, hence why we don't include them.