• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    IBM Researchers Introduce STAIR Retriever for RAG

    Elvis Saravia highlights IBM paper proposing table of contents for RAG retrieval.

    EL
    1 Source, 24d ago, first seen 24d ago

    TLDR

    Elvis Saravia posted about a paper from Vineet Kumar and IBM colleagues on STAIR, a structure-aware information retriever. The work targets common RAG problems where chunking documents by length discards existing hierarchy. It proposes encoding global structure via table of contents instead. The linked source describes STAIR as a generative retriever that stores and addresses a corpus through its table of contents, along with a novel dataset for document structure augmentation.

    Combined views

    19.5K

    1 Source, first seen 24d ago

    Combined views

    19.5K

    1 Source, first seen 24d ago

    327 likes
    327 likes
    29 comments
    335 saves
    54 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    29 comments
    335 saves
    54 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @omarsar0Great RAG paper from IBM. There are some really good ideas on how to solve common RAG issues. It's well known that retrievers chunk long documents by length, which discards the hierarchy the document already has. So they propose using a table of contents. A table of contents helps to encodes exactly the global structure that chunking throws away. STAIR uses that table of contents as the addressing scheme for a generative retriever, so the model stores and retrieves information from its own parameters against a structure the corpus supplies. On SearchTome, it reaches Recall@1 of 82.6 percent against 76.9 percent for a fine-tuned Differentiable Search Index, a statistically significant gap, with BM25 at 59.5 percent and DPR at 68.7 percent. Hallucination stays below 0.05 percent, which is the standing objection to generative retrieval and the reason grounding the address space in a real hierarchy is worth the extra structure. The ablations also show it generalizes where very few training samples exist. Paper: https://academy.dair.ai/papers/stair-structure-aware-information-retriever-a-novel-dataset-and-llm-based-retrie-2609.03874

    1 Source

    @omarsar0Great RAG paper from IBM. There are some really good ideas on how to solve common RAG issues. It's well known that retrievers chunk long documents by length, which discards the hierarchy the document already has. So they propose using a table of contents. A table of contents helps to encodes exactly the global structure that chunking throws away. STAIR uses that table of contents as the addressing scheme for a generative retriever, so the model stores and retrieves information from its own parameters against a structure the corpus supplies. On SearchTome, it reaches Recall@1 of 82.6 percent against 76.9 percent for a fine-tuned Differentiable Search Index, a statistically significant gap, with BM25 at 59.5 percent and DPR at 68.7 percent. Hallucination stays below 0.05 percent, which is the standing objection to generative retrieval and the reason grounding the address space in a real hierarchy is worth the extra structure. The ablations also show it generalizes where very few training samples exist. Paper: https://academy.dair.ai/papers/stair-structure-aware-information-retriever-a-novel-dataset-and-llm-based-retrie-2609.03874