• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    ColGrep’s blend of semantic matching and familiar code-search filters

    Weaviate Podcast says LightOn built retrieval models small enough to index a codebase locally on a laptop, then paired them with regex and file filters in ColGrep.

    WP
    1 Source, 14d ago, first seen 14d ago

    TLDR

    Weaviate Podcast describes ColGrep as combining the regex and file filters coding agents already use with semantic matching—search based on meaning. It says LightOn’s 70-million-parameter code retriever outperformed everything up to 300 million parameters, and credits the approach with fewer repeat searches, fewer tokens spent on search and better answers on the hardest queries.

    Combined views

    662

    1 Source, first seen 14d ago

    Combined views

    662

    1 Source, first seen 14d ago

    13 likes
    13 likes
    3 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @weaviatepodcastCoding agents made grep famous, but grep is retrieval too, and it turns out you can teach it semantics. 🔍 LightOn trained late interaction models small enough to index a codebase locally on a laptop, including a 70M parameter code retriever that outperformed everything up to 300M, and wired them into ColGrep: the regex and file filters agents already know, plus semantic matching. The result is fewer re-queries, fewer tokens burned on search, and better answers on the hardest queries. The conversation with Amélie Chatelain (@AmelieTabatta) and Antoine Chaffin (@antoine_chaffin) also covers reasoning-intensive retrieval, multimodal search with patch-level vectors, and ColBERT-Zero, a late interaction model trained from scratch with their open source PyLate library. 🔥 Learn more on Weaviate Podcast #134 💚: https://www.youtube.com/watch?v=44GC3E-WbHU

    1 Source

    @weaviatepodcastCoding agents made grep famous, but grep is retrieval too, and it turns out you can teach it semantics. 🔍 LightOn trained late interaction models small enough to index a codebase locally on a laptop, including a 70M parameter code retriever that outperformed everything up to 300M, and wired them into ColGrep: the regex and file filters agents already know, plus semantic matching. The result is fewer re-queries, fewer tokens burned on search, and better answers on the hardest queries. The conversation with Amélie Chatelain (@AmelieTabatta) and Antoine Chaffin (@antoine_chaffin) also covers reasoning-intensive retrieval, multimodal search with patch-level vectors, and ColBERT-Zero, a late interaction model trained from scratch with their open source PyLate library. 🔥 Learn more on Weaviate Podcast #134 💚: https://www.youtube.com/watch?v=44GC3E-WbHU