• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    KCoral aims to give kernel-writing AI agents shared GPU access for benchmarking

    Its builders say KCoral schedules a shared GPU pool and returns correctness and timing results for agents’ code.

    Tianqi ChenTC
    Yixin DongYD
    Ruihang LaiRL
    8 Sources, ,

    TLDR

    KCoral’s builders say agents can submit code for evaluation while the service handles GPU scheduling and isolation. They say it works across multiple GPUs and nodes through one API. They report 2.58× the throughput of sequential local execution on a single NVIDIA B200 by overlapping CPU compilation with GPU work across requests.

    Combined views

    3.9K

    8 Sources, first seen 2h ago

    Combined views

    3.9K

    8 Sources, first seen 2h ago

    64 likes
    2h ago
    first seen 2h ago
    64 likes
    7 comments
    21 saves
    44 reposts
    7 comments
    21 saves
    44 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    8 Sources

    Yixin Dong@yi_xin_dongKernel agents can already beat hand-tuned kernels. Spinning up more agents is easy, but scaling GPU evaluation is still hard and expensive. What happens when the agents scale faster than the GPUs? Meet KCoral 🪸: a lightweight shared GPU benchmark environment for agentic kernel development. 🧑‍💻 Agents keep writing and revising code in their own workspace. ⚡ Evaluation is sent to KCoral, which schedules a shared GPU pool and returns correctness and timing. 🌐 Scale across multiple GPUs and nodes through a unified API endpoint. 🔒 Per-request isolation with only milliseconds of execution overhead. By overlapping CPU compilation with GPU work across requests, KCoral further speeds up kernel evaluation. On a single NVIDIA B200, this translates to 2.58× the throughput of sequential local execution.2h
    Ruihang Lai@ruihanglaiWe build KCoral 🪸 to make GPU access a service for kernel agents. As kernel generation and RL training scale, giving each agent a GPU wastes capacity; leaving agents to coordinate sharing adds work and risks unreliable benchmarks. Agents submit work as needed, and KCoral handles scheduling and isolation—keeping GPUs productive and performance feedback reliable. More in our blogpost 👇2h
    Hongyi Jin@HongyiJin258very useful remote execution service when having a large batch of concurrent kernel agents with limited gpu resources1h
    Tianqi Chen@tqchenmlRT @yi_xin_dong: Kernel agents can already beat hand-tuned kernels. Spinning up more agents is easy, but scaling GPU evaluation is still ha…1h
    Genghan Zhang@zhang677Great collaboration, and excited to build KCoral together! 🚀 Robust and scalable evaluation is essential for building better kernel agents. Excited to bring what we’ve learned from building PTXBench into KCoral and push this space forward together.1h

    8 Sources

    Yixin Dong@yi_xin_dongKernel agents can already beat hand-tuned kernels. Spinning up more agents is easy, but scaling GPU evaluation is still hard and expensive. What happens when the agents scale faster than the GPUs? Meet KCoral 🪸: a lightweight shared GPU benchmark environment for agentic kernel development. 🧑‍💻 Agents keep writing and revising code in their own workspace. ⚡ Evaluation is sent to KCoral, which schedules a shared GPU pool and returns correctness and timing. 🌐 Scale across multiple GPUs and nodes through a unified API endpoint. 🔒 Per-request isolation with only milliseconds of execution overhead. By overlapping CPU compilation with GPU work across requests, KCoral further speeds up kernel evaluation. On a single NVIDIA B200, this translates to 2.58× the throughput of sequential local execution.2h
    Ruihang Lai@ruihanglaiWe build KCoral 🪸 to make GPU access a service for kernel agents. As kernel generation and RL training scale, giving each agent a GPU wastes capacity; leaving agents to coordinate sharing adds work and risks unreliable benchmarks. Agents submit work as needed, and KCoral handles scheduling and isolation—keeping GPUs productive and performance feedback reliable. More in our blogpost 👇2h
    Hongyi Jin@HongyiJin258very useful remote execution service when having a large batch of concurrent kernel agents with limited gpu resources1h
    Tianqi Chen@tqchenmlRT @yi_xin_dong: Kernel agents can already beat hand-tuned kernels. Spinning up more agents is easy, but scaling GPU evaluation is still ha…1h
    Genghan Zhang@zhang677Great collaboration, and excited to build KCoral together! 🚀 Robust and scalable evaluation is essential for building better kernel agents. Excited to bring what we’ve learned from building PTXBench into KCoral and push this space forward together.1h