• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Training platform aims to automate research work and speed up training iteration

    A builder says observing training and inference rollouts helps show how models behave and which behaviors are undesirable.

    YP
    LL
    2 Sources, ,

    TLDR

    A person building a training platform says the team is adding intelligence to automate research work and enable much faster training iteration. They say observing training and inference rollouts helps them understand model behavior and identify undesirable behaviors. They argue that models cannot be “programmed” and built up like software applications, so building them requires new tools and infrastructure.

    Combined views

    4.2K

    2 Sources, first seen 6h ago

    Combined views

    4.2K

    2 Sources, first seen 6h ago

    52 likes
    6h ago
    first seen 6h ago
    52 likes
    3 comments
    10 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 comments
    10 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @ypatil125We're building a lot of intelligence into our training platform to automate research work and enable much faster training iteration. Good observability on training and inference rollouts helps us understand exactly how models are behaving and what behaviors are undesirable. The reality is, you can't "program" models and build them up like you can with software applications. There's a whole new set of tools and infrastructure needed to build these complex, hard to interpret, AI systems.
    @lindensliReward hacking has become a more challenging and prevalent alignment problem for us to solve as (1) the average world size of our jobs has increased dramatically -> inference compute increases, so the model gets more shots on goal to solve a task; (2) tasks have gotten more sophisticated and challenging to grade; (3) open weight base models have gotten increasingly powerful. Reward curves improving does not necessarily mean that the model is better in production. We’ve started automating this detection on all runs to make data and verifier iteration much faster.

    2 Sources

    @ypatil125We're building a lot of intelligence into our training platform to automate research work and enable much faster training iteration. Good observability on training and inference rollouts helps us understand exactly how models are behaving and what behaviors are undesirable. The reality is, you can't "program" models and build them up like you can with software applications. There's a whole new set of tools and infrastructure needed to build these complex, hard to interpret, AI systems.
    @lindensliReward hacking has become a more challenging and prevalent alignment problem for us to solve as (1) the average world size of our jobs has increased dramatically -> inference compute increases, so the model gets more shots on goal to solve a task; (2) tasks have gotten more sophisticated and challenging to grade; (3) open weight base models have gotten increasingly powerful. Reward curves improving does not necessarily mean that the model is better in production. We’ve started automating this detection on all runs to make data and verifier iteration much faster.