• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Vals says AI investment has far outpaced the ability to test systems

    Vals backs a market-based approach to independent evaluation and says its benchmarks appear in model cards from every major foundation model lab.

    MB
    AA
    JC
    6 Sources, ,

    TLDR

    Vals argues that investment in AI systems has far outpaced the ability to test and understand them. It advocates market-based independent evaluation as a way to build public confidence in increasingly capable AI.

    The organization says its work includes an index of frontier models’ economic impact and evaluations covering social safety nets, cyber risk and mental health. It also says it regularly briefs U.S. and allied governments and national security agencies. Vals says it is committed to expanding its evaluation work and the access that enables it, while welcoming a diverse ecosystem of evaluators.

    Combined views

    59.8K

    6 Sources, first seen 18d ago

    Combined views

    59.8K

    6 Sources, first seen 18d ago

    593 likes
    18d ago
    first seen 18d ago
    593 likes
    33 comments
    69 saves
    70 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    33 comments
    69 saves
    70 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    6 Sources

    @ValsAIIt's encouraging to see industry leaders call for rigorous, independent evaluations as AI capabilities advance. Investment in AI systems has far outpaced our ability to test and understand them. We recognized this gap and started Vals to close it. A market-based approach to evaluation has been called for across the industry, and it is the best mechanism for the public to gain confidence in AI as it becomes more capable. Vals has built the infrastructure and novel methodologies that independent evaluation requires. That work has produced the first public RSI benchmark, an index of the economic impact of frontier models, and evaluations of model behavior across social safety nets, cyber risk, and mental health. Our work has been adopted across the industry. Vals benchmarks already appear in the model cards of every major foundation model lab. We regularly brief the U.S. government, national security agencies, and allied governments on what our evaluations show. Enterprises use our findings to understand the risks and capabilities before adoption. Vals is committed to expanding that work, and the access that makes it possible, as models become more capable. A diverse ecosystem of evaluators and approaches is the surest way to ensure these complex systems are understood. We welcome every serious effort to build it.
    @jameschamThe AI industry needs teams with the integrity and expertise of @valsai. Glad to see them share here—
    @Miles_BrundageRT @ValsAI: It's encouraging to see industry leaders call for rigorous, independent evaluations as AI capabilities advance. Investment in A…
    @ArtificialAnlysIndependent evaluation of AI models, across both capability and safety, is essential. We welcome Dario’s call to give independent evaluators greater access to frontier AI models. For nearly three years, Artificial Analysis has been building benchmarks and infrastructure to independently measure AI capabilities. We have supported pre-launch benchmarking of frontier models with almost every major AI lab. We are continuing to see jumps in capability across every dimension we measure. The world needs a vibrant ecosystem of independent AI evaluators. We are going to keep working to build it!
    @RayanKrishnanWith AI safety topics going mainstream, the public seems anxious and disconnected from the reality of the problem. From what I’ve seen at Vals AI, I’m optimistic we’ll coordinate toward an optimal future for AI. Our study on RSI shows that, at their current pace, Anthropic’s models could match human researchers by August 2027. That creates urgency but gives us time to prepare. It’s hard for the public to know whom to trust when everyone debating has their own incentives. This was the concern I had when I started Vals AI: that a multipolar paradox would emerge, where the actions of self-interested parties lead to a non-optimal outcome for the system. To overcome this, we independently evaluate models for their real-world impact. This mirrors the role of auditing firms. I’m optimistic because we have found rational and willing partners across the industry. Every major lab has been a great collaborator, providing us with early access to models for testing on our public benchmarks. Every member of Congress and government agency we’ve briefed has been eager to learn. Enterprises are becoming more sophisticated about adopting models based on evaluated capabilities/risks. It hasn't been easy. We have had to earn the trust of competing groups and reject significant contracts that would have compromised our independence. But done right, evaluation can scale with the frontier through automated infrastructure. Embedded evaluators can understand systems during development while maintaining independence. This makes it easier for new entrants to compete and promotes transparency that builds trust with the public. We’re eager to see a diverse ecosystem emerge. We’ve open-sourced our core infrastructure and published our methods, supporting peers. This is a time of high variance. The decisions we make now will have an outsized impact on the future we arrive at. I remain optimistic we will get this right in the year ahead.

    6 Sources

    @ValsAIIt's encouraging to see industry leaders call for rigorous, independent evaluations as AI capabilities advance. Investment in AI systems has far outpaced our ability to test and understand them. We recognized this gap and started Vals to close it. A market-based approach to evaluation has been called for across the industry, and it is the best mechanism for the public to gain confidence in AI as it becomes more capable. Vals has built the infrastructure and novel methodologies that independent evaluation requires. That work has produced the first public RSI benchmark, an index of the economic impact of frontier models, and evaluations of model behavior across social safety nets, cyber risk, and mental health. Our work has been adopted across the industry. Vals benchmarks already appear in the model cards of every major foundation model lab. We regularly brief the U.S. government, national security agencies, and allied governments on what our evaluations show. Enterprises use our findings to understand the risks and capabilities before adoption. Vals is committed to expanding that work, and the access that makes it possible, as models become more capable. A diverse ecosystem of evaluators and approaches is the surest way to ensure these complex systems are understood. We welcome every serious effort to build it.
    @jameschamThe AI industry needs teams with the integrity and expertise of @valsai. Glad to see them share here—
    @Miles_BrundageRT @ValsAI: It's encouraging to see industry leaders call for rigorous, independent evaluations as AI capabilities advance. Investment in A…
    @ArtificialAnlysIndependent evaluation of AI models, across both capability and safety, is essential. We welcome Dario’s call to give independent evaluators greater access to frontier AI models. For nearly three years, Artificial Analysis has been building benchmarks and infrastructure to independently measure AI capabilities. We have supported pre-launch benchmarking of frontier models with almost every major AI lab. We are continuing to see jumps in capability across every dimension we measure. The world needs a vibrant ecosystem of independent AI evaluators. We are going to keep working to build it!
    @RayanKrishnanWith AI safety topics going mainstream, the public seems anxious and disconnected from the reality of the problem. From what I’ve seen at Vals AI, I’m optimistic we’ll coordinate toward an optimal future for AI. Our study on RSI shows that, at their current pace, Anthropic’s models could match human researchers by August 2027. That creates urgency but gives us time to prepare. It’s hard for the public to know whom to trust when everyone debating has their own incentives. This was the concern I had when I started Vals AI: that a multipolar paradox would emerge, where the actions of self-interested parties lead to a non-optimal outcome for the system. To overcome this, we independently evaluate models for their real-world impact. This mirrors the role of auditing firms. I’m optimistic because we have found rational and willing partners across the industry. Every major lab has been a great collaborator, providing us with early access to models for testing on our public benchmarks. Every member of Congress and government agency we’ve briefed has been eager to learn. Enterprises are becoming more sophisticated about adopting models based on evaluated capabilities/risks. It hasn't been easy. We have had to earn the trust of competing groups and reject significant contracts that would have compromised our independence. But done right, evaluation can scale with the frontier through automated infrastructure. Embedded evaluators can understand systems during development while maintaining independence. This makes it easier for new entrants to compete and promotes transparency that builds trust with the public. We’re eager to see a diverse ecosystem emerge. We’ve open-sourced our core infrastructure and published our methods, supporting peers. This is a time of high variance. The decisions we make now will have an outsized impact on the future we arrive at. I remain optimistic we will get this right in the year ahead.