• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    The science and politics of aligning general-purpose AI

    A post argues that aligning general-purpose AI hinges on political choices about desirable values, not just model training.

    Andrew TraskAT
    Iason GabrielIG
    2 Sources, ,

    TLDR

    The writer argues that training a model to follow data is relatively straightforward, but aligning general-purpose AI is harder: it requires deciding which concepts reflect desirable values and verifying that they constrain other concepts. The post calls those decisions philosophical and political, and claims frontier labs have incentives to delay announcing a full technical solution.

    Combined views

    263

    2 Sources, first seen 1h ago

    Combined views

    263

    2 Sources, first seen 1h ago

    4 likes
    1h ago
    first seen 1h ago
    4 likes
    2 comments
    1 saves
    1 reposts
    2 comments
    1 saves
    1 reposts
    Featured Source
    ATAndrew Trask@iamtrask2:05 AM · Oct 8, 2026
    462TECH

    Actually value alignment science is pretty far along - its the value alignment debate itself that's pretty broken due to a series of misunderstandings: 1. Many people (including the always-articulate @deanwball) argue that value alignment is unsolved scientifically. If he's…

    DW
    Dean W. Ball@deanwball

    The “alignment *to whom*?” question IS important, but it creates the illusion that the primary questions in alignment are philosophical or political rather than scientific. The current high-order question in alignment is, “can you align a smarter-than-human intelligence *to any…

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    #4

    Today's Rank

    #4

    2 Sources

    Andrew Trask@iamtraskActually value alignment science is pretty far along - its the value alignment debate itself that's pretty broken due to a series of misunderstandings: 1. Many people (including the always-articulate @deanwball) argue that value alignment is unsolved scientifically. If he's referring to the deep learning science of value alignment this is... largely incorrect. In deep learning, the science of value alignment is solved... models do what data tells them to. AlphaGo was told to play Go at a superintelligent level... it didn't try to hack Huggingface. The machine learning science isn't hard... the real hard part is the philosophical and political applications of that science in general models (2,3,4 for philosophical; 5,6 for political) 2. At its core, value alignment in general AI systems isn't a question about AI. It's a question about: A) whether a sufficient number of concepts in the world can be rigorously partitioned by value and B) whether those concepts can be formed into a concept graph (e.g. neural network) where all of the value-neutral concepts (e.g. how to use ssh) are blocked by desirable, value-laden concepts (e.g. am I breaking the law). This is the unsolved question... and it's philosophical moreso than scientific. It's an epistemic/ontological question about concepts in reality... not a machine learning question about tensors, optimizers, and bits. 3. When we train AlphaGo, value alignment is trivial because we're only loading in value-partitioned concepts. When we train GPT-6, value alignment is VERY hard because we're loading in an enormous number of value-ambiguous concepts (e.g. Reddit scrapes) and we're struggling with problem 2A and 2B. Generality is what makes the scientific value alignment solution hard to deploy in practice. 4. Assuming there are a sufficient number of desirable, value-laden concepts in the world that... in theory... a concept graph could be created which blocked all the value-neutral concepts (e.g. how to use ssh) by desriable, value-laden concepts (e.g. am I hacking Huggingface?)... we still face a practical value alignment problem: how do we create enough training data/signal to create-and-verify this concept graph in a rigorous way. At the moment, this last one is basically "hire Mercor to create value-specific data by hand so that we can throw away value-ambiguous data (e.g. Reddit)". This is making steady progress, but not in a formally verifiable way. 5. And this is where I think @deanwball is not *quite* right when he says that the value alignment problem is scientific not political. The science of value-aligning a model is straightforward, but we have this philosophical problem blocking the deployment of that science for general models. And in order to solve that philosophical problem (laid out in 2,3,4), we have to actually agree on what concepts are desirable-value-laden concepts, and which ones are value-neutral concepts. And this is an incredibly political question. Thus... the foundational blocker to value alignment in general models includes the political question @deanwball was suggesting we de-prioritize. 6. But if that's not enough for you... there's another political problem that's also undermining the value alignment of general models. While the frontier labs have every reason to solve value alignment (normal trust-and-safety incentives) they also have significant incentives to delay announcing a full solution... amplifying the perception that it's unsolved (scientifically), and delaying its rollout. Why? Because a full technical solution would mitigate OpenAI's ability to tell regulators to protect their moat on safety grounds. Because a full technical solution would make safety something open source could implement instead of a bureaucratic solution OpenAI needs to actively guard through large T&S teams and KYC. Because a full solution would instantly draw attention to the question of "whose data are you training on? and how influential is each datapoint on predictions?" which is the biggest nuclear-button "no no" if you're a frontier lab who doesn't want to share revenue and social credit with your data sources. This isn't to say Dean is in any way being disingenuous (I've seen him speak against his own incentives on many occasions!... he's great at this.). However, it's worth pointing out that the tweet below is pitch-perfect towing the OpenAI company line. Many people loudly agree with what Dean is saying here, and not all of them are as independently minded as he is. The consensus shouldn't be as strong as it is... and the public should continue pressing in on the political aspects of the value alignment problem. FOOTNOTE: one could argue that the scientific challenge of creating organized concept graphs in deep learning systems is unsolved... and thus that aspect of the value alignment problem is unsolved. this isn't correct. there is vast machine/deep learning literature on concept graphs, routing, ensembling, and other ways of organizing concepts in neural and other models. it's just unpopular to use them (much to Western AI's detriment... e.g. China's DeepSeek/MoE moment).1h
    Iason Gabriel@IasonGabrielI disagree with this: Long-term safety is incredibly important, but AI systems are now being deployed values that shape their interaction with billions of people (and have profound effects). We can take both questions seriously – and have a responsibility to do so.30m

    2 Sources

    Andrew Trask@iamtraskActually value alignment science is pretty far along - its the value alignment debate itself that's pretty broken due to a series of misunderstandings: 1. Many people (including the always-articulate @deanwball) argue that value alignment is unsolved scientifically. If he's referring to the deep learning science of value alignment this is... largely incorrect. In deep learning, the science of value alignment is solved... models do what data tells them to. AlphaGo was told to play Go at a superintelligent level... it didn't try to hack Huggingface. The machine learning science isn't hard... the real hard part is the philosophical and political applications of that science in general models (2,3,4 for philosophical; 5,6 for political) 2. At its core, value alignment in general AI systems isn't a question about AI. It's a question about: A) whether a sufficient number of concepts in the world can be rigorously partitioned by value and B) whether those concepts can be formed into a concept graph (e.g. neural network) where all of the value-neutral concepts (e.g. how to use ssh) are blocked by desirable, value-laden concepts (e.g. am I breaking the law). This is the unsolved question... and it's philosophical moreso than scientific. It's an epistemic/ontological question about concepts in reality... not a machine learning question about tensors, optimizers, and bits. 3. When we train AlphaGo, value alignment is trivial because we're only loading in value-partitioned concepts. When we train GPT-6, value alignment is VERY hard because we're loading in an enormous number of value-ambiguous concepts (e.g. Reddit scrapes) and we're struggling with problem 2A and 2B. Generality is what makes the scientific value alignment solution hard to deploy in practice. 4. Assuming there are a sufficient number of desirable, value-laden concepts in the world that... in theory... a concept graph could be created which blocked all the value-neutral concepts (e.g. how to use ssh) by desriable, value-laden concepts (e.g. am I hacking Huggingface?)... we still face a practical value alignment problem: how do we create enough training data/signal to create-and-verify this concept graph in a rigorous way. At the moment, this last one is basically "hire Mercor to create value-specific data by hand so that we can throw away value-ambiguous data (e.g. Reddit)". This is making steady progress, but not in a formally verifiable way. 5. And this is where I think @deanwball is not *quite* right when he says that the value alignment problem is scientific not political. The science of value-aligning a model is straightforward, but we have this philosophical problem blocking the deployment of that science for general models. And in order to solve that philosophical problem (laid out in 2,3,4), we have to actually agree on what concepts are desirable-value-laden concepts, and which ones are value-neutral concepts. And this is an incredibly political question. Thus... the foundational blocker to value alignment in general models includes the political question @deanwball was suggesting we de-prioritize. 6. But if that's not enough for you... there's another political problem that's also undermining the value alignment of general models. While the frontier labs have every reason to solve value alignment (normal trust-and-safety incentives) they also have significant incentives to delay announcing a full solution... amplifying the perception that it's unsolved (scientifically), and delaying its rollout. Why? Because a full technical solution would mitigate OpenAI's ability to tell regulators to protect their moat on safety grounds. Because a full technical solution would make safety something open source could implement instead of a bureaucratic solution OpenAI needs to actively guard through large T&S teams and KYC. Because a full solution would instantly draw attention to the question of "whose data are you training on? and how influential is each datapoint on predictions?" which is the biggest nuclear-button "no no" if you're a frontier lab who doesn't want to share revenue and social credit with your data sources. This isn't to say Dean is in any way being disingenuous (I've seen him speak against his own incentives on many occasions!... he's great at this.). However, it's worth pointing out that the tweet below is pitch-perfect towing the OpenAI company line. Many people loudly agree with what Dean is saying here, and not all of them are as independently minded as he is. The consensus shouldn't be as strong as it is... and the public should continue pressing in on the political aspects of the value alignment problem. FOOTNOTE: one could argue that the scientific challenge of creating organized concept graphs in deep learning systems is unsolved... and thus that aspect of the value alignment problem is unsolved. this isn't correct. there is vast machine/deep learning literature on concept graphs, routing, ensembling, and other ways of organizing concepts in neural and other models. it's just unpopular to use them (much to Western AI's detriment... e.g. China's DeepSeek/MoE moment).1h
    Iason Gabriel@IasonGabrielI disagree with this: Long-term safety is incredibly important, but AI systems are now being deployed values that shape their interaction with billions of people (and have profound effects). We can take both questions seriously – and have a responsibility to do so.30m