• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Schmidhuber Spotlights Economies of Minds in NLSOMs Paper

    AI researcher promotes his award-winning 2023 paper on natural language-based societies of mind.

    JS
    1 Source, 28d ago, first seen 28d ago

    TLDR

    Jürgen Schmidhuber posted that his award-winning 2023 paper Mindstorms in Natural Language-Based Societies of Mind discusses learning Economies Of Minds which might have a spectacular future. He links the work to his earlier learning to think framework from 2015 and 2018 on reinforcement learning. The paper draws from Minsky's society of mind concept and explores how diverse societies of large multimodal neural networks solve problems by interviewing each other.

    Combined views

    3K

    1 Source, first seen 28d ago

    Combined views

    3K

    1 Source, first seen 28d ago

    14 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    14 likes
    2 comments
    10 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 comments
    10 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @SchmidhuberAIOur award-winning 2023 paper on "Mindstorms in Natural Language-Based Societies of Mind" (NLSOMs) discusses learning "Economies Of Minds" (EOMs) which might have a spectacular future ahead of them. Quote: The original “learning to think” framework (Schmidhuber, 2015, 2018) addresses Reinforcement Learning (RL), the most general type of learning (it’s trivial to show that any problem of computer science can be formulated as an RL problem). A neural controller C learns to maximize cumulative reward while interacting with an environment. To accelerate reward intake, C can learn to interview, in a very general way, another neural net (NN) called M, which has itself learned in a segregated training phase to encode/predict all kinds of data, e.g., videos. In the present paper, however, we have so far considered only zero-shot learning. So let us now focus on the general case where at least some NLSOM members use RL techniques to improve their reward intakes. How should one assign credit to NLSOM modules that helped to set the stage for later successes of other NLSOM members? A standard way for this uses policy gradients for LSTM networks (Wierstra et al., 2010) to train (parts of) NLSOM members to maximize their reward (just like RL is currently used to encourage LLMs to provide inoffensive answers to nasty questions [43][65]). However, other methods for assigning credit exist. As early as the 1980s, the local learning mechanism of hidden units in biological systems inspired an RL economy called the Neural Bucket Brigade (NBB, Schmidhuber, 1989) [35] for neural networks with fixed topologies [36]. There, competing neurons that are active in rare moments of delayed reward translate the reward into “weight substance” to reinforce their current weights. Furthermore, they pay weight substance to “hidden” neurons that earlier helped to trigger them. The latter, in turn, pay their predecessors, and so on, such that long chains of credit assignment become possible. This work was inspired by even earlier work on non-neural learning economies such as the bucket brigade (Holland, 1985) [66] (see also later work [67][68]). How can we go beyond such simple hardwired market mechanisms in the context of NLSOMs? A central aspect of our NLSOMs is that they are human-understandable since their members heavily communicate through human-invented language. Let’s now generalize this and encode rewards by another concept that most humans understand: money. Some members of an NLSOM may interact with an environment. Occasionally, the environment may pay them in the form of some currency, say, USD. Let us consider an NLSOM member called M. In the beginning, M is endowed with a certain amount of USD. However, M must also regularly pay rent/taxes/other bills to its NLSOM and other relevant players in the environment. If M goes bankrupt, it disappears from the NLSOM, which we now call an Economy of Minds (EOM), to reflect its sense of business. M may offer other EOM members money in exchange for certain services (e.g., providing answers to questions or making a robot act in some way). Some EOM member N may accept an offer, deliver the service to M, and get paid by M. The corresponding natural language contract between M and N must pass a test of validity and enforceability, e.g., according to EU law. This requires some legal authority, possibly an LLM (at least one LLM has already passed a legal bar exam [69][70]), who judges whether a proposed contract is legally binding. In case of disputes, a similar central executive authority will have to decide who owes how many USD to whom. Wealthy NLSOM members may spawn kids (e.g., copies or variants of themselves) and endow them with a fraction of their own wealth, always in line with the basic principles of credit conservation. An intriguing aspect of such LLM-based EOMs is that they can easily be merged with other EOMs or inserted into - following refinement under simulations - existing human-centred economies and their marketplaces from Wall Street to Tokyo. Since algorithmic trading is an old hat, many market participants might not even notice the nature of the new players. Note that different EOMs (and NLSOMs in general) may partially overlap: the same agent may be a member of several different EOMs. EOMs (and their members) may cooperate and compete, just like corporations (and their constituents) do. To maximize their payoffs, EOMs and their parts may serve many different customers. Certain rules will have to be obeyed to prevent conflicts of interest, e.g., members of some EOM should not work as spies for other EOMs. Generally speaking, human societies offer much inspiration for setting up complex EOMs (and other NLSOMs), e.g., through a separation of powers between legislature, executive, and judiciary. Today LLMs are already powerful enough to set up and evaluate NL contracts between different parties [69]. Some members of an EOM may be LLMs acting as police officers, prosecutors, counsels for defendants, and so on, offering their services for money. The EOM perspective opens a rich set of research questions whose answers, in turn, may offer new insights into fundamental aspects of the economic and social sciences... From the 2023 arXiv preprint: https://arxiv.org/abs/2305.17066

    1 Source

    @SchmidhuberAIOur award-winning 2023 paper on "Mindstorms in Natural Language-Based Societies of Mind" (NLSOMs) discusses learning "Economies Of Minds" (EOMs) which might have a spectacular future ahead of them. Quote: The original “learning to think” framework (Schmidhuber, 2015, 2018) addresses Reinforcement Learning (RL), the most general type of learning (it’s trivial to show that any problem of computer science can be formulated as an RL problem). A neural controller C learns to maximize cumulative reward while interacting with an environment. To accelerate reward intake, C can learn to interview, in a very general way, another neural net (NN) called M, which has itself learned in a segregated training phase to encode/predict all kinds of data, e.g., videos. In the present paper, however, we have so far considered only zero-shot learning. So let us now focus on the general case where at least some NLSOM members use RL techniques to improve their reward intakes. How should one assign credit to NLSOM modules that helped to set the stage for later successes of other NLSOM members? A standard way for this uses policy gradients for LSTM networks (Wierstra et al., 2010) to train (parts of) NLSOM members to maximize their reward (just like RL is currently used to encourage LLMs to provide inoffensive answers to nasty questions [43][65]). However, other methods for assigning credit exist. As early as the 1980s, the local learning mechanism of hidden units in biological systems inspired an RL economy called the Neural Bucket Brigade (NBB, Schmidhuber, 1989) [35] for neural networks with fixed topologies [36]. There, competing neurons that are active in rare moments of delayed reward translate the reward into “weight substance” to reinforce their current weights. Furthermore, they pay weight substance to “hidden” neurons that earlier helped to trigger them. The latter, in turn, pay their predecessors, and so on, such that long chains of credit assignment become possible. This work was inspired by even earlier work on non-neural learning economies such as the bucket brigade (Holland, 1985) [66] (see also later work [67][68]). How can we go beyond such simple hardwired market mechanisms in the context of NLSOMs? A central aspect of our NLSOMs is that they are human-understandable since their members heavily communicate through human-invented language. Let’s now generalize this and encode rewards by another concept that most humans understand: money. Some members of an NLSOM may interact with an environment. Occasionally, the environment may pay them in the form of some currency, say, USD. Let us consider an NLSOM member called M. In the beginning, M is endowed with a certain amount of USD. However, M must also regularly pay rent/taxes/other bills to its NLSOM and other relevant players in the environment. If M goes bankrupt, it disappears from the NLSOM, which we now call an Economy of Minds (EOM), to reflect its sense of business. M may offer other EOM members money in exchange for certain services (e.g., providing answers to questions or making a robot act in some way). Some EOM member N may accept an offer, deliver the service to M, and get paid by M. The corresponding natural language contract between M and N must pass a test of validity and enforceability, e.g., according to EU law. This requires some legal authority, possibly an LLM (at least one LLM has already passed a legal bar exam [69][70]), who judges whether a proposed contract is legally binding. In case of disputes, a similar central executive authority will have to decide who owes how many USD to whom. Wealthy NLSOM members may spawn kids (e.g., copies or variants of themselves) and endow them with a fraction of their own wealth, always in line with the basic principles of credit conservation. An intriguing aspect of such LLM-based EOMs is that they can easily be merged with other EOMs or inserted into - following refinement under simulations - existing human-centred economies and their marketplaces from Wall Street to Tokyo. Since algorithmic trading is an old hat, many market participants might not even notice the nature of the new players. Note that different EOMs (and NLSOMs in general) may partially overlap: the same agent may be a member of several different EOMs. EOMs (and their members) may cooperate and compete, just like corporations (and their constituents) do. To maximize their payoffs, EOMs and their parts may serve many different customers. Certain rules will have to be obeyed to prevent conflicts of interest, e.g., members of some EOM should not work as spies for other EOMs. Generally speaking, human societies offer much inspiration for setting up complex EOMs (and other NLSOMs), e.g., through a separation of powers between legislature, executive, and judiciary. Today LLMs are already powerful enough to set up and evaluate NL contracts between different parties [69]. Some members of an EOM may be LLMs acting as police officers, prosecutors, counsels for defendants, and so on, offering their services for money. The EOM perspective opens a rich set of research questions whose answers, in turn, may offer new insights into fundamental aspects of the economic and social sciences... From the 2023 arXiv preprint: https://arxiv.org/abs/2305.17066