• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Impersona-env released as a proof of concept for training personalized AI assistants

    Its creator says assistants can learn preferences by repeatedly questioning a simulated version of the user.

    CLSCL
    Sarah CatanzaroSC
    Ben ShiBS
    3 Sources, ,

    TLDR

    The developer released impersona-env, a proof-of-concept reinforcement-learning environment built from user models. They say assistants can ask simulated users questions and use their feedback to refine their understanding of personal preferences. Evaluating those models across users and applications, handling long-term personal data and protecting privacy remain open questions, the developer says.

    Combined views

    5.9K

    3 Sources, first seen 2h ago

    Combined views

    5.9K

    3 Sources, first seen 2h ago

    113 likes
    2h ago
    first seen 2h ago
    113 likes
    12 comments
    37 saves
    16 reposts
    12 comments
    37 saves
    16 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    Ben Shi@BenShi34Hello my 406 twitter followers. Hopping on to share two things: 1. I’ve started a CS PhD at Stanford! For the love of the game ofc. 2. I’m releasing impersona-env, a proof of concept to train personalized assistants by building an RL environment out of user models. Your assistant learns to help you by repeated asking a simulated version of you questions, probing its internal latent thinking and feedback, and updating its priors accordingly. I think this unlocks a new dimension to personalization beyond discrete fact recall. My personal assistant model now understands nuances of my preferences that even I didn’t know about. Full blog: https://benshi34.github.io/#/blog/impersona-env It’s still a proof of concept and there’s still a ton of things to figure out in this space. It’s unclear to me how to build user model evals that generalize across users and downstream applications, how to best represent longitudinal personal data at training time // inference time, the interplay between population level and individual level preferences, the immense questions surrounding privacy and hillclimbing in decentralized data access regime… if you want to chat on anything in this area please dm me!2h
    CLS@ChengleiSiRT @BenShi34: Hello my 406 twitter followers. Hopping on to share two things: 1. I’ve started a CS PhD at Stanford! For the love of the ga…1h
    Sarah Catanzaro@sarahcat21Personalized models mean personalized RL environments1h

    3 Sources

    Ben Shi@BenShi34Hello my 406 twitter followers. Hopping on to share two things: 1. I’ve started a CS PhD at Stanford! For the love of the game ofc. 2. I’m releasing impersona-env, a proof of concept to train personalized assistants by building an RL environment out of user models. Your assistant learns to help you by repeated asking a simulated version of you questions, probing its internal latent thinking and feedback, and updating its priors accordingly. I think this unlocks a new dimension to personalization beyond discrete fact recall. My personal assistant model now understands nuances of my preferences that even I didn’t know about. Full blog: https://benshi34.github.io/#/blog/impersona-env It’s still a proof of concept and there’s still a ton of things to figure out in this space. It’s unclear to me how to build user model evals that generalize across users and downstream applications, how to best represent longitudinal personal data at training time // inference time, the interplay between population level and individual level preferences, the immense questions surrounding privacy and hillclimbing in decentralized data access regime… if you want to chat on anything in this area please dm me!2h
    CLS@ChengleiSiRT @BenShi34: Hello my 406 twitter followers. Hopping on to share two things: 1. I’ve started a CS PhD at Stanford! For the love of the ga…1h
    Sarah Catanzaro@sarahcat21Personalized models mean personalized RL environments1h