• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    David Friedberg alleges AI labs used his chat history to train later models

    A post quotes Friedberg saying a later model, accessed through a different account, described the same novel scientific idea discussed in an earlier chat. He calls these “a handful of anecdotal experiences” but believes his team's work was used for training.

    CP
    @J
    TA
    4 Sources, ,

    TLDR

    David Friedberg, quoted in a user's post, alleges that ideas from his team's AI conversations resurfaced in a later model. He describes the experiences as anecdotal but concludes that the conversations or analyses were used for training. Friedberg calls the work his organization's intellectual property and says it had no NDA or confidentiality protections with the provider. He ties his interest in open source to that concern: he doesn't want chat logs used for training that could spread his organization's ideas to the wider market.

    Combined views

    401.4K

    4 Sources, first seen 19d ago

    Combined views

    401.4K

    4 Sources, first seen 19d ago

    1.9K likes
    19d ago
    first seen 19d ago
    1.9K likes
    118 comments
    740 saves
    209 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    118 comments
    740 saves
    209 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    @theallinpodDavid Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @Jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” @friedberg: “I have had experiences where we've asked some fairly novel scientific questions, and (the AI model) identifies it as a novel insight. It's like, ‘Oh, never thought about that, interesting, blah, blah, blah.’ And then using a different account, asking the next version (of the model) later, I've now experienced this. It's like, ‘Oh, well, you could do this,’ and it actually just describes this exact thing that we had in our chat in the previous version. Now, these are a handful of anecdotal experiences, but I know the domain that we work in, and the niche of it, and the ideation of this stuff, and the novelty of this stuff, and the lack of papers being published, and so on. So I know that there isn't some new corpus of information out there that's training the new model. So all I can say at that point is that my conversation or our analyses have been used for training.” @DavidSacks: “Okay, this does raise a really good question. What does it mean that the model is allowed to train on unidentifiable data?” Friedberg: “Well, that's my point. So it doesn't use any of my personal information, but it can use an insight derived from our chat, which it can then say is some training data that is unrelated. But the truth is, it's actually a piece of IP that's our organization’s IP, and our engagement back and forth. We don't have any NDA or confidentiality provisions or protections with them being a service provider back to us. This is why I care a lot about open source because I don't want them having my chat logs because they can use it for training to create an IP advantage that is now diffused to the rest of the market.”
    @JasonRT @theallinpod: David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @Jason: “Should they trust…
    @dnapwayDavid Friedberg reveals frontier labs are taking novel insights from his chat history and packaging them as their own "I have had experiences where we've asked some fairly novel scientific questions, and it identifies it as a novel insight, 'oh, never thought about that, might be interesting,' blah blah blah..." "And the using a different account, asking the next model version, it's like, 'oh, you could do this'. And it actually just describes this exact thing that we had in our chat in the previous version." "These are a handful of anecdotal experiences. But I know the domain that we work in, and the novelty of this stuff. So I know that there isn't some new corpus of information out there that's training the new model." "So all I can say at that point is that my conversation or our analyses have been used for training. The truth is, it's actually a piece of IP, that's our organizational IP. We don't have any NDA or confidentiality protections, with them being a service provider back to us." "This is why I care a lot about open source, because I don't want them having my chat logs, because they can use them for training to create an IP that is now diffused to the rest of the market. I find this very unfair."
    @chamathRT @dnapway: David Friedberg reveals frontier labs are taking novel insights from his chat history and packaging them as their own "I have…

    4 Sources

    @theallinpodDavid Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @Jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” @friedberg: “I have had experiences where we've asked some fairly novel scientific questions, and (the AI model) identifies it as a novel insight. It's like, ‘Oh, never thought about that, interesting, blah, blah, blah.’ And then using a different account, asking the next version (of the model) later, I've now experienced this. It's like, ‘Oh, well, you could do this,’ and it actually just describes this exact thing that we had in our chat in the previous version. Now, these are a handful of anecdotal experiences, but I know the domain that we work in, and the niche of it, and the ideation of this stuff, and the novelty of this stuff, and the lack of papers being published, and so on. So I know that there isn't some new corpus of information out there that's training the new model. So all I can say at that point is that my conversation or our analyses have been used for training.” @DavidSacks: “Okay, this does raise a really good question. What does it mean that the model is allowed to train on unidentifiable data?” Friedberg: “Well, that's my point. So it doesn't use any of my personal information, but it can use an insight derived from our chat, which it can then say is some training data that is unrelated. But the truth is, it's actually a piece of IP that's our organization’s IP, and our engagement back and forth. We don't have any NDA or confidentiality provisions or protections with them being a service provider back to us. This is why I care a lot about open source because I don't want them having my chat logs because they can use it for training to create an IP advantage that is now diffused to the rest of the market.”
    @JasonRT @theallinpod: David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @Jason: “Should they trust…
    @dnapwayDavid Friedberg reveals frontier labs are taking novel insights from his chat history and packaging them as their own "I have had experiences where we've asked some fairly novel scientific questions, and it identifies it as a novel insight, 'oh, never thought about that, might be interesting,' blah blah blah..." "And the using a different account, asking the next model version, it's like, 'oh, you could do this'. And it actually just describes this exact thing that we had in our chat in the previous version." "These are a handful of anecdotal experiences. But I know the domain that we work in, and the novelty of this stuff. So I know that there isn't some new corpus of information out there that's training the new model." "So all I can say at that point is that my conversation or our analyses have been used for training. The truth is, it's actually a piece of IP, that's our organizational IP. We don't have any NDA or confidentiality protections, with them being a service provider back to us." "This is why I care a lot about open source, because I don't want them having my chat logs, because they can use them for training to create an IP that is now diffused to the rest of the market. I find this very unfair."
    @chamathRT @dnapway: David Friedberg reveals frontier labs are taking novel insights from his chat history and packaging them as their own "I have…