Executives working on AI at Microsoft and OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing theft of unprecedented proportions,” and the “largest theft of labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire web.”

Those statements and a series of other mask-off moments feature heavily in an unredacted court filing that was unsealed Thursday in the behemoth New York Times vs OpenAI copyright lawsuit that has been winding its way through the court system for years. In a filing asking for summary judgment (basically, a filing with the court asking it to rule), lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained either sealed or redacted at the request of Microsoft and OpenAI.

It’s easy to see why the AI companies wanted to hide this from the public. The statements, taken together, are some of the most damning indictments of the ways LLMs were trained, how they worked, and the immediate threat they pose to human labor. It is a reminder that even as AI becomes more powerful and companies try to shift the narrative to the supposed existential risk of “superintelligent” AI, the tools they have already built were created by stealing from human creativity and labor and are by definition existential threats to the human labor market.


  • CosmoNova@lemmy.world
    link
    fedilink
    English
    arrow-up
    23
    ·
    2 hours ago

    They knew what they were doing every step of the way. They are criminals that need to be disarmed and locked away. And we need to create a new Internet from scratch somehow thanks to these donkeys.

  • Grandwolf319@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    17
    ·
    3 hours ago

    “Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the document said.

    So basically the ouroboros

  • kablez@lemmy.world
    link
    fedilink
    English
    arrow-up
    29
    ·
    4 hours ago

    Gonna get real weird soon when they run out of rich new training data and they begin to consume their own shit. When that happens their entire model will collapse and if the bubble hasn’t popped already that may be what causes it.

    • MalReynolds@slrpnk.net
      link
      fedilink
      English
      arrow-up
      5
      ·
      1 hour ago

      Pretty sure they mostly use the pre-AI internet (that they scraped and kept) and synthetic data currently. Probably trying (and failing so far or we’d have heard) to adapt to using video as training material at the moment, but developments there will likely apply to robotics at some point. Here’s hoping the current chuds have crashed and burned before then and that some sanity has taken over from unfettered capitalist oligarchs dreams of computer slavery.

    • RepleteLocum@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      19
      ·
      4 hours ago

      They’re already doing it. They call it distilling when they take it from another llm. Pretty sure most content was already stolen in the early days and they now rely on distillation and stealing new content.

      • kablez@lemmy.world
        link
        fedilink
        English
        arrow-up
        8
        ·
        3 hours ago

        So it’s like if everyone combined the backwash from their water bottles into a new drink…

        Sounds super appealing!

      • LifeInMultipleChoice@lemmy.world
        link
        fedilink
        English
        arrow-up
        5
        ·
        4 hours ago

        That’s my thought. So long as someone inputs anything new to the internet they will be able to scrape it and sell it as their product.

    • Thorry@feddit.org
      link
      fedilink
      English
      arrow-up
      2
      arrow-down
      1
      ·
      2 hours ago

      Why do you think there has been such an emphasis on hacking with LLMs lately (especially by OpenAI). They figured out all those vulnerability databases were an excellent source for training material. In the past they scraped those, but just for general language training. Now they’ve specifically trained the models on the information within. Some team figured out how to use that data to train a model and have testing scenarios automated so they could write a good reward function. It wasn’t that they figured out the models are good at hacking, they ran out of content and found a new source of good data.

      With all the books they’ve been scanning I wonder if the next thing is going to be a writing assistant or editor or something like that. Even though writing good books is an art form and the actual writing down of the words is the easiest part (still not easy tho).

      These companies are starving for content and they’ve not just poisoned buy absolutely destroyed the content well that is the internet. Given they were already hitting diminishing returns hard, it doesn’t matter too much to them probably. But more compute and storage has also been hitting diminishing returns hard and customers are complaining about the cost. So they are getting a bit desperate on how to improve these things at all.

  • Serinus@lemmy.world
    link
    fedilink
    English
    arrow-up
    1
    ·
    1 hour ago

    They’re more of a threat to life than labor. Think about how much easier it is to kill you than to do your job.

  • velma@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    115
    ·
    6 hours ago

    “Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the document said.

    Microsoft executives, including CEO Satya Nadella, testified under oath that after ripping content from the New York Times and other news sites, clicks to those news sites fully cratered, falling by more than 90 percent on Bing.

    Documents obtained during the court proceedings found that OpenAI created “a hack to get around nytimes paywall,” to which OpenAI cofounder Greg Brockman said “ah, nice.” Microsoft executive Brent Hecht wrote that LLMs steal content “without ways of distributing economic value down the supply chain, [which] necessarily threatens the economic stability of those who create the content.”

    Fuck these guys.

    • TeaWithDani@lemmy.world
      link
      fedilink
      English
      arrow-up
      11
      ·
      3 hours ago

      That’s kind of the funniest part in fact: these companies are destroying their own viable business segments. Bing was a huge growth driver for Microsoft. Less clicks is bad for them. They make more money on Bing ads than they do on LLMs. Same with Google.

      Reddit is getting crushed atm after it sold access to its data to train models. Chat bots make visiting these websites pointless, without replacing that traffic with anything they can meaningfully monetize.

      The more popular Gemini is, the less money Google will make. The market has already shown how much people are willing to pay for Ai subscriptions, and it isn’t all that much. None of these companies have found a way to make ads viable in LLMs either. They are beyond self sabotaging themselves at this point. It really is a doom loop.

    • tangeli@piefed.social
      link
      fedilink
      English
      arrow-up
      22
      ·
      5 hours ago

      That’s because, thus far, they get away with choosing not to distribute any of their trillions of dollars to the suppliers of the information they consume - money has only gone to the suppliers of hardware and power, and to influencing politicians and rewarding investors. That’s their choice, and they should not be allowed to continue to make that choice. Good luck to the NYT.

    • belochka@lemmy.world
      link
      fedilink
      English
      arrow-up
      7
      ·
      4 hours ago

      At least they are going to pick one between copyright and using all the Web as accumulated material for their answering machine.

  • Arancello@aussie.zone
    link
    fedilink
    English
    arrow-up
    54
    arrow-down
    2
    ·
    6 hours ago

    Relax guys, the very stable high IQ president of the united states will protect you. No meed for guardrails or regulations.

    • AmyAye@nord.pub
      link
      fedilink
      English
      arrow-up
      32
      ·
      5 hours ago

      God the guardrails things. These stupid companies are all hyping up “We need to slow down, we need guard rails!”

      Ok.

      No one is fucking stopping you. Just… Slow yourself down, guardrail yourself.

      Oh wait, it’s just an excuse to create regulatory capture.

      • Fluke@feddit.uk
        link
        fedilink
        English
        arrow-up
        2
        ·
        39 minutes ago

        It’s the theftbot manufacturers trying to have excuses in the public perception for why all their claims about their product don’t ever happen.

        “We had to slow down for safety. All those things we promised are coming, just one more round of funding bro.”

    • Archangel1313@lemmy.ca
      link
      fedilink
      English
      arrow-up
      11
      ·
      6 hours ago

      That clown is legalizing all kinds of white collar crime, already. This kind of shit is nothing to him.

  • gravitas_deficiency@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    43
    ·
    6 hours ago

    Microsoft, in a policy document, wrote that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained […] LLMs are a product that destroys its supply chain.”

    If that’s not the very definition of a categorically unsustainable business model, I don’t know what the fuck is.

      • belochka@lemmy.world
        link
        fedilink
        English
        arrow-up
        3
        ·
        4 hours ago

        That’s not quite the thing. They know that people in various societies in the course of history could create artifacts getting much less in return than they’d get today in a market without LLMs. They just want those people to become crops in their fields, so to say. To reap and not sow.

        There’s the obvious problem with this, that there still are alternative business models of closed circles and ordering artifacts, not accepting offers to buy them. And, of course, that either it’s IP violation of ultimate proportions, or if it’s not, then they are going to have too much competition to get anything out of it anyway.

        So it’s either oligopoly with regulatory capture, or too much competition to make this profitable.

        They apparently want to make some theft legal and some not, a bit like copyright protection entities in ex-Soviet states, which are usually controlled by former pirates legalized (with ex-pirate electronic libraries which are now legitimate stores, except I’m confident they didn’t get consent of most authors whose books are being sold, similarly with movies and such), hunting their “honest pirate” competition.

        So - no, at the point this is mainstream, the only regulation of this should be aimed at enforcing existing laws against them. Then one can think of some, but before existing laws are being firmly enforced - no.

        Except enforcing existing laws can mean fines bigger than their capitalization and some life sentences for all chief people involved, and these are all “too big to fail” companies, so I don’t even know.

        If these companies are, figuratively, sentenced to death for doing this, then it’ll be a process of the century. It’ll define future. A bit like with Standard Oil and United Fruit Company. Or so I think.

        Except it might go differently, in their favor, then it won’t be a very nice future.

        • Valmond@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          1
          ·
          2 hours ago

          Or they will pop all by themselves?

          I mean there is not very much new data to train on, and only so much refining you can do (and you don’t need to be one of the big endebted companies to do that).

  • calcopiritus@lemmy.world
    link
    fedilink
    English
    arrow-up
    10
    ·
    5 hours ago

    Anyone with 2 braincells saw this from day 1. Unfortunately the timeline lined up with USAians electing a president with less than 1 braincell. So there was no one to stop them when it was time to stop them.

    They destroyed the web forever for the entire world, but owned the libs for a couple years, so it’s worth it I guess.

  • Arcanepotato@crazypeople.online
    link
    fedilink
    English
    arrow-up
    6
    ·
    6 hours ago

    Like, intellectual property doesn’t matter to me at all with respect to anything I may have produced that has been scraped. I am, however, really not liking the idea that parodies on highly niche technical subjects that my me or my friends might have been created might have been added to their “collective knowledge” or whatever and it will be used to synthesize answers in the future.

    Not because it’s anything so powerful (this is not a flex, I swear) but that’s got to be an infinitesimal fraction of all the absolute trash that’s been fed to the machines without scrutiny.

    I hate that