Those claiming AI training on copyrighted works is “theft” misunderstand key aspects of copyright law and AI technology. Copyright protects specific expressions of ideas, not the ideas themselves. When AI systems ingest copyrighted works, they’re extracting general patterns and concepts - the “Bob Dylan-ness” or “Hemingway-ness” - not copying specific text or images.

This process is akin to how humans learn by reading widely and absorbing styles and techniques, rather than memorizing and reproducing exact passages. The AI discards the original text, keeping only abstract representations in “vector space”. When generating new content, the AI isn’t recreating copyrighted works, but producing new expressions inspired by the concepts it’s learned.

This is fundamentally different from copying a book or song. It’s more like the long-standing artistic tradition of being influenced by others’ work. The law has always recognized that ideas themselves can’t be owned - only particular expressions of them.

Moreover, there’s precedent for this kind of use being considered “transformative” and thus fair use. The Google Books project, which scanned millions of books to create a searchable index, was ruled legal despite protests from authors and publishers. AI training is arguably even more transformative.

While it’s understandable that creators feel uneasy about this new technology, labeling it “theft” is both legally and technically inaccurate. We may need new ways to support and compensate creators in the AI age, but that doesn’t make the current use of copyrighted works for AI training illegal or unethical.

For those interested, this argument is nicely laid out by Damien Riehl in FLOSS Weekly episode 744. https://twit.tv/shows/floss-weekly/episodes/744

  • @General_Effort
    link
    English
    89 days ago

    Let’s engage in a little fantasy. Someone invents a magic machine that is able to duplicate apartments, condos, houses, … You want to live in New York? You can copy yourself a penthouse overlooking the Central Park for just a few cents. It’s magic. You don’t need space. It’s all in a pocket dimension like the Tardis or whatever. Awesome, right? Of course, not everyone would like that. The owner of that penthouse, for one. Their multi-million dollar investment is suddenly almost worthless. They would certainly demand that you must not copy their property without consent. And so would a lot of people. And what about the poor construction workers, ask the owners of constructions companies? And who will pay to have any new house built?

    So in this fantasy story, the government goes and bans the magic copy machine. Taxes are raised to create a big new police bureau to monitor the country and to make sure that no one use such a machine without a license.

    That’s turned from magical wish fulfillment into a dystopian story. A society that rejects living in a rent-free wonderland but instead chooses to make itself poor. People work to ensure poverty, not to create wealth.

    You get that I’m talking about data, information, knowledge. The first magic machine was the printing press. Now we have computers and the Internet.

    I’m not talking about a utopian vision here. Facts, scientific theories, mathematical theorems, … All such is free for all. Inventors can get patents, but only for 20 years and only if they publish them. They can keep their invention secret and take their chances. But if they want a government enforced monopoly, they must publish their inventions so that others may learn from it.

    In the US, that’s how the Constitution demands it. The copyright clause: [The United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries.

    Cutting down on Fair Use makes everyone poorer and only a very few, very rich people richer. Have you ever thought about where the money goes if AI training requires a license?

    For example, to Reddit, because Reddit has rights to all those posts. So do Facebook and Xitter. Of course, there’s also old money, like the NYT or Getty. The NYT has the rights to all their old issue about a century back. If AI training requires a license, they can sell all their old newspapers again. That’s pure profit. Do you think they will their employees raises out of the pure goodness of their heart if they win their lawsuits? They have no legal or economics reason to do so. The belief that this would happen is trickle-down economics.

    • @Emerald
      link
      English
      3
      edit-2
      9 days ago

      Thanks for a comment like this. It’s interesting how everyone steps in to endorse piracy (unauthorized copying of copyrighted works), yet when a business does it for AI purposes everyone freaks out.

      • @[email protected]
        link
        fedilink
        English
        89 days ago

        Because most people pirating are doing it for their own personal entertainment while these companies are doing it to build a commercial product for sale. Pirates that sell access to their collections get a lot of negative attention, even from other people who pirate like me.

        • @General_Effort
          link
          English
          29 days ago

          Copyright is utterly corrupted. Besides, I believe it is corrosive and outright dangerous in the age of the internet. Every time you open a website or a stream or anything, that is copied to your device. In the age of the printing press, it was about what happened in a few “factories”/printing houses. Libraries were fine because they didn’t copy, but online libraries do. Now, copyright is about all our communications. Total enforcement would mean total surveillance.

          So this is not a defense of copyright. It is simply an explanation.

          Building products for sale is what US-copyright is all about. Think about the copyright clause: To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries.

          Without copyright, everything would be public domain. Everyone would be free to share any book or movie. That makes it hard to make money, to monetize your product, to recoup your investment. Copyright is supposed to be a way to enable that. It’s supposed to create an incentive to entertain you. If you have to pay for your entertainment, then someone will come along and entertain you to get your money. Piracy is an attack on that system.

          If AI companies have to buy licenses, that would not incentivize much of anything. Licensing curated datasets for AI training would be one thing, but paying for individual books or even Reddit posts makes no sense. It would just make development slower and much more expensive. That makes it an unconstitutional use of copyright.

      • @General_Effort
        link
        English
        18 days ago

        The copyright industry wants money. So, 4 legs good, 2 legs better. It’s depressing to see how easily people are led around by the nose.