• mechoman444
    link
    fedilink
    English
    arrow-up
    1
    ·
    13 days ago

    I will admit that it is an interesting argument. As far as I can tell from some toilet-bowl research, that appears to be the tack some of the lawsuits are taking.

    My counter would be this: if someone reads ten thousand books and then writes their own story, they are not committing copyright infringement.

    In the same sense, someone who paints in the style of Michelangelo or Studio Ghibli does not pay royalties to either of those entities.

    It is, in my opinion, a weak argument, even if it is a legally valid one.

    • balsoft@lemmy.ml
      link
      fedilink
      English
      arrow-up
      1
      ·
      edit-2
      12 days ago

      someone who paints in the style of Michelangelo or Studio Ghibli does not pay royalties to either of those entities

      I think it’s actually not entirely true. There is a thing called “trade dress” (part of the trademark law), which can protect certain design or “look&feel” elements of any products a company produces; the requirements for this are a bit strict, but if the “style” of your painting is sufficiently similar to a Ghibli animation to cause confusion for customers (e.g. someone may reasonably think that the painting is by Studio Ghibli), there is a possibility that it’s a trademark violation. But this is also beside the point.

      if someone reads ten thousand books and then writes their own story, they are not committing copyright infringement.

      There is a legal distinction between human learning and LLM training.

      The neural connections in someone’s brain formed by reading a book are not considered to be a derivative work, because they are not a “work of authorship” as they are not “fixed in any tangible medium … from which it can be perceived, reproduced, or otherwise communicated", and they are not “sufficiently permanent or stable to permit it to be perceived, reproduced, or otherwise communicated for a period of more than transitory duration".

      LLM weights meanwhile totally fit the definition of a “work”, stored in the medium of a digital file (fixation) and produced by humans through a computational process (human authorship), making it a derivative work of the training material by definition.

      Once again, I agree that this is actually unfair, but it is the consequence of the law as written, because the law itself is unfair.

      • mechoman444
        link
        fedilink
        English
        arrow-up
        1
        ·
        12 days ago

        The conclusion doesn’t follow from the premises. Being fixed in a tangible medium is a prerequisite for copyright protection, not a test for derivative works. To show that LLM weights are derivative, you would have to demonstrate that they recast, transform, or adapt the protected expression of specific copyrighted works. Whether neural-network weights satisfy that standard is precisely the legal issue currently before the courts, so it is not “by definition” settled. The distinction between human neurons and digital weights addresses fixation, but it does not establish that the weights themselves contain copyrightable expression.

        More importantly, both humans and LLMs generate novel responses by drawing on what they have previously learned. A human uses patterns encoded in neural connections to formulate thoughts, while an LLM uses statistical patterns encoded in its weights to generate text. In both cases, the underlying information influences the output rather than being reproduced verbatim. The distinction between biological neurons and machine weights does not, by itself, answer the copyright question. The relevant legal issue is whether the resulting output or the weights themselves contain protectable expression from the original works, not whether the learning system is biological or computational.

        Your final sentence also makes an unsupported leap: “…making it a derivative work of the training material by definition.” Nothing in the statutory definition of a derivative work says that any fixed artifact produced after analyzing copyrighted material is automatically derivative. That conclusion is asserted rather than demonstrated. A derivative work must recast, transform, or adapt the protected expression of an existing work. Simply being created through exposure to copyrighted material is not sufficient.

        As an aside, I agree that trade dress can, in limited circumstances, protect a company’s distinctive visual identity under trademark law. However, that is a separate area of law from copyright and does not materially affect the point I was making. My example concerned copyright royalties for learning and creating in a similar style, not trademark claims based on consumer confusion.

        Moreover, I think I’ll end the discussion here. You and I are both unqualified to truly determine or judge whether this is copyright infringement. We’ll see what the courts ultimately decide.

        On a personal level, I honestly couldn’t care less how it turns out. AI and LLMs don’t impact my life in any meaningful way, so I don’t have a personal stake in the outcome.

        My main concern is that many people don’t really understand copyright law itself. It’s not as simple as saying, “I made this, therefore it’s protected.” That’s simply not how copyright works.