The world’s biggest AI companies are buying antiquarian books en masse so they can be scanned to train their large language models (LLMs) before being destroyed.
I heard someone on another social media site say that if the only physical copies are gone then it’s hard to prove copyright infringement.
So they are intentionally shreading books knowing that they can’t be sued for copyright infringement if there is no copy of the og material to prove it.
I have no clue if that’s true, but that sounds pretty far fetched imo.
As already mentioned, the books being purchased are not like, the only sole remaining copies of the work. Many of the authors are probably still alive, and there’s likely thousands or millions of each book distributed all over the country of origin, if not the world, in the hands of individuals, bookshops, libraries, archives, etc.
A judge said it was fair use when they purchased and solely used the books for AI training without redistributing them or a copy of them afterwards, so they’re just doing that in order to not get sued again. Not much more to it.
Aren’t some of the books considered rare ?
I mean that’s why people are mad.
And yeah I agree they already use stolen works.
I’m just saying I heard someone say that.
I’m curious, myself, if it holds true or not. I don’t know enough about copyrights.
But I do know Google has made quite a lot of effort to scan books. Ive found stuff on Google books that is legit 100 years old in German research on optics. (I study depth perception).
And I was really surprised they had a photocopy of it.
But maybe Google wanted a hefty price for access. Maybe a subscription.
Aren’t some of the books considered rare ? I mean that’s why people are mad.
I looked it up further, and while I’ve seen some people claim they’re using rare books, as far as I can tell that’s just based on a single bookseller saying that some of the books he distributed to ISBNdb (which the AI companies are buying through) were “rare or out of print”, but he didn’t provide any details on what those books were or how rare they were exactly, and he’s also the only source I’ve seen for that entire claim, so I’m not really sure how common that would actually be if he’s literally the one single person they could find who sold rare books to them.
Regardless, I still obviously am not a fan of that happening, I’d much rather that those books could be archived, even if that meant destruction but then digitizing them in, say, the Internet Archive instead, but at the end of the day I just wanted people to know that there wasn’t exactly no reason for them doing it the way they are. It’s not just malice for the hell of it, they got a court order that said how they could do it legally, so they did.
I heard someone on another social media site say that if the only physical copies are gone then it’s hard to prove copyright infringement.
So they are intentionally shreading books knowing that they can’t be sued for copyright infringement if there is no copy of the og material to prove it.
I have no clue if that’s true, but that sounds pretty far fetched imo.
As already mentioned, the books being purchased are not like, the only sole remaining copies of the work. Many of the authors are probably still alive, and there’s likely thousands or millions of each book distributed all over the country of origin, if not the world, in the hands of individuals, bookshops, libraries, archives, etc.
A judge said it was fair use when they purchased and solely used the books for AI training without redistributing them or a copy of them afterwards, so they’re just doing that in order to not get sued again. Not much more to it.
Aren’t some of the books considered rare ? I mean that’s why people are mad.
And yeah I agree they already use stolen works.
I’m just saying I heard someone say that.
I’m curious, myself, if it holds true or not. I don’t know enough about copyrights.
But I do know Google has made quite a lot of effort to scan books. Ive found stuff on Google books that is legit 100 years old in German research on optics. (I study depth perception). And I was really surprised they had a photocopy of it.
But maybe Google wanted a hefty price for access. Maybe a subscription.
I looked it up further, and while I’ve seen some people claim they’re using rare books, as far as I can tell that’s just based on a single bookseller saying that some of the books he distributed to ISBNdb (which the AI companies are buying through) were “rare or out of print”, but he didn’t provide any details on what those books were or how rare they were exactly, and he’s also the only source I’ve seen for that entire claim, so I’m not really sure how common that would actually be if he’s literally the one single person they could find who sold rare books to them.
Regardless, I still obviously am not a fan of that happening, I’d much rather that those books could be archived, even if that meant destruction but then digitizing them in, say, the Internet Archive instead, but at the end of the day I just wanted people to know that there wasn’t exactly no reason for them doing it the way they are. It’s not just malice for the hell of it, they got a court order that said how they could do it legally, so they did.