• @[email protected]
    link
    fedilink
    English
    8210 months ago

    Can we get a list of companies NOT doing this? I’d assume it’s going to be much shorter.

    • @[email protected]
      link
      fedilink
      English
      4110 months ago

      All these AI and machine learning companies are taking content directly from websites and ignoring robot.txt files.

      If your content is able to be crawled, even without being listed on search engines, I don’t think it really matters.

      • @T156
        link
        English
        810 months ago

        It might help proof an AI company against legal issues that might be brought about by their using the content. If they’re ever sued by Automattic, then they can just point to the deal and say that they bought the data from them. There’s much less ambiguity.

        • @[email protected]
          link
          fedilink
          English
          310 months ago

          You are correct, about the legal stuff. These companies are being sued all the time.

          Doing this deal also makes processing the data a lot easier. Being handed a big ass database would be a lot easier than crawling for content.

          What I posted was about how they operate. These companies showed time and time again that they don’t really care what data they are taking or from whom. They will even take their own AI or machine learning content and put it in their own system.