TL;DR: In the last year, the Wikimedia Foundation has fired several union organizers, including those that worked on the Community Tech team - a team dedicated to building features for the volunteer community that edits Wikipedia.

As the Wiki Workers Union tries to get the Wikimedia Foundation to recognize their union, it is worth remembering that this is not the first time that the Foundation has worked against the community.

Wikimedia Enterprise is a betrayal of the volunteer movement community of Wikipedia editors, as the Wikimedia Foundation is providing privileged access to big tech AI companies to the Wikipedia corpus - a body of work that the Foundation does not own.

Movement volunteer communities contributed to Wikipedia under copyleft licenses - licenses that work to ensure that the work remains free (as in speech). The big tech AI companies do not license derivative works under copyleft licenses and often do not even attribute where the works came from.

This means that volunteers are working for big tech for free, and the Wikimedia Foundation is selling privileged access to that free labor.

It is against that backdrop that the current unionization struggle unfolds.

  • TheTechnician27
    link
    fedilink
    English
    arrow-up
    8
    arrow-down
    2
    ·
    5 days ago

    Hmm: “Source: ~50,000+ contributions to Wikimedia projects.”

    My actual sources for statements of fact were already listed. It’s a shorthand for “This is my personal position as an experienced editor that this is an overreaction” that you’re warping in bad-faith.

    • ivanvector@piefed.ca
      link
      fedilink
      English
      arrow-up
      4
      arrow-down
      1
      ·
      5 days ago

      The volunteers’ work is in no way being stolen. All contributors agree to release their content in perpetuity under CC-BY-SA (or similar licences that preceded it), meaning that all content is free for anyone to use, including commercial uses and derivatives, with the only restrictions being requiring attribution (credit the original contributor) and releasing under an equivalent license. The WMF can’t restrict access to any use complying with that license. Meaning that if AI companies want to scrape Wikipedia’s content for a compliant use, they will and they are fully legally permitted to.

      The Disney example is not apt: Disney produces content to make money, and reserves all rights to that content. It is not intended to be free nor to make money for third parties (absent a licensing agreement), and they were right to sue when AI companies misused their content. Wikipedia’s goals of creating a large collection of quality information and making it available to everyone for free are far from the same. The WMF does not own the content anyway, only the servers where the content resides.

      The WMF negotiating paid privileged access to data streams for these large clients is a win for everyone. The purpose of Wikipedia is to disseminate information, not gatekeep it, and the WMF has literally no right to decide who can access it and who cannot. AI agents’ ridiculous server demands threatened to impede access for everyone else, and the deals they have made sidestep that impending problem while also generating some revenue for the WMF.

      The Wikipedia community (its editors) could decide not to allow its content to be used by AI, but it has not. It would be very legally complicated, anyway, given the content’s license.

      • yoasif@fedia.ioOP
        link
        fedilink
        arrow-up
        2
        arrow-down
        1
        ·
        5 days ago

        All contributors agree to release their content in perpetuity under CC-BY-SA (or similar licences that preceded it), meaning that all content is free for anyone to use, including commercial uses and derivatives, with the only restrictions being requiring attribution (credit the original contributor) and releasing under an equivalent license. The WMF can’t restrict access to any use complying with that license.

        The Wikipedia community (its editors) could decide not to allow its content to be used by AI, but it has not. It would be very legally complicated, anyway, given the content’s license.

        Why do you think the license allows for big tech to produce derivative works that are not licensed CC-BY-SA?

        The WMF negotiating paid privileged access to data streams for these large clients is a win for everyone. The purpose of Wikipedia is to disseminate information, not gatekeep it, and the WMF has literally no right to decide who can access it and who cannot.

        If WMF has no right to decide who can access it and who cannot, how can they sell privileged access to who can access it? Frankly, that assertion fails on its face.

    • yoasif@fedia.ioOP
      link
      fedilink
      arrow-up
      2
      arrow-down
      1
      ·
      4 days ago

      My actual sources for statements of fact were already listed. It’s a shorthand for “This is my personal position as an experienced editor that this is an overreaction” that you’re warping in bad-faith.

      I have no idea why you think I am “warping” my reaction in bad faith - my reaction is based on the license text and what is written in the post. I also think it is ironic that you ask me to extend you grace in accepting your shorthand, and you clearly don’t bother to accept mine (that your statement was a reference to your own authority as an experienced editor).

      Material published to Wikimedia is CC BY-SA 4.0; thus, those current Enterprise customers have every right to use the material basically however they see fit regardless of Enterprise.

      You don’t actually tell us why this is the case - I argue that the companies are violating the license by not licensing their derivative works reciprocally - you don’t even bother to respond to that and just posit that they have “every right to use the material however they see fit”. Do you really believe that? Are the CC-BY-SA and GFDL licenses just completely worthless?

      Wikimedia isn’t selling access to the material, because it’s literally nobody’s to sell; they’re selling access to stream the data on their servers which they host.

      I’m not sure how much that matters. Would Warner Brothers not have an issue with me “streaming” access to their movies via my home server for payment?

      I also don’t know how much pointing out that this access has been granted for other users matters - the current CEO very clearly states that this was built to support the scraping use cases; Wikimedia knows that the scrapers are violating the licenses with every derivative work created not licensed reciprocally, and designed the feature to make that happen faster.

      Enabling non-violating use cases don’t erase the violating ones.