• reddfugee
    link
    fedilink
    English
    arrow-up
    25
    ·
    20 days ago

    I think it’s directed at the “I don’t use commercial LLMs, I just run small local models” crowd

    • Skullgrid
      link
      fedilink
      English
      arrow-up
      10
      arrow-down
      3
      ·
      20 days ago

      yes and what books are they destroying?

      • tyler@programming.dev
        link
        fedilink
        English
        arrow-up
        18
        arrow-down
        2
        ·
        20 days ago

        Those smaller models are trained on the larger models. They take just as many resources to train and then you add on more training, so they’re literally worse than the big models.

        • Skullgrid
          link
          fedilink
          English
          arrow-up
          4
          arrow-down
          6
          ·
          20 days ago

          What if your power comes from renewables? And if you’re training private data?

          • ayyy@sh.itjust.works
            link
            fedilink
            English
            arrow-up
            4
            ·
            19 days ago

            And what if tech oligarchs weren’t evil and used this new technology to make lives better instead of concentrating capital? We don’t have that either so there’s no use speculating.

            Snark aside, I think you’re underestimating just how much data it takes to make these models work. There is no such thing as enough “private” data for training.

            • Skullgrid
              link
              fedilink
              English
              arrow-up
              3
              ·
              19 days ago

              There is no such thing as enough “private” data for training.

              what the fuck are you on about? You can build local datasets and train additional LORAs locally to have more relevant outputs, and it doesn’t go back to the creators of the original weights.

              • ayyy@sh.itjust.works
                link
                fedilink
                English
                arrow-up
                1
                ·
                19 days ago

                You are describing fine tuning, no? Which still requires a base model that was trained on a massive amount of other data.

                • Skullgrid
                  link
                  fedilink
                  English
                  arrow-up
                  1
                  ·
                  19 days ago

                  Yes and those models already exist, for free.

                  It’s not the unit 731 data ffs

          • tyler@programming.dev
            link
            fedilink
            English
            arrow-up
            3
            ·
            19 days ago

            Then they still used just as many resources? All these local models are distillations of the big models.