Big day for people who use AI locally. According to benchmarks this is a big step forward to free, small LLMs.

  • just another dev
    link
    fedilink
    English
    282 months ago

    128k token context is pretty sweet. Mistral nemo also just launched with a similar context. Good times.

      • @brucethemoose
        link
        English
        7
        edit-2
        2 months ago

        At long context (close to the full 128K), Nemo is way better than llama 8B in my testing.

        Turns out they are both very sensitive to quantization though.

        TBH I didn’t know people here were running LLMs. Seems like most of Lemmy is very broadly anti AI?

        • @ObsidianZed
          link
          English
          82 months ago

          My impression is the general consensus is we don’t want huge corporations stealing data to train their AI models only to turn around and cram it down our throats anywhere they can with increasingly negative experiences. That being said, while I would generally agree with that, I still find it interesting and especially if I can host it myself.

        • just another dev
          link
          fedilink
          English
          72 months ago

          Yeah, there’s a massive negative circlejerk going on, but mostly with parroted arguments. Being able to locally run a model with this kind of context is huge. Can’t wait for the finetunes that will result from this (*cough* NeverSleep’s *-maid models come to mind).

          • @brucethemoose
            link
            English
            2
            edit-2
            2 months ago

            I am looking into doing it on the 12B for myself, not so much for RP but novel style prose.

            I am thinking literature + a fanfic dump as a dataset?

            • just another dev
              link
              fedilink
              English
              12 months ago

              Ah, that’s a wonderful use case. One of my favourite models has a storytelling lora applied to it, maybe that would be useful to you too?

              At any rate, if you’d end up publishing your model, I’d love to hear about it.

                • just another dev
                  link
                  fedilink
                  English
                  12 months ago

                  Oof - not on my 12gb 3060 it doesn’t :/ Even at 48k context and the Q4_K quantization, it’s ollama its doing a lot of offloading to the cpu. What kind of hardware are you running it on?

        • Bilb!
          link
          fedilink
          English
          52 months ago

          If forced to characterize the attitude of lemmy towards LLM/“AI,” I’d say people here are broadly interested in the tech but critical of the way it’s often used.

          • @General_Effort
            link
            English
            52 months ago

            If by interested you mean willing to bullshit… Talking about AI here is like talking about evolution at bible camp in the deep south.

          • @brucethemoose
            link
            English
            2
            edit-2
            2 months ago

            I dunno, with image models specifically it seems like they’re the devil because of the datasets they’re trained on, killing artists, and… that’s that. And LLMs to a lesser extent. There’s truth to all that, but there’s also a lot more.

            I think most people don’t realize how much of an inflection point local running vs. corporate hosting could be, which is especially ironic on Lemmy.

      • just another dev
        link
        fedilink
        English
        32 months ago

        I haven’t given it a very thorough testing, and I’m by no means an expert, but from the few prompts I’ve ran so far, I’d have to hand it to Nemo concerning quality.

        Using openrouter.ai, I’ve also given llama3.1 405B a shot, and that seems to be at least on par with (if not better than) Claude 3.5 Sonnet, whilst being a bit cheaper as well.

        • @brucethemoose
          link
          English
          22 months ago

          Llama 70B is probably where its at, if you go the API route. It’s distilled from 405B, and its benchmarks are pretty close.