• brucethemoose
    link
    fedilink
    arrow-up
    3
    ·
    5 days ago

    I use an exl3, with 4 bit MLPs but higher bit depth attention layers. And I force some custom sampling so I can lower the temperature a bit while keeping it out of loops.

    This won’t work in LM Studio though. You have to run such a thing in TabbyAPI or some other backend that supports exllamav3.

    • stuner
      link
      fedilink
      arrow-up
      2
      ·
      5 days ago

      Ok, thanks! Thaf sounds quite advanced, but I’ll have a read afterwards :)