Breaking the 1.58-bit Barrier for Ternary LLMs

(arxiv.org)

93 points | by matt_d 2 hours ago

7 comments

  • infogulch 3 minutes ago
    So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.

    If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.

  • om8 41 minutes ago
    Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
    • janalsncm 8 minutes ago
      PTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU.

      If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.

    • om8 39 minutes ago
      If you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS
  • plqbfbv 42 minutes ago
    Very interesting, I was just exploring this to hopefully fit one of the latest quantized models in 16GB of VRAM.
  • NooneAtAll3 42 minutes ago
    This is the only time "1.58 bit" phrase makes more sense than "1 trit"

    Who knew that if you actually look at information entropy you can pack stuff better!

  • Kevcmk 1 hour ago
    Woah. Good science.
  • kadushka 21 minutes ago
    [dead]
  • kadushka 59 minutes ago
    [flagged]