PTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU.
If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.
If you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS
If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.
If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.
Who knew that if you actually look at information entropy you can pack stuff better!