papersSEP 10 04:00 UTC
Researchers scale post-training ternarisation to Qwen3-8B with 1.58-bit packed execution
An arXiv paper reports scaling an aggressive post-training ternarisation pipeline to the 8B-parameter Qwen3 model, measuring how much capability is retained after conversion to ultra-low-bit weights. The authors argue that a nominal 1.58-bit label alone does not specify the actual deployed representation or its runtime cost, so they additionally present a lossless packing scheme and a method for executing the model directly on packed weights. The work includes reproduction details so that the conversion results can be independently replicated.