4-bit quantisation
AI4-bit quantisation: Compressing model weights to just four bits each, roughly a quarter of the usual size, so large models fit on smaller or cheaper hardware.
4-bit quantisation: Compressing model weights to just four bits each, roughly a quarter of the usual size, so large models fit on smaller or cheaper hardware.