← Lcoalhost
How 8-bit quantization shrinks an LLM to a quarter of its size — and why a single outlier weight can quietly ruin it
Source :
DEV · #llm
See it live in context on Lcoalhost →