quantization, local-llms, small-models, ai-development
Ninety-Seven to Two: What Quantization Does to a Small Model
We shipped a 1.5B router that produced valid JSON 97% of the time. One quantization step took it to two. Here is the day we spent finding out why — and why the rules everyone repeats about quantization were written for models ten times the size. It started with a