Does a smaller model file guarantee lower latency?
Show answer
No. Kernel support, dequantization, workload, and hardware can make a smaller artifact slower.
Does weight quantization automatically shrink KV cache?
Show answer
No. KV-cache precision and weight precision are separate choices.
What evidence closes this lesson?
Show answer
A measured result with assumptions, exact revisions, and a quality check.