Systems session 2/3 · Quiz & review

CPU vs GPU and quantization

Does a smaller model file guarantee lower latency?

Show answer

No. Kernel support, dequantization, workload, and hardware can make a smaller artifact slower.

Does weight quantization automatically shrink KV cache?

Show answer

No. KV-cache precision and weight precision are separate choices.

What evidence closes this lesson?

Show answer

A measured result with assumptions, exact revisions, and a quality check.