Systems session 2/3 · Exercises

CPU vs GPU and quantization

  1. Estimate weights for a 7B model at 16, 8, and 4 bits.
  2. Choose CPU and GPU benchmark cases at concurrency 1 and 8.
  3. Define a quality rollback threshold before testing.
Show answer

No. Kernel support, dequantization, workload, and hardware can make a smaller artifact slower. No. KV-cache precision and weight precision are separate choices.