Applied AI · advanced · Session 12
Quiz and review — Q/K/V attention: from projections to causal output
← Back to courseFrançaisMarkdown source

Quiz with answers — Q/K/V attention: from projections to causal output

Answer all eight questions, then check the score. Open only the explanations needed for remediation.

1. Solve the worked-case variant: Increase q’s first component from 2 to 2.4 without changing the keys. Recompute both scores divided by √4 and the attention weights.

Show answer

D. Scaled scores move from [1.5,1] to [1.7,1]. Softmax weights move from about [0.622,0.378] to [0.668,0.332]: the Maya key receives more mass without becoming certain. The correct answer executes the requested change and gives a checkable result; the other texts do not close this calculation or trace.

2. Which causal order correctly connects the first three stages of “Q/K/V attention: from projections to causal output”?

Show answer

A. Three learned projections → Query-key compatibility → Scaling The chain follows the taught progression; reversing stages consumes a representation or state before it is produced.

3. If “Scaling” is removed, which diagnostic method is defensible?

Show answer

B. Keep the same input, predict the first output that depends on “Scaling,” then compare the before/after trace. One intervention and a prior prediction make the delta attributable to the removed mechanism.

4. Which verdict respects this session’s validity boundary?

Show answer

C. Attention maps alone do not prove a causal explanation of the model’s overall behavior. The correct answer bounds the conclusion; the others turn a local relation into a global guarantee.

5. Which evidence best matches the stated status of “Q/K/V attention: from projections to causal output”?

Show answer

D. Established mechanisms; numerical simplifications are pedagogical. Product or mechanism evidence must remain attributed and measured; availability and completion do not prove value.

6. When should a simpler baseline be preferred to “Weighted value mixture”?

Show answer

B. When a controlled test shows equivalent quality with lower memory, latency, or complexity. The choice depends on a measured trade-off on the real workload, not novelty or one isolated metric.

7. A learner gets the right result but cannot explain “Query-key compatibility.” Which remediation is most useful?

Show answer

A. Rebuild the first missing transformation, label its inputs and outputs, then test a neighboring case. The remediation targets the first causal break and then requires transfer instead of rewarding a guessed result.

8. Which submission actually demonstrates the outcome “Connect prefill, decode, and KV cache.”?

Show answer

C. A trace with starting data, transformations, observed result, boundary, and next experiment. The correct submission makes the reasoning reproducible and the verdict revisable by future measurement.

No answers checked yet.