# Quiz with answers — Mixture-of-Experts: routing and capacity

### Question 1

Solve the worked-case variant: Increase the first expert score from 2.1 to 2.52 and keep [1.8,0.2] for the others. Recompute softmax and check whether top-2 changes.

- A. Routing example reduced to three experts for the calculation: scores [2.1,1.8,0.2] give softmax about [0.529,0.392,0.079]. With top-2, experts 1 and 2 are active. If expert 1 has reached capacity, routing must apply the overflow policy.
- B. Activation names and exact expert counts attributed to a named model remain source-reported until primary evidence is attached.
- C. Weights move from about [0.529,0.392,0.079] to [0.631,0.307,0.062]. Experts 1 and 2 remain top-2, but expected load concentrates more on expert 1; capacity and overflow policy become more critical.
- D. The router is a small projection: h_t → one logit per expert. Worked example over 3 experts: [2.1, 1.8, 0.2] → softmax [0.529, 0.392, 0.079]. These scores are decisions learned through the global loss — not human categories.

**Answer: C.** Weights move from about [0.529,0.392,0.079] to [0.631,0.307,0.062]. Experts 1 and 2 remain top-2, but expected load concentrates more on expert 1; capacity and overflow policy become more critical. The correct answer executes the requested change and gives a checkable result; the other texts do not close this calculation or trace.

---

### Question 2

Which causal order correctly connects the first three stages of “Mixture-of-Experts: routing and capacity”?

- A. Top-k and mixture → Router → Why experts
- B. Router → Why experts → Top-k and mixture
- C. Why experts → Top-k and mixture → Router
- D. Why experts → Router → Top-k and mixture

**Answer: D.** Why experts → Router → Top-k and mixture The chain follows the taught progression; reversing stages consumes a representation or state before it is produced.

---

### Question 3

If “Top-k and mixture” is removed, which diagnostic method is defensible?

- A. Keep the same input, predict the first output that depends on “Top-k and mixture,” then compare the before/after trace.
- B. Also change the data to amplify the difference.
- C. Observe only the final output and invent the cause.
- D. Conclude that the whole system fails before measuring.

**Answer: A.** Keep the same input, predict the first output that depends on “Top-k and mixture,” then compare the before/after trace. One intervention and a prior prediction make the delta attributable to the removed mechanism.

---

### Question 4

Which verdict respects this session’s validity boundary?

- A. The mechanism guarantees accuracy, speed, and stability for every workload.
- B. Activation names and exact expert counts attributed to a named model remain source-reported until primary evidence is attached.
- C. One successful example proves the whole architecture is superior.
- D. The mechanism name alone is enough for a production choice.

**Answer: B.** Activation names and exact expert counts attributed to a named model remain source-reported until primary evidence is attached. The correct answer bounds the conclusion; the others turn a local relation into a global guarantee.

---

### Question 5

Which evidence best matches the stated status of “Mixture-of-Experts: routing and capacity”?

- A. The route loads without an error.
- B. Every learner opened the file.
- C. Mixed: established mechanisms + source-reported Kimi K3-style choices.
- D. The same result is assumed on every hardware target.

**Answer: C.** Mixed: established mechanisms + source-reported Kimi K3-style choices. Product or mechanism evidence must remain attributed and measured; availability and completion do not prove value.

---

### Question 6

When should a simpler baseline be preferred to “Load balancing”?

- A. Never: the newest mechanism wins by default.
- B. As soon as one memory metric falls, regardless of quality.
- C. As soon as the diagram contains fewer components.
- D. When a controlled test shows equivalent quality with lower memory, latency, or complexity.

**Answer: D.** When a controlled test shows equivalent quality with lower memory, latency, or complexity. The choice depends on a measured trade-off on the real workload, not novelty or one isolated metric.

---

### Question 7

A learner gets the right result but cannot explain “Router.” Which remediation is most useful?

- A. Rebuild the first missing transformation, label its inputs and outputs, then test a neighboring case.
- B. Accept the answer because the final number is correct.
- C. Provide the final result a second time.
- D. Change several variables and ask for an intuition.

**Answer: A.** Rebuild the first missing transformation, label its inputs and outputs, then test a neighboring case. The remediation targets the first causal break and then requires transfer instead of rewarding a guessed result.

---

### Question 8

Which submission actually demonstrates the outcome “Separate total and active parameters.”?

- A. A list of terms without causal relations.
- B. A trace with starting data, transformations, observed result, boundary, and next experiment.
- C. A screenshot without values or interpretation.
- D. A confident claim without a baseline or threshold.

**Answer: B.** A trace with starting data, transformations, observed result, boundary, and next experiment. The correct submission makes the reasoning reproducible and the verdict revisable by future measurement.
