Instructor notes: Frame the problem before naming the mechanism. Collect an initial prediction and retain it for the exit ticket.
Instructor notes: Connect every step to the next with a causal verb. Flag any merely decorative arrow.
Instructor notes: Write the full announcement on the board — equation, ratio, benchmark — and ask: “what can we verify from here, offline?”. The room’s spontaneous sort gives the session its outline; keep it posted.
Instructor notes: Open by writing the mechanism on the board WITHOUT the product name, and introduce the name only after the arithmetic. If the name comes first, everything after inherits its prestige — exactly the bias this session fights.
Instructor notes: Answer: the whole core remains — read-compare-correct, gates, half-lives: recomputable algebra, independent of any product. Expected wrong answer: believing that without the name only “theory” remains — every number in the trace can be redone with a pencil.
Instructor notes: Hook question: “with ONE forgetting dial, how do you keep the customer’s name and flush the weather?”. The felt impossibility creates the need for the vector — the exact limit of session 16’s scalar α.
Instructor notes: Have them compute both half-lives (6.58 and 0.43), then ask what a 128-gate vector would look like. The jump from the 2-channel toy to real scale must be made by them, not announced.
Instructor notes: Answer: half-lives ln(0.5)/ln(0.9) ≈ 6.58 tokens and ln(0.5)/ln(0.2) ≈ 0.43; cost: 128 gate values per token versus 1. No, not free: production, bounding within [0,1], and per-channel debugging. Expected wrong answer: forgetting the bounding constraint (stability, session 16).
Instructor notes: Have the room fill in the t=4 column (0.656 and 0.0016), then ask which channel they would pick for a first name, which for a sentence’s tone. Choosing by use case makes the vector concrete.
Instructor notes: Take bets: “a perfect correction on a fast channel — what is it worth two tokens later?”. Collect numeric bets before computing 4×0.2² — the gap between the bets and 0.160 is the beat.
Instructor notes: Have them recite the read-compare-correct chain from memory before adding forgetting. If session 14 is shaky, the KDA stack becomes a magic word instead of a composition.
Instructor notes: Answer: the correction did NOT fail — it landed exactly on [4,4]; retention is what is short (0.160 after two steps of a 0.2 gate). Correcting = the state reached at t; retaining = what survives at t+n. Expected wrong answer: “β should have been larger” — β changes nothing here.
Instructor notes: Ask who remembers session 15’s verdict (chunking is exact). If nobody, redo it in sixty seconds with o₅ = [3,5]: this beat only makes sense leaning on that result.
Instructor notes: Explicitly recall the session-15 result: chunking is exact, therefore it adds no quality. Have the group articulate why “chunkwise computation” in an announcement is not a performance argument.
Instructor notes: Answer: exactness proves chunkwise breaks nothing — same numbers — hence also improves nothing: any quality gain comes from elsewhere. Expected wrong answer: counting “parallel prefill” as a quality argument — it is a cost argument.
Instructor notes: Hand out three printed sentences — equation, ratio, benchmark — and have them pinned on a wall-mounted evidence ladder. The placement disagreements ARE the beat’s content; do not adjudicate too fast.
Instructor notes: Run it aloud: each person phrases a sentence about K3-style, the group votes “established / reported / not evaluated”. Correct the phrasing, not the person. This is language training, not a knowledge test.
Instructor notes: Model answer: “the vendor technical report states a ratio of N KDA layers per attention layer; that is reported, not independently replicated — I treat it as an architecture hypothesis”. Have the group vote on two or three phrasings; reject those that drop the attribution.
Instructor notes: Have two claims from outside the course placed on the ladder (one from the last press release the room read). The ladder is only learned once it works on fresh material.
Instructor notes: Role-play the meeting: one learner reads the fused sentence, another politely interrupts and sorts it into three statuses. Two minutes, then swap roles — the target skill is verbal, not conceptual.
Instructor notes: Close by applying the labelling to the course itself: ask which claim on YOUR slides is weakest in evidence. An instructor who welcomes that question teaches evidence discipline better than ten slides about it.
Instructor notes: Answer: (a) established — redoing the arithmetic suffices; (b) reported — require the cited technical report and its exact scope; (c) not evaluated — require protocol, test sets, ablations, and variance before repeating it. Expected wrong answer: demanding the same evidence level for all three.
Instructor notes: Have everyone rewrite the defensible sentence in their own words, then read three aloud. The target skill is a sentence speakable in a meeting, not a taxonomy.
Instructor notes: Walk line by line. Locate an inconsistency at the first faulty step, not only on the final line.
Instructor notes: Have learners assign the column-2 statuses claim by claim before revealing the column. The expected disagreement is on row 3: that is exactly where “reported” separates from “established”.
Instructor notes: Retain initial and final values. Do not allow simultaneous changes that make the delta impossible to attribute.
Instructor notes: For each claim, have the group produce the smallest counterexample before giving the correction.
Instructor notes: Separate verifiable mechanism, reported implementation choice, and experimental result. Evidence precision must match claim precision.
Instructor notes: Expected: (1) e.g. “diag(a)·S decays each channel” / “the report states a 3:1 ratio” / “the long-context gain is not evaluated here”; (2) independent replication + published protocol (test sets, budget, ablations); (3) the established mechanism is enough to prototype — the key point: you can build on (a) without certifying (b) or (c). Misconception to harvest: “until it is verified, nothing can be done”. Eight minutes, solo then debrief.
Instructor notes: Rebuild the chain without looking at the slides, then fill the ticket in at most six lines. Compare with the opening prediction and name what actually changed.