Noise Lab, Chapters 5 through 8. Five bakers, six judges, one troublesome cake.
For LexiI made you a laboratory, doamnă. This seemed like a perfectly reasonable response to four chapters of a book.
Self-contained browser lab · no external libraries

Noise Lab

Learn Chapters 5–8 of Noise through one recurring judging system: build the distinctions, manipulate the sources of variability, retrieve them without prompts, then stress-test what the setup can actually support. Reset the entire experience whenever you want.

1 · ContextMeet the castGive the recurring judges, bakers, and criteria names before introducing the model.
2 · DistinguishName the taskDecide whether the judgment is predictive or evaluative before choosing an error or noise model.
3 · Predict + manipulateCommit, then change one causeState what you expect before moving a control; then compare expectation with the observed pattern.
4 · RetrieveName the phenomenonUse Test to recover the distinction without the worked example in view.
5 · CritiqueDiagnose a new caseApply the setup when more than one source of noise is present.
6 · TransferMake it defend itselfCommit to the evidence first, then compare your conclusion with a published criticism.
Step 1 of 6 · Context
Recommended path: Learn → Explore → Test → Critic Lab. In Explore, change one dimension at a time. In Critic Lab, make your own narrow claim before revealing anyone else’s argument. Explanations are delayed until after a prediction or response whenever possible.
Your brief · Lexi

The results are not certified yet.

You have been invited to audit a chocolate-cake competition before the final rankings are released. Everyone used the same rubric, yet the scores do not behave as neatly as the organizers expected. Your task is not to pick a winner. It is to determine why competent judges disagree, which disagreement the system should tolerate, and which procedures are manufacturing avoidable variability.

Margin note

Kira has no voting rights.

She does, however, reserve the right to disapprove of causal claims made from one mean, one panel, or one suspiciously persuasive senior judge. Gremlins 👹 appear only when a mistake deserves a little mockery; they never reduce the score.

Current stageLEARN · Build the model
Coverage0%
Practice0
Scenarios0 / 12
Learning cycles0
ReadingWandering
Mastery0%
Correct0 / 0
Guided path0 / 4
Cumulative learning record Nothing to prove yet, doamnă. Start anywhere useful. 0 concepts · 0 scenarios · 0 retrieval chapters · 0 critic encounters
0 / 280%
Suggested next stepStart with LearnMeet the cast first, then let the chapter examples give the controls meaning.
01
LearnMeet the system, then build the modelCast first. Criteria second. Concepts and worked examples only after the people have names.
Learn · local routeGive the ideas something concrete to attach to.
1 of 3
Do now: follow the highlighted step. The rail changes as the relevant section enters view.
i

Before the judging starts

This is a teaching simulation, not a dashboard. The story gives the variables meaning; the laboratory lets you change them.

The scene: the tasting is over, but the scorecards are still on the table. The organizers hand them to you, doamnă, because several rankings changed depending on who judged, when they tasted, and who spoke first in deliberation. You will reconstruct the failure mechanism before recommending any procedural fix.
Learn · first things first People → criteria → model → worked examples Meet the recurring cast before the machinery. Then the variables have somebody to belong to.
Meet the bakers
Color is a recognition cue, not a score.Hover, focus, or tap ? for why each contestant matters.
Home baker · 46

Marta Kowalska

Marta learned by watching her mother and grandmother rather than by measuring everything. She adjusts batter by sight and texture, and cares first about whether people want another slice.

In the simulation: deep flavor and moisture, but uneven layers create a sensory-success versus precision conflict.
Recent pastry graduate · 24

Leo Chen

Leo came to baking through formal training. He weighs precisely, tracks temperatures, and treats reproducibility as evidence that a result is deserved rather than lucky.

In the simulation: beautiful structure and control, but restrained flavor make him a useful counterpoint to Owen.
Traditional home baker · 61

Nina Petrescu

Nina has baked for decades and prefers darker, less sweet cakes with a denser crumb. She does not regard contemporary sweetness or airiness as neutral standards.

In the simulation: judges must decide whether difference from current expectations is a defect or an intentional style.
Experimental baker · 35

Owen Mercer

Owen experiments the way some people annotate books: relentlessly. His chocolate cake uses espresso, olive oil, sea salt, restrained frosting, and a deliberately soft crumb.

Anchor case: originality and flavor pull some judges upward while conventional technique pulls others downward.
Methodical recipe follower · 29

Sara Malik

Sara trusts tested recipes and executes them carefully. Her cake is balanced, clean, and difficult to fault, though few judges find it unforgettable.

In the simulation: she tests whether “no obvious defects” is the same thing as “excellent.”
Meet the judges
Each judge keeps a stable palette identity.Hover, focus, or tap ? for the learner-relevance note.
Pastry instructor · technical lens

Helen Ward

Helen has spent years teaching students to diagnose crumb, structure, emulsification, symmetry, and finishing technique. She sees execution errors quickly.

Tendency: technical departures carry more weight, so Leo often benefits and Owen may not.
Food writer · sensory lens

Marcus Bell

Marcus begins with the eating experience: aroma, flavor development, bitterness, sweetness, and finish. A technically imperfect cake can still win him over.

Tendency: especially useful for Chapter 7 because his judgment changes with the local tasting context.
Competition judge · control lens

Priya Shah

Priya is interested in whether the result looks controlled and reproducible. Avoidable irregularity matters because she treats it as evidence about the process.

Tendency: rewards consistency and penalizes results that seem difficult to reproduce.
Pastry chef · originality lens

Daniel Ruiz

Daniel works with unconventional flavor combinations and is willing to tolerate departures from convention when the departure produces something distinctive.

Tendency: often becomes Helen’s opposite on Owen and Leo, making pattern noise visible.
Community judge · lenient scale

Claire Bennett

Claire has judged local competitions for years and uses the upper end of the scale relatively freely. She cares strongly about whether a cake is pleasurable to eat.

Tendency: her stable leniency makes level noise easy to see against Thomas.
Senior judge · severe scale

Thomas Reed

Thomas reserves the top of the scale for unusually complete work. He can agree with Claire about ranking while placing the entire field lower.

Tendency: stable severity illustrates level noise; his seniority also makes his Chapter 8 comments socially consequential.
Competition rubric
Flavor30%
Texture25%
Technique20%
Appearance15%
Originality10%
Now build the machinery

What the tool actually does

You will watch the same judging system from four angles. Chapter 5 asks whether the panel is wrong on average or merely inconsistent. Chapter 6 asks where the inconsistency comes from. Chapter 7 holds the judge and cake constant and changes the occasion. Chapter 8 lets judges influence one another and asks whether consensus has improved the judgment or only synchronized it.

Your job: predict first, manipulate one variable at a time, read the consequence, then test whether you can diagnose a new pattern without hints.

What is fixed—and what can move

Every score belongs to one Judge × Cake × Occasion cell. The Explore laboratory lets you amplify stable severity differences, case-specific interactions, occasion effects, and—only in discussion-first mode—social influence.

Warm surfaces = story, interpretation, coachingCool surfaces = controls, data, scoring
Chapter 8 is deliberately an overlay, not a fourth statistical dimension: judges begin changing one another’s information environment.
Why baking?Subjective enough to require judgment; structured enough to compare.

Baking keeps the problem emotionally neutral while preserving exactly what Noise needs: shared criteria, professional discretion, repeatable cases, and social influence.

What the simulation is not doingIt is not deciding which aesthetic preference is morally correct.

The numerical model is illustrative. It asks when a competition would want competent judges to agree more closely—not whether all human difference should be eliminated.

Scores and privacyThree progress signals answer three different questions.

Mastery is cumulative Test accuracy. Coverage is bounded at 100% across eight concepts. Practice points reward repetition and can keep growing. Learning cycles count complete feedback loops: a solved Scenario run or a finished randomized Test. Mastery remains cumulative Test accuracy. All are stored only in this browser when local storage is available. Nothing is transmitted.

Persistence status will be checked when the app starts.
5

Chapter 5 · The error equation needs a target

Use the predictive case only where a true value can later be recovered.

Prediction station · Owen’s final cake weight

Before Owen’s cake is weighed, six judges independently predict its final weight: 2.10, 2.62, 2.25, 2.70, 2.55, and 2.18 kg. The scale later reveals 2.40 kg.

Here there is an independently recoverable target. The predictions average exactly 2.40 kg, yet the individual predictions are widely dispersed.

Boundary check

This is where the mean-squared-error decomposition belongs: the target exists independently of the judgments.

Predictive takeaway: for this recoverable-target task, MSE = Bias² + Noise². Low bias can coexist with substantial noise even when the average prediction is correct.
Do not carry this equation across the task boundary: the later 0–10 quality scores are evaluative judgments. They do not have a hidden true score of 7.0.
6

Chapter 6 · Switch back to evaluation and decompose the noise

The cake-quality scores have no independently correct value; disagreement can still be analyzed.

Task switch: “What will the cake weigh?” was predictive. “How good is this cake?” is evaluative. For the evaluative task, inspect unwanted variability—especially level and pattern noise—without applying the predictive error equation.

Level noise

Claire and Thomas rank cakes similarly, but Claire uses a consistently higher portion of the scoring scale.

Look for the geometry

If two judges’ lines are roughly parallel but vertically separated, which part of the disagreement belongs to the judge rather than the cake?

Research analogue: rater severity / leniency; similar to a rater random intercept.

Pattern noise

Helen and Daniel can have similar average severity but cross repeatedly when different cakes activate different evaluative priorities.

Look for the interaction

If their average scores are similar, what does repeated line-crossing tell you that a mean comparison cannot?

Research analogue: judge × case interaction. Different feature weighting is one possible mechanism.
Common mistake: defining pattern noise as “different preferences.” Preferences are one possible cause; the observable phenomenon is the judge × case interaction.
7

Chapter 7 · Hold judge and cake constant; change the occasion

Now the same measuring instrument changes over time.

Marcus re-tastes Owen

Early tasting: 7.8. After several unusually sweet cakes: 8.5. Later, after technically exceptional entries: 7.3.

Hold two dimensions constant

Judge = Marcus and Cake = Owen. Only Occasion changes. What source of variability remains available to explain the movement?

Takeaway: reliability can fail within the same judge, not only between judges.
8

Chapter 8 · Add dependence between judges

Consensus can increase while informational independence decreases.

Two panels, same cake, different first frames

One panel hears Thomas’s technical criticism first; another hears Daniel’s originality praise first. Both groups become internally consistent, but their final conclusions diverge.

Do not equate agreement with reliability

If within-panel spread shrinks in both groups, what additional comparison is needed before concluding that noise fell?

Judgment hygiene: aggregate first; discuss second.
Transfer: ask whether independently formed groups reproduce one another, not merely whether one group reaches consensus.
Σ

Cumulative model

The framework should grow as one nested system.

Predictive error

When a true target exists, mean error can be decomposed into bias and noise. Do not import that equation into the cake-quality ratings.

Chapter 5

Level noise

Stable severity or leniency differences between judges.

Chapter 6

Pattern noise

Judge × cake interaction: similar means, different case reactions.

Chapter 6

Occasion noise

Same judge, same case, different occasion.

Chapter 7
Bounded learning measureConcept coverageEight ideas, each counted once when demonstrated in a scenario.
0%
Chapter 8 is not another sibling variance component. It is a social process that can make individual errors correlated.
Learn checkpoint You now know who is judging, what they are judging, and which distinctions the Lab will ask you to manipulate. Explore is the next step because prediction is more useful once the cast and worked examples have given the controls meaning.
Guided path: Learn → Explore → Test → Critic Lab. You can wander at any time; this route follows the pedagogical dependency chain.
02
ExploreManipulate the systemChange one mechanism at a time, predict what should move, then inspect the consequence.
Explore · local routeRun one small experiment, not a control-panel safari.
1 of 3
Default: Start with the matrix. I will keep the extra machinery out of your way until you ask for it.
LAB

Explore the judging system

The model still contains Judge × Cake × Occasion. Instead of making you decode all three spatially, choose the view that best answers the question you are asking.

Use the right view for the right question Matrix gives the overview. Score Strip exposes Chapter 5 spread. Score Lines reveal Chapter 6 level and pattern noise. Occasion Timeline isolates Chapter 7. Panel Comparison makes Chapter 8’s consensus trap visible.

Judge × Cake matrix

Rows are judges; columns are cakes; occasion is fixed. Click a cell to inspect it.
Recommended default: position + number
Why this view: the matrix provides the broadest overview without hiding the actual score behind color.
dot position = score number = exact value Heat mode uses the uploaded blue-to-navy subset rather than red/green “good/bad” semantics.
Selected score
Judge mean
Cake mean
Within-judge SD

What just happened?

The laboratory waits until you pause, then translates the visual change back into the judging story.

Choose a view or change one control. The explanation will appear here after you pause.

Scenario bank

Each card now loads a real laboratory configuration, asks you to diagnose it, and contributes to cumulative practice.

0 / 12 completedFirst-time diagnoses earn the most practice points; completed runs count as learning cycles. Contrast sequences deliberately place similar mechanisms beside one another.
03
TestRetrieve and transferCommit before feedback: distinguish similar mechanisms and apply them to new cases.
Test · local routeAnswer first. I explain after you commit.
1 of 3
One question at a time. The interface keeps the current retrieval task primary.
T

Adaptive test

Questions are drawn from a larger bank and scored cumulatively. Wrong answers trigger targeted explanations.

The default follows your reading position.
The audit hearing: I stop showing you the chapter labels now. You get twelve shuffled findings and have to decide what each one means. Answer first; explanation comes after. There are three Gremlins in the jar. Kira denies everything, which is exactly what Kira would say.
Why I make you answer first: I learn more from your answer if you have to produce it before I show you mine. Read the explanation even when you are right. It tells you which feature separated the right answer from the tempting wrong one.
Question 1

Test complete

Your mastery profile is stored locally in this browser until you reset it.

Chapter 5
Chapter 6
Chapter 7
Chapter 8

What to do next

The tool is designed for cycling, not one-pass completion. Review the weak distinction, manipulate it once in Explore, then take a fresh randomized test. When you want a completely clean run, use Reset Everything.

04
Critic LabNow you get to be difficult on purposeThe framework has to earn its conclusions too: compare first, make your own call, then let the published argument into the room.
Critic Lab · local routeMake the evidence speak before the citation does.
1 of 4
Rule: no citation gets to do your thinking first, doamnă.
?

Now I let four people argue with the book.

You know what the book is trying to do. Now I want you to see what happens when serious people disagree with it.

Why you are here

This is the stress-test room. Each encounter changes one feature of the case and asks what still follows. First make the narrowest claim the evidence supports; only then reveal the published challenge.

In other words: Do your own thinking before I show you the citation, doamnă. Yes, I know. Terribly inconvenient.

Four moves, always in this order

  1. Compare the two cases.
  2. Commit to your narrowest defensible conclusion.
  3. Reveal the critic’s specific objection.
  4. Keep the answer, counterevidence, or remaining unknown visible.
Warm = argument, provenance, interpretationCool = evidence to inspect before anyone argues with you
Encounter 1 of 4Compare → Commit → Reveal → Keep
Encounter 1 · Causation

Same disagreement, different causes

The number tells you that two judges differ. It does not automatically tell you why. This encounter makes that gap visible before anyone names it for you.

Exemplar A

The rule was misunderstood.

Helen6.2
Daniel7.8

Daniel misreads the scoring rule and doubles originality.

Observed: 1.6-point disagreement.
Known mechanism: misunderstanding.
Contrast B

The rule was understood.

Helen6.2
Daniel7.8

Both understand the rubric. They consistently weight originality differently.

Observed: same 1.6 points.
Known mechanism: stable judgment policy.
Your call before the critic appears

What have the scores alone established?

Gilhooly + Sleeman enter the bakery
G+S

Kenneth Gilhooly & Derek Sleeman

Gilhooly is a cognitive psychologist; Sleeman works in artificial intelligence and knowledge engineering. Together they published a direct scholarly commentary on Noise.

Why they are here nowYou just saw identical disagreement produced by different mechanisms. Their objection belongs exactly at that boundary: disagreement is observable; its cause requires more evidence.
The narrow challengeObserved disagreement does not itself identify whether its cause is error, bias, stable judgment-policy differences, or another mechanism.
What you can safely take from itThe scores establish disagreement. They do not, by themselves, establish what produced it.
Encounter 2 · Evidence

A good framework still has to survive its own evidence standard

Two findings can look equally tidy on a screen while resting on very different amounts of evidence. Here you are judging the support beneath the result, not the prettiness of the result.

Exemplar A

Repeated and broad.

Four hundred judgments recur across judges, cakes, and occasions.

Support: repeated measurement over multiple observations.
Contrast B

Small and striking.

Seven judges make one set of judgments in one narrow setting.

Support: one small sample.
Your call before the critic appears

Same graph. Same confidence?

Sood + Gelman enter the bakery
S+G

Gaurav Sood & Andrew Gelman

Sood and Gelman are statisticians who reviewed Noise in CHANCE. They focus on whether some behavioral claims are supported by evidence strong enough to carry the generalization.

Why they are here nowYou just compared broad repeated evidence with a small striking sample. Their challenge is about how much confidence the evidence can actually bear.
The narrow challengeA book that teaches caution about variable human judgment should also be cautious when broad claims lean on small or fragile behavioral studies.
What you can safely take from itA striking result is not automatically a stable generalization.
Encounter 3 · Useful difference

The outlier might be broken. It might also have noticed the cream.

Reducing disagreement sounds attractive until the disagreeing judgment contains information the group missed. The question is not whether variation exists, but whether you know enough to call it unwanted.

Exemplar A

Useless outlier.

Panel≈7.0
Thomas3.1

Thomas was distracted and tapped the wrong score.

Contrast B

Informative outlier.

Panel≈7.0
Thomas3.1

Thomas noticed spoiled cream everyone else missed.

Your call before the critic appears

If you suppress the outlier before asking why it exists, what could happen?

Krakauer + Wolpert enter the bakery
K+W

David Krakauer & David Wolpert

Krakauer and Wolpert are Santa Fe Institute researchers whose work spans complex systems, collective intelligence, information, computation, and optimization.

Why they are here nowThomas may be a nuisance or the only judge who noticed spoiled cream. Their exchange with the authors asks whether reducing variation can sometimes erase useful independent information.
The narrow challengeVariation can sometimes preserve independent information or aid collective exploration. Convergence can remove information as well as inconsistency.
The authors’ answerThey define noise as unwanted variability. If variation is useful, it is not noise in their technical sense.
What remains liveCan a decision system reliably know which variation is unwanted before it understands what produced the variation?
Encounter 4 · Machine consistency

The machine agrees with itself every time. Suspiciously well behaved.

Repeatability removes one kind of variability. It does not tell you whether the rule, training data, or objective deserves your confidence.

Exemplar A

Explicit rule.

Taste 50%, texture 30%, appearance 20%. The rule is visible and repeatable.

Repeatability: high.
Rule: directly inspectable.
Contrast B

Learned rule.

A model learns from 20,000 historical scores. Those scores overreward decoration and disagree about subtle flavors.

Repeatability: high.
Inherited structure: training data and development choices.
Your call before the critic appears

The model gives the same score twice. What has that established?

Lauren Yu enters the bakery
LY

Lauren J. Yu

Yu is the author of a Michigan Law Review critique focused on how Noise treats artificial intelligence as a route to more consistent judgment.

Why she is here nowThe model is perfectly repeatable. That is precisely why this is the right moment to ask whether consistency has been mistaken for correctness.
The narrow challengeMachine-learning systems do not remove human choices. Training examples, objectives, model development, and testing can carry consequential error even when outputs are repeatable.
What you can safely take from itRepeatability establishes output consistency. It does not establish correctness or independence from human judgment.

One last nuisance, doamnă

Six judges disagree. You do not yet know whether they understood the rubric differently, whether the evidence base is too small, whether an outlier noticed something real, or whether a scoring system inherited a bad rule from earlier judgments.

So what do you need next? More information about the mechanism, the evidence, or the system that produced the result.

Sometimes the most disciplined answer is also the least theatrical one: I cannot classify it yet. Annoying. Useful. Very Chapter 6 of you.

0 notes