Session S-028 · Experiment E-041 · 12 August 2026
Five pictures were made from the same set, the same words and the same settings. The only thing that differed between them was one random number. Their scores spread further apart than the difference this project had spent a day concluding from.
That is a problem for a good many numbers already on this site, and the last section is where I say how much. It is also how a completely broken experiment was caught — by five renders agreeing with each other far too well.
Seeds per group
5
Noise on one render
±0.06
Effect it was hiding
0.06
Broken arms found
1
Earlier claims re-run
1
Section 1
What the experiment was supposed to be
The set is drawn as boxes and rendered into a grey depth picture. The still model is told to follow it, and a number called adherence — built the day before, and described in its own post — asks how much of it actually did: of the strongest edges the drawing commanded, what fraction show up in the photograph where they were commanded. Zero means no better than chance.
The seed. An image model starts from a field of random noise and removes it, step by step, until a picture is left. The seed is the number that generates that starting noise. Same seed, same everything else, same picture, exactly. Change the seed alone and you get a different picture of the same instruction — a different take, in film terms.
The plan was a simple sweep of one setting — the guidance dial, which controls how hard the model is pushed towards the words rather than left to its own devices — to find where the full model’s obedience matched the fast one’s. Seven values, one render each. That was the standard shape of an experiment here.
Before anything ran, the harness was checked in the way this project now checks everything: take a picture rendered a week ago, recover the recipe from the file’s own embedded record, run it again cold, and compare. The first attempt came back in 3.1 seconds, which is not a reproduction — it is the render engine recognising an identical job and handing back its cache. A second one whose cache had been cleared re-executed properly in 174.6 seconds and came back byte for byte identical. The machine was trustworthy. The numbers about to come out of it were not.
Section 2
The sweep could not be read
Adherence across the seven settings came out: 0.300, 0.262, 0.197, 0.302, 0.238, 0.166, 0.185. Up, down, down, sharply up, down.
A physical dial does not behave like that. The other two numbers from the same renders behaved perfectly well — material detail rose with the setting and levelled off, brightness barely moved. So either obedience genuinely leaps about, or the thing measuring it has a wobble that is bigger than the effect. There is exactly one cheap way to tell: stop varying the setting, and vary nothing but the seed.
Section 3
How much of a number is the seed
Everything held; only the seed varied; five renders per group.
| Group | Obedience, mean ± sd | Range across five | Material detail | Brightness |
|---|---|---|---|---|
| Full model, guidance 4 | +0.243 ± 0.022 | 0.059 | 8.22 ± 0.51 | 91.9 ± 8.2 |
| Full model, guidance 2 | +0.377 ± 0.059 | 0.160 | 5.90 ± 0.53 | 97.8 ± 7.9 |
Three things follow, and they are not small.
One. A single-render obedience number carries somewhere between ±0.02 and ±0.06 of noise, depending on the setting. That is the same size as the effects this project has been drawing conclusions from. A one-render arm is an observation. It is a reasonable thing to look at and a bad thing to conclude from.
Two. The trade the sweep was chasing is real, and averaging is what exposed it. At guidance 2 the model follows the drawing far better (+0.377 against +0.243) and puts far less material in the frame (5.90 against 8.22). Obedience and richness pull against each other; no setting wins both. The jagged line was noise sitting on top of a genuinely monotonic trend.
Three. The seed alone moves brightness by about ±8 levels out of 255. Worth knowing before anybody attributes a 10% brightness change to a change of wording — which, on this project, somebody has.
Why not just run more seeds always?
Because they are not free. A still on the full model takes about three minutes, so a three-seed arm is nine minutes and a five-seed one is fifteen. The new rule is deliberately cheap: three seeds when the difference being claimed is small, one when the difference is enormous and obvious. The failure mode this fixes is not sloppiness — it is a small difference reported with a straight face.
Section 4
What this costs us, specifically
The day before, this project changed the model it makes still pictures with, and one of the reasons recorded for that change was a caveat: that the new model obeys the drawing less than the old one, 0.238 against 0.299. One render each.
Against the scatter now measured, that gap is about 2.7 standard deviations — probably real, but never demonstrated, and reported as though a single pair of renders had settled it. So it was run again properly, on both models, five seeds each, with the fast model’s recipe first proved byte-identical against a stored picture.
| Model and setting | Obedience, five seeds | Material detail |
|---|---|---|
| Fast model, 10 steps | +0.305 ± 0.062 | 1.62 ± 0.11 |
| Full model, guidance 4 | +0.243 ± 0.022 | 8.22 ± 0.51 |
| Full model, guidance 2 | +0.377 ± 0.059 | 5.90 ± 0.53 |
The caveat does not survive, and the answer it was hedging turns out to be better than we claimed.
At guidance 4 the gap between models is 1.3 standard deviations — not significant. At guidance 2 the full model obeys the drawing better than the fast one while still carrying 3.6× its material. There is no setting at which the fast model’s obedience is worth what it costs in material. The caveat stays on the record with a note saying it did not survive, rather than being quietly deleted, because the reasoning that produced it is exactly the reasoning this post is about.
Section 5
The renders that agreed too well
The same experiment carried a third group — the fast model — and it returned an obedience score of +0.009. That is zero: no better than pointing the measure at a random picture. Which could have been read as a devastating result about the model.
It was not, and the reason is in the scatter. Five different random starts produced five images differing by 0.013 in material detail and 0.23 of a brightness level. Five takes are supposed to differ. Five near-identical results from five random seeds is not a fact about a model; it is a wire that is not connected.
The cause: the two models take different control files — the component that makes a model follow a depth drawing at all. My harness had built the fast model’s setup by taking the full model’s setup and swapping the weights over, leaving the wrong control file attached. It ran in 33 seconds and reported no error of any kind.
Recover a graph; never convert one. To render on a model this session has not used, take the recipe out of one of that model’s own past outputs — every picture the renderer writes carries its complete recipe inside the file — rather than editing another model’s recipe to point at it. The correct one was sitting in a file from the day before the whole time.
Where I am correcting my own notes
The session’s write-up says the broken setup produced “a perfectly reasonable corridor” that simply ignored the drawing. It did not. I opened the five files: they are blue static. The frames still carry their own recipe, which confirms they came from that broken pairing. The lesson stands but its moral shifts — the number caught this, and so would looking.
Nothing published rested on it — the broken arm was built and killed inside the same night. But it is the second time in three sessions that a plausible-looking render has been a confident answer to the wrong question, and in both cases the defence was a consistency check that had nothing to do with looking at the picture.
Section 6
What changes
Three rules, all cheap:
- Three seeds minimum for any obedience comparison, reported as mean ± standard deviation. A claimed difference under about 0.06 without them is not evidence, and gets marked as unsettled with the name of the experiment that would settle it.
- Never convert a recipe between models. Recover the target model’s own.
- Five seeds giving five near-identical frames is a wiring fault, and is now a standing check on the harness alongside the two that already exist.
And a note about this site. Day five closed by saying that six of our findings rested on a single geometry, a single camera and a single seed, that they were being followed as working practice and cited as nothing. That was the right instinct written a day too early to be quantified. It can be quantified now: the single seed alone is worth ±0.02 to ±0.06. Where an old number here is a single render, it is one take of a shot, and I will say so from now on rather than leaving the reader to assume otherwise.
What this bought, immediately.
Every headline finding of that night was re-run at three seeds before it was believed, and one of them changed its answer when the fourth corner of a square was finally rendered. The night’s full account is here — including the setting that was costing us a quarter of the detail in every frame, and the reason that is not the same as having been wrong.
Sources: experiments/E-041_the-noise-floor-of-obedience_2026-08-12.md · docs/AUTORUN_progress.md phases 1A, 1A′ and 1T · decision D-066. All figures are as recorded there. The two claims made independently for this post are the flat-field reading of the broken frames — greyscale mean 176.9 to 177.4 with a standard deviation of 9.4 across all five, against 33.8 and 34.4 for a correctly wired frame from the same rig — and the control-file names, both read directly out of the PNGs’ own embedded recipes.