Day four · 11 August 2026 · sessions S-020 to S-025
Around fifty renders, six new measuring instruments, and the discovery that five of those six were lying to us. One had already published a false failure before it was caught.
It was also the day the pipeline got genuinely good. We can now direct an actor across a room, move a camera on a dial, hold a costume and light a set. And we still cannot do the one thing that sounds easiest of all.
Renders
~50
Instruments that lied
5 of 6
Of one frame, empty
53%
Record nearly lost
263 rows
Predictions correct
2 of 5
Research doc that never existed
1
The question that opened the day
Would a better model fix this? No — and the benchmark was measuring the wrong thing
Gabe asked a sharp question. Yesterday we found we had to choose between a beautiful picture and an obedient one. Would a stronger image model need less constraint to hold a layout, giving back the richness we were paying for control?
The honest answer required admitting our own reasoning was broken. We had been ranking candidate models by a benchmark score for spatial positioning. That benchmark measures how well a model turns text into layout — which is precisely the pathway we abandoned the previous day when we started injecting geometry directly. Buying the top-scoring model would have been buying a better version of the thing we had stopped using.
And there is no benchmark for what we need
There is no public measurement of depth adherence for any image model. Anyone who tells you one model obeys geometry better than another is reporting an impression, not a number. We could not shop our way out of this.
Then the ranking inverted completely: the best-scoring candidate has no depth control at all. The best model on paper was the worst possible fit for this pipeline. The proposed upgrade would have cost weeks and moved us backwards.
Continuity does not come from the model. It comes from the blockout — which is literally the same object across every camera.
What measuring found
Over half that frame was empty, and there was no wall in it
The previous day’s conclusion had been that the model cannot ornament a wall the depth map declares to be flat. Sensible, and wrong. When I measured the map instead of looking at it, 53% of the frame was pure black — and the reason is pure optics.
That produced the most useful piece of engineering knowledge on the project. A depth map is eight bits and linear, so the currency is grey levels per metre — on that room, 26.8. Whether a protrusion survives depends on the angle you view it from, and the effect is dramatic.
| Viewing angle | Effect on relief | Practical threshold |
|---|---|---|
| Grazing — the surface nearly edge-on | amplifies enormously: a 0.18 m aisle kerb reads as 1.90 m of depth shift | tiny details survive |
| Frontal — facing the surface | one to one, no amplification | needs ≥ 0.30 m to render |
| Past the far limit | nothing at all | impossible at any strength |
Yesterday’s default was wrong
0.65 holds one camera. A set needs four
We had adopted a control strength of 0.65 as the production default on the strength of a sweep that used one camera. When I ran four cameras at it, they did not describe one place — and the tell was the floor. The base camera’s dark polished composite became pale grey tile with cyan inlays in all three others. Desk counts differed. One arm grew a bank of wall consoles present nowhere else.
At 0.90 the spread in floor brightness across cameras roughly halves. So 0.65 is a fine setting for a single plate and the wrong setting for a set, and the blanket default was superseded the day after it was adopted.



And then the limit the whole method cannot pass. The window sits 18 to 22 metres out, while the room’s defining masses sit at 5 to 14. Every honest bracket clips the window to black. Its aperture is therefore reinvented on every render — a chamfered letterbox, then a rounded rectangle, then a full mullioned wall. The most identity-defining feature of the set is the one thing depth cannot carry, at any strength.
We tried the obvious fallback: lock the window in words, with a full four-part description. It failed — three different apertures across three plates, and three more across three videos. Words do not carry what depth omits.
The long watch
The room held. The camera did not
Three shots in one room: a wide down the centre line, a four-metre dolly to port, and a 22° pan. Deliberately small — big enough to stress two cameras in one space, a translation, an actor across a cut and a carried prop.
The room held beautifully. Floor brightness across the three cameras spread by a standard deviation of 1.8, against 13.5 in the earlier four-camera test at the same strength — measured with masks cast from the box model rather than eyeballed. One floor, one desk styling, one lamp vocabulary, one chair.
The dolly did not happen. Its plate carries the four-metre translation correctly. The render does not contain it: the glass is confined to the left third where the geometry spans the full frame, the central aisle and its rails are missing entirely, and the actor sits centre-frame where the depth map puts him at 78% across. He is centre-frame at the end of the previous shot too — so across the cut, the camera appears not to have moved at all.
The plate fixes a room. It does not fix a lens.
We then tried to fix it with language — 764 characters of composition description stating the finished frame explicitly, nothing else changed, same seed. It moved the actor from 49.3% to 47.7% across frame, against a plan of 78%. In the wrong direction, by a rounding error. Neither the plate nor composition language commands the lens.
Measuring all four renders together produced the finding that cost nothing and mattered most.
Two things nobody was looking for
A frame-difference scan found an unrequested cut inside a shot — at frame 77 of 124, nine times the median frame-to-frame change — in a render given a single shot marker, no timestamps, and the instruction “the camera is a static shot”. It turned out to be about length, not description: the same shot at 124 frames cut at frame 77 again, harder, while a 90-frame version did not cut at all. Not elapsed time either — the cut point sits comfortably inside the shorter clip and nothing happens there.
Both cautions were recorded with it, because the result is not free: 124 frames is the bottom of the model’s trained range and 90 is below it — and the 90-frame version turned the actor around, so he walks away up the aisle with his back to the lens where the text says he faces camera. Shortening bought the cut and cost the staging.
The second was stranger. The armour drifted through three different shoulder treatments across three renders although the subject definition text was byte-identical in all three — while the face, which carries two photographs, held perfectly. Identity is anchored by pictures; wardrobe was anchored only by words, and words are not an anchor.
Where the frontier now sits
Four things we can do that we could not that morning
| Now possible | What unlocked it |
|---|---|
| Direct an actor across a room | describe a journey, not a position |
| Move the camera on a dial | the model’s own closed camera vocabulary — which we had never actually used |
| Hold a costume between shots | a faceless photograph of the armour: the first thing that ever worked |
| Light a set | the prompt owns light outright; depth conditioning costs 1–1.4 stops and a cyan cast, and one high-key sentence recovers about 90% of it |
The single largest lever found all week was hiding in plain sight and has nothing to do with cleverness: plate resolution dominates everything. Rendering the source plate at 2048 instead of 1344 pixels is worth +26% detail on its own and +89% stacked with other changes — a bigger effect than removing the depth map altogether.
One documented belief was also inverted. We had written that shots sharing a single render get continuity “for free”. Scored properly against all eleven continuity vectors, that is backwards: a shared render shares the noise. It wins on actor position and loses on identity, wardrobe, geometry, appearance and lighting.
The real story of the day
Five of six new instruments were lying
To answer any of the above you have to build something that measures it. We built six measures that day. Five were caught giving wrong answers before they became claims. One was not.
- A travel measure that summed sub-pixel drift to zero, reporting a moving camera as stationary.
- A detail metric that made the winning plate look worst, until resolution was normalised.
- A structure measure that could not distinguish correct depth maps from wrong ones.
- A galaxy locator that kept finding desk lamps. Brightness alone finds the pale desk tops; warm-and-bright alone finds every amber indicator. A first pass with neither correction produced a tidy, entirely false conclusion that the galaxy and the actor had swapped places.
- A costume measure that scored a known-good control at −0.073, i.e. as a failure.
- And the one that got away: a camera-travel measure that saturates at ±160 pixels and, once saturated, wraps onto the repeating architecture of the set and returns a confident wrong number — sometimes with the wrong sign.
- That last one published a false failure for the tracking shot, since retracted, and understated a camera arc by a factor of ten.
A bounded search does not fail loudly when the true answer lies outside its bounds. It returns the best wrong answer inside them.
This generalises well past this project, and it is the most useful thing on this page. A measurement that cannot represent the real answer will hand you a plausible one instead — and a plausible number is far more dangerous than a missing one, because you will act on it. Every magnitude in the pipeline is now chained sub-pixel, and two of the eleven continuity vectors are recorded as having no instrument at all rather than being given a bad one.
Five predictions had been written down before any of these renders ran, precisely so they could be scored rather than rationalised afterwards. Two were right, one was right for the wrong reason, and two were wrong. The wrong ones: that high strength would hold all four cameras (it holds the geometry, but materials still drift) and that 0.65 would hold three of four (it holds one).
And the folder audited itself
A research document we had been citing for three days does not exist
The written record had grown to the point where nobody could verify it by reading, so I wrote a checker for it. Its first run found eight faults, four of which nobody knew about — and the worst is a good lesson about trusting your own paperwork.
Research document R-008 has never existed. The research index carries a full row for it marked current. An experiment write-up says a decision was taken “after research (R-008) ruled out the alternatives”. A document written days later cites it again. There is no file, and there never was one. The conclusions survive elsewhere; the sources and reasoning behind them are simply gone.
Not one of those four new faults is findable by reading. You have to try to open a file, or compare two timestamps. Prose alone re-invents — so draw the thing instead of describing it. That was true of sets, and it turns out to be equally true of a folder.
The same audit found that the most-used asset in the entire production — the face photograph appearing in 49 renders, more than twice anything else — existed only inside the renderer’s own scratch folder, protected by nothing.
And a scan of the render ledger reconstructed the previous day’s disaster without being told it had happened: fourteen renders carrying zero references, all on one evening between 17:18 and 21:24. Hours of machine time and two confidently wrong conclusions had been spent on a fault whose evidence was sitting inside the output files the whole time.
Then the record nearly ate itself
The ledger tool, run under the wrong Python interpreter, rewrote the render record from 385 rows to 122 — every still silently dropped, no error, exit status zero. The cause was an image library import sitting inside the same error handler used for “this file has no embedded workflow”, so a missing library looked exactly like a file with nothing in it.
It was caught only because the printed count disagreed with a number in another document, and recovered from version control. Two guards now exist: a missing library aborts and names the correct interpreter, and a rebuild that would shrink the record refuses to write at all without an explicit override.
A related bug had been quietly waiting to do worse. The same tool reset the human verdict field on every rebuild — so the project’s first properly graded sequence would have survived exactly until the next time anyone ran it. A rebuild must never destroy a judgement. That is now a rule with a test behind it.
Next: the door was in the drawing all along.
Day five discovers that the map the renderer was reading was not the map on disk, opens a reference channel that had been failing on a single number, and builds an instrument that immediately overturns the decision which asked for it.
Sources: docs/01_SESSION_LOG.md S-020 to S-025 · experiments E-023 to E-034 · research R-014 · decisions D-043 to D-051 · docs/PLAN_2026-08-11_model-vs-method.md for the pre-registered predictions. Tooling written this day: verdict.py, make_corridor.py and its MK II and MK III successors. Forty-four renders were driven through the renderer’s HTTP API in about five minutes of GPU time when the usual connector was unavailable — worth remembering that the convenience layer is not a dependency.