Day four: five of our six new instruments were lying

A contact sheet. At the top, the grey reference plate of an empty operations hall; beneath it, four frames from a single take of a man in white armour standing before the galaxy window, holding a tablet in the first frame and empty-handed in the rest.

Day four · 11 August 2026 · sessions S-020 to S-025

Around fifty renders, six new measuring instruments, and the discovery that five of those six were lying to us. One had already published a false failure before it was caught.

It was also the day the pipeline got genuinely good. We can now direct an actor across a room, move a camera on a dial, hold a costume and light a set. And we still cannot do the one thing that sounds easiest of all.

Renders

~50

Instruments that lied

5 of 6

Of one frame, empty

53%

Record nearly lost

263 rows

Predictions correct

2 of 5

Research doc that never existed

1

The question that opened the day

Would a better model fix this? No — and the benchmark was measuring the wrong thing

Gabe asked a sharp question. Yesterday we found we had to choose between a beautiful picture and an obedient one. Would a stronger image model need less constraint to hold a layout, giving back the richness we were paying for control?

The honest answer required admitting our own reasoning was broken. We had been ranking candidate models by a benchmark score for spatial positioning. That benchmark measures how well a model turns text into layout — which is precisely the pathway we abandoned the previous day when we started injecting geometry directly. Buying the top-scoring model would have been buying a better version of the thing we had stopped using.

And there is no benchmark for what we need

There is no public measurement of depth adherence for any image model. Anyone who tells you one model obeys geometry better than another is reporting an impression, not a number. We could not shop our way out of this.

Then the ranking inverted completely: the best-scoring candidate has no depth control at all. The best model on paper was the worst possible fit for this pipeline. The proposed upgrade would have cost weeks and moved us backwards.

Continuity does not come from the model. It comes from the blockout — which is literally the same object across every camera.

What measuring found

Over half that frame was empty, and there was no wall in it

The previous day’s conclusion had been that the model cannot ornament a wall the depth map declares to be flat. Sensible, and wrong. When I measured the map instead of looking at it, 53% of the frame was pure black — and the reason is pure optics.

Why there was no side wall to ornament A wide lens in a wide room sees mostly nothing. Plan view, to scale. SIDE WALL (port) SIDE WALL (starboard) 24 m wide 24 mm lens bracket ends · 14.9 m cone reaches the wall · 17.4 m the wall only exists 2.5 m beyond where the map stops So the map did not say “this wall is flat”. It said “empty, and far”.
The conclusion survived; the mechanism was wrong. And a new precondition appeared: you can only model detail onto a surface that is inside the frame and inside the bracket.

That produced the most useful piece of engineering knowledge on the project. A depth map is eight bits and linear, so the currency is grey levels per metre — on that room, 26.8. Whether a protrusion survives depends on the angle you view it from, and the effect is dramatic.

Viewing angleEffect on reliefPractical threshold
Grazing — the surface nearly edge-onamplifies enormously: a 0.18 m aisle kerb reads as 1.90 m of depth shifttiny details survive
Frontal — facing the surfaceone to one, no amplificationneeds ≥ 0.30 m to render
Past the far limitnothing at allimpossible at any strength
Those 0.18 m kerbs came back as waist-height rails. This is Gabe’s original “boxes render as boxes” objection arriving through a route neither of us predicted — not through the box, but through the viewing angle.

Yesterday’s default was wrong

0.65 holds one camera. A set needs four

We had adopted a control strength of 0.65 as the production default on the strength of a sweep that used one camera. When I ran four cameras at it, they did not describe one place — and the tell was the floor. The base camera’s dark polished composite became pale grey tile with cyan inlays in all three others. Desk counts differed. One arm grew a bank of wall consoles present nowhere else.

At 0.90 the spread in floor brightness across cameras roughly halves. So 0.65 is a fine setting for a single plate and the wrong setting for a set, and the blanket default was superseded the day after it was adopted.

Contact sheet of takes for shot one, the aisle wide
Shot 1 — the wide down the centre line
Contact sheet of takes for shot two, the dolly to port
Shot 2 — the 4 m dolly, the hard one
Frame from shot three, the pan to port
Shot 3 — the pan, which worked
A fault caught mid-experiment: three of the four camera arms had inherited the first camera’s depth bracket. One arm’s map turned out to contain only a wall, a floor and one desk slab — and the model dutifully filled the void with a staircase with handrails that exists nowhere in the blockout. Re-bracketing per camera cut one arm’s unconstrained frame from 46.5% to 21.6%. Near and far are per-camera properties.

And then the limit the whole method cannot pass. The window sits 18 to 22 metres out, while the room’s defining masses sit at 5 to 14. Every honest bracket clips the window to black. Its aperture is therefore reinvented on every render — a chamfered letterbox, then a rounded rectangle, then a full mullioned wall. The most identity-defining feature of the set is the one thing depth cannot carry, at any strength.

We tried the obvious fallback: lock the window in words, with a full four-part description. It failed — three different apertures across three plates, and three more across three videos. Words do not carry what depth omits.


The long watch

The room held. The camera did not

Three shots in one room: a wide down the centre line, a four-metre dolly to port, and a 22° pan. Deliberately small — big enough to stress two cameras in one space, a translation, an actor across a cut and a carried prop.

The room held beautifully. Floor brightness across the three cameras spread by a standard deviation of 1.8, against 13.5 in the earlier four-camera test at the same strength — measured with masks cast from the box model rather than eyeballed. One floor, one desk styling, one lamp vocabulary, one chair.

The dolly did not happen. Its plate carries the four-metre translation correctly. The render does not contain it: the glass is confined to the left third where the geometry spans the full frame, the central aisle and its rails are missing entirely, and the actor sits centre-frame where the depth map puts him at 78% across. He is centre-frame at the end of the previous shot too — so across the cut, the camera appears not to have moved at all.

The plate fixes a room. It does not fix a lens.

We then tried to fix it with language — 764 characters of composition description stating the finished frame explicitly, nothing else changed, same seed. It moved the actor from 49.3% to 47.7% across frame, against a plan of 78%. In the wrong direction, by a rounding error. Neither the plate nor composition language commands the lens.

Measuring all four renders together produced the finding that cost nothing and mattered most.

The room lands roughly where planned. The man never does Error against the planned position, in percentage points across the frame. Four renders, same set. on the mark +30 −30 shot 1shot 2 t1shot 2 t2shot 3 +26.9−28.7−16.0−30.3 −10.1−17.1−1.3−16.6 the actor the room (measured on the galaxy) He is never within 16 points of his mark, under any wording. The set is in the depth map; the actor never was.
Which points at the fix: put the man in the depth map as a mass, exactly as we do for walls. It is our own rule, applied to the one element we had never applied it to.

Two things nobody was looking for

A frame-difference scan found an unrequested cut inside a shot — at frame 77 of 124, nine times the median frame-to-frame change — in a render given a single shot marker, no timestamps, and the instruction “the camera is a static shot”. It turned out to be about length, not description: the same shot at 124 frames cut at frame 77 again, harder, while a 90-frame version did not cut at all. Not elapsed time either — the cut point sits comfortably inside the shorter clip and nothing happens there.

Both cautions were recorded with it, because the result is not free: 124 frames is the bottom of the model’s trained range and 90 is below it — and the 90-frame version turned the actor around, so he walks away up the aisle with his back to the lens where the text says he faces camera. Shortening bought the cut and cost the staging.

The second was stranger. The armour drifted through three different shoulder treatments across three renders although the subject definition text was byte-identical in all three — while the face, which carries two photographs, held perfectly. Identity is anchored by pictures; wardrobe was anchored only by words, and words are not an anchor.

Directing an actorHe crosses the room from 77% to 24% across frame, as instructed — because the instruction described a journey rather than a position.
Moving the cameraA truck at large amplitude — a real 3.2× dial, once we used the model’s own closed camera vocabulary instead of describing the move.
An arcRotation remains far better behaved than translation — the finding from day three, holding up with an actor in the frame.
Nine experiments in one stretch moved the frontier further than anything else in the week.

Where the frontier now sits

Four things we can do that we could not that morning

Now possibleWhat unlocked it
Direct an actor across a roomdescribe a journey, not a position
Move the camera on a dialthe model’s own closed camera vocabulary — which we had never actually used
Hold a costume between shotsa faceless photograph of the armour: the first thing that ever worked
Light a setthe prompt owns light outright; depth conditioning costs 1–1.4 stops and a cyan cast, and one high-key sentence recovers about 90% of it
Still impossible, after four separate instruments failed at it: specifying where the camera stands. Also still impossible: putting detail on a surface using words.

The single largest lever found all week was hiding in plain sight and has nothing to do with cleverness: plate resolution dominates everything. Rendering the source plate at 2048 instead of 1344 pixels is worth +26% detail on its own and +89% stacked with other changes — a bigger effect than removing the depth map altogether.

One documented belief was also inverted. We had written that shots sharing a single render get continuity “for free”. Scored properly against all eleven continuity vectors, that is backwards: a shared render shares the noise. It wins on actor position and loses on identity, wardrobe, geometry, appearance and lighting.


The real story of the day

Five of six new instruments were lying

To answer any of the above you have to build something that measures it. We built six measures that day. Five were caught giving wrong answers before they became claims. One was not.

  • A travel measure that summed sub-pixel drift to zero, reporting a moving camera as stationary.
  • A detail metric that made the winning plate look worst, until resolution was normalised.
  • A structure measure that could not distinguish correct depth maps from wrong ones.
  • A galaxy locator that kept finding desk lamps. Brightness alone finds the pale desk tops; warm-and-bright alone finds every amber indicator. A first pass with neither correction produced a tidy, entirely false conclusion that the galaxy and the actor had swapped places.
  • A costume measure that scored a known-good control at −0.073, i.e. as a failure.
  • And the one that got away: a camera-travel measure that saturates at ±160 pixels and, once saturated, wraps onto the repeating architecture of the set and returns a confident wrong number — sometimes with the wrong sign.
  • That last one published a false failure for the tracking shot, since retracted, and understated a camera arc by a factor of ten.

A bounded search does not fail loudly when the true answer lies outside its bounds. It returns the best wrong answer inside them.

This generalises well past this project, and it is the most useful thing on this page. A measurement that cannot represent the real answer will hand you a plausible one instead — and a plausible number is far more dangerous than a missing one, because you will act on it. Every magnitude in the pipeline is now chained sub-pixel, and two of the eleven continuity vectors are recorded as having no instrument at all rather than being given a bad one.

Five predictions had been written down before any of these renders ran, precisely so they could be scored rather than rationalised afterwards. Two were right, one was right for the wrong reason, and two were wrong. The wrong ones: that high strength would hold all four cameras (it holds the geometry, but materials still drift) and that 0.65 would hold three of four (it holds one).

And the folder audited itself

A research document we had been citing for three days does not exist

The written record had grown to the point where nobody could verify it by reading, so I wrote a checker for it. Its first run found eight faults, four of which nobody knew about — and the worst is a good lesson about trusting your own paperwork.

Research document R-008 has never existed. The research index carries a full row for it marked current. An experiment write-up says a decision was taken “after research (R-008) ruled out the alternatives”. A document written days later cites it again. There is no file, and there never was one. The conclusions survive elsewhere; the sources and reasoning behind them are simply gone.

Not one of those four new faults is findable by reading. You have to try to open a file, or compare two timestamps. Prose alone re-invents — so draw the thing instead of describing it. That was true of sets, and it turns out to be equally true of a folder.

The same audit found that the most-used asset in the entire production — the face photograph appearing in 49 renders, more than twice anything else — existed only inside the renderer’s own scratch folder, protected by nothing.

And a scan of the render ledger reconstructed the previous day’s disaster without being told it had happened: fourteen renders carrying zero references, all on one evening between 17:18 and 21:24. Hours of machine time and two confidently wrong conclusions had been spent on a fault whose evidence was sitting inside the output files the whole time.

Then the record nearly ate itself

The ledger tool, run under the wrong Python interpreter, rewrote the render record from 385 rows to 122 — every still silently dropped, no error, exit status zero. The cause was an image library import sitting inside the same error handler used for “this file has no embedded workflow”, so a missing library looked exactly like a file with nothing in it.

It was caught only because the printed count disagreed with a number in another document, and recovered from version control. Two guards now exist: a missing library aborts and names the correct interpreter, and a rebuild that would shrink the record refuses to write at all without an explicit override.

A related bug had been quietly waiting to do worse. The same tool reset the human verdict field on every rebuild — so the project’s first properly graded sequence would have survived exactly until the next time anyone ran it. A rebuild must never destroy a judgement. That is now a rule with a test behind it.

Next: the door was in the drawing all along.

Day five discovers that the map the renderer was reading was not the map on disk, opens a reference channel that had been failing on a single number, and builds an instrument that immediately overturns the decision which asked for it.

Sources: docs/01_SESSION_LOG.md S-020 to S-025 · experiments E-023 to E-034 · research R-014 · decisions D-043 to D-051 · docs/PLAN_2026-08-11_model-vs-method.md for the pre-registered predictions. Tooling written this day: verdict.py, make_corridor.py and its MK II and MK III successors. Forty-four renders were driven through the renderer’s HTTP API in about five minutes of GPU time when the usual connector was unavailable — worth remembering that the convenience layer is not a dependency.