Day three · 10 August 2026 · sessions S-018 and S-019
Gabe watched the cut and said the galaxy looked mirrored between two scenes — then added that he might be wrong about why. He was right about the symptom and right to doubt the cause. Nothing was mirrored anywhere. Two entirely separate bugs were wearing one costume.
Then we found out why one room had never looked right in any shot ever made of it — and the answer had been sitting inside a PNG file for two days.
Bugs behind one symptom
2
Cameras right by luck
4 of 5
Boxes to rebuild a room
185
3D applications installed
0
Ray-caster, in lines
~120
Production strength
0.65
Bug one
Left and right belong to the camera, not to the room
The model resolves every bare left and right against the current shot’s frame, and it holds no persistent model of the set. So a phrase that reads like geography — “the left-hand end of the room” — is actually read as screen-left of this camera, and therefore names a different physical place in every differently-facing shot.
Across eighteen shots this was mostly invisible, for a reason that is worse than a straightforward bug: four of the five cameras faced forward, and when you face forward, screen-left simply is port. The convention was broken the whole time and returned the right answer anyway.
Then one camera turned ninety degrees to face the window wall, took “beyond him and to the left of the frame” literally, and put the second actor behind him instead of beside him.
Bug two
In one shot, the galaxy had nothing to copy
The continuity lock described the spiral as reading like a wide oval lying level. The galaxy in the actual reference picture is tilted about twenty degrees. Every render carrying that reference faithfully reproduced the reference — tilted.
Except one. The opening shot carried no hall reference at all, so the room glimpsed through the opening door came purely from the text — and the text won: level, warm-cored, off-centre. A “/” cut against a “\” reads unmistakably as a mirror.
This is one of our own rules firing where we did not expect it: a description that contradicts its reference beats the reference. Here it fired in the single render where there was no reference to beat.
And a third thing, found while checking the first two
One shot’s camera position and its own dramatic requirement contradicted each other inside a single sentence. The bible wanted the actor’s skull to eclipse the galaxy’s core — but that arithmetic treated the galaxy as a picture hanging on the far wall, rather than as a distant object whose direction from anywhere in the room is essentially straight ahead. From a camera at the starboard end of the window apron, he cannot cover it. It is not a hard shot. It is an impossible one.
What the model did with that is the finding. It did not fail, and it did not ignore the instruction. It traded — it dropped the eclipse and pulled both actors back off the glass so the window could sit screen-right, putting a console desk between them and the view. Nobody asked for that desk.
An impossible instruction is not ignored. It is traded against the instructions around it — and you pay in details you never specified.
Coincidental agreement is the most expensive kind of passing test.
When a rule is obeyed in most shots, check why it was obeyed. Four cameras here pointed the same way, so a broken convention returned the right answer four times and the wrong answer once — and that single wrong answer looked like a rendering defect rather than a specification defect. We spent a session chasing a mirror that did not exist.
The room that never looked right
The reference had been contradicting us since the day it was made
The operations hall had a recurring, unexplained problem: it kept rendering as a domed theatre with sweeping curved tiers, when the specification called for two straight parallel banks of consoles under a flat ceiling at 4.5 metres. The bible blamed a vague layout description. We had been rewriting prompts against it for two days.
Then I read the reference picture’s own metadata — because every image this pipeline generates carries the recipe that made it — and found the prompt that had produced it. It asks, in as many words, for “three sweeping tiers of crew stations” curving toward the glass, under a ceiling “six storeys high”.
The layout was never vague. It was in the picture. We had been attaching an image that argued against our own specification to every single render, and then blaming the words.



The fix, and Gabe’s objection to it
Draw the room as boxes and hand the model a depth map
If words cannot place a wall, geometry can. Render the ground plan as plain grey boxes, output the true depth at every pixel, and feed that to the image model as a spatial constraint. The renderer already shipped the necessary parts — no plugins, no compiler, nothing to install.
Gabe went straight to the weak point:
The depth map needs to already have all the meshes inside — you are making boxes instead of chairs and desktops, so the render will be boxes, no?
Half right, and the half that was right produced the working rule. He was correct that at high control strength the model follows the silhouette it is given. He was wrong about the mechanism: a depth map describes mass and distance, not form. The model is told where the volume sits and left to decide what fills it. These control models were also trained on depth estimated from photographs — which is soft and blobby — so a rough blockout is actually closer to what they expect than a crisp render would be. The documented risk is the opposite of his concern: rendered depth can be too sharp, and the fix is to blur it.
Whatever you want to be there goes in the blockout. Whatever you want the model to invent stays out.
Both halves were proved the same day. Chairs left out of the blockout came back as proper padded swivel chairs on chrome pedestals. Wall vents left out never appeared at any usable strength — because at high strength the model cannot ornament a wall the depth map declares to be a flat plane. Part of what we had admired as art direction in the old reference was never art direction at all. It was the absence of constraint.
The research recommended installing a full 3D application to make the depth maps. It was not needed. Axis-aligned boxes are about 120 lines of arithmetic — analytic ray-box intersection — and blockout_depth.py now turns a locale specification into a depth map, a plan view and a labelled framing thumbnail in one pass, with no 3D software, no GPU and nothing installed on the machine. Every pixel of look still comes from the image model. The blockout contributes no pixels; it only states where the masses are.
The best thing we learned all week
The desks were in the map. They were worth four grey levels
The first attempt half-worked. The camera sat correctly on the centre line, the window wall filled the far wall — and then the second bank of consoles simply refused to appear. It would have been easy to conclude the method had a ceiling.
Instead I dumped the actual numbers in the depth map, and the geometry was unmistakably there: clean discontinuities at both desks. The problem was not the geometry. It was how the range had been scaled.
The next question was whether any of this reaches the film, or only improves the library of reference pictures. Our standing rule said the reference governs appearance while the text governs layout — which would predict a correct plate changing very little.
It changed a great deal. Same seed, byte-identical prompt, only the plate swapped: the layout followed the plate. A V of desks along the side walls became one straight desk across the foreground, correct for a camera between the banks. The vaulted ceiling went flat with recessed panels. Oval portholes became the wall consoles the lock had always specified and text alone had never once produced.
The bill, in the same render
The plate also imported its own colour temperature — cooler and darker than the lock asked for, walls reading blue-grey against a specification of glossy pearl-white. Appearance obedience is not selective. A plate must be lit to the lighting plan, not merely built to the ground plan, or it quietly re-grades every shot that carries it.
The one real trade-off
The most beautiful setting is the wrong setting for a film
Gabe’s note on the new plates was fair: the rooms are pretty empty and naked. So we swept the control strength on one depth map at one seed, changing nothing else.
| Strength | Geometry | Dressing | Verdict |
|---|---|---|---|
| 0.90 | perfect | thin, bare | obedient and dull |
| 0.65 | fully intact | ornament returns — indicator lamps, vent grilles, cyan runs, deep floor reflections | the production default |
| 0.45 | the central aisle is gone | dense, genuinely beautiful | a mood board, not a set |



A sequence is more than one frame, so the setting that yields the single most beautiful image is the wrong setting for a film.
One asymmetry nobody predicted came out of the same sweep, and it has governed camera choices ever since: a 22° pan rendered correctly off-axis, while a 4-metre lateral dolly came back re-centred — as though the camera had never moved. The model reasserts its own framing against a translation but not against a rotation. A pan is safer than a dolly.
And the answer to “so we have to design everything now?”
That was Gabe’s real worry, and it deserved a real test. So the old reference picture — the beautiful, wrong, domed theatre — was transcribed by eye into 185 axis-aligned boxes. No depth estimator, no 3D application: curved window arc, curved soffit ring, three concentric tiers, all approximated with boxes. Then re-rendered with the original’s own prompt and seed.
The look survived. Three things follow, and together they are the answer. The design you must supply is coarse — 185 boxes, none under half a metre, no materials, no fixtures, no lights, with all the beauty still arriving free from the model. You do not have to invent it, because a look you already like can be absorbed in an afternoon. And the amount of design scales with what must stay the same between shots.
Gabe then checked that result too, and corrected it again: the rebuilt version at low strength has no roof. The map contains a soffit ring at 4.1 m and a flat lid at 5.6 m, and neither rendered — the top of frame opens into dark nothing. So it demonstrates a look transfer, not a set transfer: a large piece of stated geometry was silently dropped while the picture still looked good.
Below roughly 0.6, a depth map is a mood board, not a set.
A locale only counts as a set if a second camera placed inside it produces the same place. That is the whole test, and it is stricter than looking at one frame and liking it. What the boxes buy is not appearance — it is identity across cameras, which a single photograph of a room that exists nowhere else can never provide.
There is a structural lesson underneath all of this, which became the design of everything built afterwards. The failures kept producing rules that were universal, while the fixes kept being written as facts about one film. Separating those was the day’s quiet output: craft rules that apply to any film, an empty schema of the questions every shot must answer, and one place where proper nouns are allowed. What you harvest from a failure is a rule, not a fact — and then every past failure protects every film that comes after it.
Next: the long watch.
Day four asks whether it is the model or the method that limits us, discovers that 53% of one frame was empty, and runs a shot long enough to find out what a camera does when nobody is stopping it.
Sources: docs/01_SESSION_LOG.md S-018, S-019 · experiments E-016 to E-022 · research R-012, R-013 · decisions D-034 to D-042 · tooling tools/blockout_depth.py. Control-strength sweeps run on one depth map at one seed; every arm verified one-variable from the output files’ embedded workflows.