Two locations, and three rules we already had

A stone farm steading seen from above: a slate-roofed barn with black silage bales stacked under its eaves, a house beyond it, blue barrels against the wall, and deep mud tracks curving through the yard.

Session S-029 · Experiments E-044 to E-047 · 13 August 2026

We built two places from scratch — a mountain farm in late winter and a shopping street in an Italian old town — and shot in both. The pictures are the best this project has made. Three things went wrong, and all three were rules we had already written down and then did not follow.

The fourth session in a row to find a bug in its own measuring instrument. And the one question the session was set — can you have the exact set and the exact face in the same shot — has an answer, and it is no.

Locations built

2

Camera positions

13

Drawing faults caught before rendering

11

Rules we already had and ignored

3

Instruments repaired

1

A pan across the yard at Cascina Corvara. Every set this project had built before today was a spaceship corridor. This one is a list of 93 boxes with coordinates in metres — the buildings, the walls, the water trough — and everything else in the frame, the mud and the snow lying only in the shade and the bare trees on the fellside, comes out of a paragraph of description. The man was not asked for in this shot. He arrived anyway, in a jacket nobody specified, and that turns out to be the session’s central problem.

Section 1

Reading the manual, one session too late

The previous session characterised the image model across roughly a hundred and twenty renders and never once read its documentation. This one read it first. Most of what came back was reassuring: the authors’ recommended control-strength band is 0.65 to 1.00, and we had independently measured 0.70 as the best setting and found the picture collapsing at 0.60 — immediately below their published floor, discovered without knowing the floor was there.

Better than agreement, the documentation explained something we had only observed. Our fast model ignores negative instructions and our slow one obeys them, and we had recorded that as a quirk. The reason is that the fast model has the guidance mechanism baked in during training, so at render time there is no branch for a negative instruction to act on. It does not ignore them — there is nowhere to put one.

Where the documentation is simply silent

We hand the renderer a greyscale picture in which brightness means distance. No maker, anywhere, publishes how that picture is supposed to be encoded — not which end of the scale is near, not whether it should be linear in metres, not which tool produced the training examples. Two of our decisions rest entirely on our own measurements and no amount of further reading will settle them. Worth knowing which parts of a practice are load-bearing and unverifiable.

And one correction I owe the record. Halfway through writing the research file I found that a research file from the previous session had already found the most important thing in it — that our descriptions run three times longer than the model was ever tested on — and had made testing it its number-one recommendation. Nobody ran it. I nearly presented it as a discovery. The useful fact is not that the problem exists; it is that we knew and did not test it.

A stone hill farm in late winter seen from above, terraced fields behind
The establishing shot of the farm. It took four camera positions to find, all of them rejected by looking at the grey-box drawing rather than by rendering — because the yard is eleven metres wide between buildings six to eight metres tall, and from anywhere inside it no camera can hold both buildings and the horizon.
Inside a stone byre looking out through the open door to a muddy yard
The hardest single frame in the set: it shows one place through another. The yard, the house wall and the outdoor stone stair are all visible through the doorway, and a viewer compares them directly against the outdoor shots.

Section 2

What the grey boxes can and cannot do outdoors

Every set this project had built was a corridor, eight to thirteen metres end to end. A farm runs from a wall two metres away to a valley two hundred metres away, and the method was built for boxes. Two things came out of stressing it, and both are arithmetic rather than opinion.

Minimum relief that registers = 8 × (far − near) ÷ 255 metres. The depth picture spends 255 shades of grey across whatever distance range you give it, and anything under about eight shades is lost. In a corridor that is 26 centimetres, so latches and hoses and panel lines all register. On the street’s long view it is 5.2 metres.

So at that scale a window reveal does not exist. Nor does a balcony, a cornice, a sill or a door surround — every one of them is under a metre deep. Modelling the street’s façades in detail would have cost several hundred blocks and changed nothing at all. They belong in the words.

But blocking still registers, and that is the trick. Relief on a distant surface is worthless; a thin object standing in front of a distant surface is worth its whole depth step. A lamp column nine inches thick, fifteen metres from the lens against a façade thirty metres away, is a fifteen-metre step — about twenty-nine shades, four times the floor. So a large set is massing plus silhouettes: building lines, columns, bollards, islands, monuments. Nothing that merely bumps out of a wall.

The second finding reversed a decision we had thought settled. There are two ways to write the depth picture, and a previous session compared them on a corridor and found no difference. That result was right, and it was a fact about corridors — on a nine-metre range the two methods are very nearly the same function. On the farm, one of them gives the buildings 143 shades and the other gives them 28. On the street, the answer reverses, because there the thing that matters is at the far end and the near foreground is empty road. Neither method is correct in general. The rule is: put the shades where the detail you care about is.

Section 3

Three failures, all of them ours

Writing the rulebook is not following it

We have a rule that says: before writing any prompt, give every object in the set one fixed sentence, and paste that exact sentence into every shot the object appears in. Not a similar sentence — the same one, pasted. Paraphrase is how the door changes.

The farm was built to that rule. Every object got its locked sentence, written before anything was drawn, in a proper register. Then six pictures were rendered from paragraphs that described the same objects in their own words, and not one locked sentence was ever pasted in. The result was exactly the failure the rule exists to prevent: the place holds beautifully across all six cameras — the light, the palette, the materials are unmistakably one farm — and the objects wander. The outdoor stone staircase moves from the house to the barn. The roofs change material. The two indoor shots are two different barns.

The register is not the artefact

The pasted sentence is the artefact. Writing the register and then prompting from memory produces the drift the register exists to stop, and it looks like diligence while it is happening. That is the part worth remembering.

A description of a place cannot say “on the left”

Our rule for screen direction is strict: anything belonging to the set gets a compass fact and its translation into this camera’s frame, in the same sentence. The street’s description says the arcade is “on one side” and the open pavement “on the other”. It never says which side of the picture, and it cannot — a description of a place has no camera in it.

The arcade came out on the wrong side. The fix is structural rather than verbal: the place description stays general, and each camera adds one paragraph describing the place from here. That paragraph is then the only part that differs between two shots of the same set.

Saying a thing is off-screen does not keep it off-screen

That per-camera paragraph also stated, in as many words, that there was no fountain anywhere in the frame — the fountain being at the other end of the street, behind the camera. The fountain rendered anyway, standing in front of the obelisk at the far end, because the place description still described it in loving detail as the far end’s monument.

We had a rule that a thing which must not appear must not be named, and a later session lifted it for the slower model. This narrows that: it was lifted for things that are in shot, and not for things that are not. A named thing gets rendered somewhere. The only reliable way to keep something out of frame is to not mention it.

Looking down an Italian shopping street to a stone obelisk in a circus
With the fountain deleted from the description rather than declared absent, the obelisk arrives — tall, on its plinth, closing the street. The reverse shot down the same pavement correctly shows no obelisk at all.

And the one I nearly filed as a failure

The arcade is drawn on the east side of the street. Looking north it came out on the left of the picture; looking south it came out on the right. Read carelessly, both are wrong. Worked through the camera geometry properly, both shots put the arcade on the same physical side — the model chose the opposite side to the one I asked for, and then held it across a reverse angle.

Consistency beats obedience

It does not obey which side. It does stay on the same side between two opposed cameras. For cutting, only the second matters — what breaks a scene is a set that changes between shots, not a set mirrored from the plan. This street cuts. So the rule is: take the handedness from the first picture and rewrite the plan to match. Fighting it costs renders and loses.

Inside a stone arcade with vaulted ceiling and lit shopfronts
Under the arcade. None of the façade detail here — no window reveal, no cornice, no shutter — exists in the depth drawing at all; at this scale they are all far below the threshold at which the drawing can carry them. Every bit of it comes from the words.

Section 4

The instrument, checked before it was believed

One experiment needed to know whether a detail described near the end of a very long description lands as reliably as the same detail near the start. The test object was a bright red fuel can, in two otherwise colourless palettes, and the measure was simply how much red is in the frame.

Run against five pictures that contained no red fuel can at all, that measure returned between 0.19 and 0.35 per cent — rust primer on a harrow, a no-entry sign, terracotta plaster. A fuel can occupying a few tenths of a per cent of the frame would have sat inside the noise, and the experiment would have measured nothing while appearing to measure something.

The fix follows from what the object actually is: the test object is one compact thing and the noise is scattered, so the statistic is the largest connected red region, not the total. Re-measured on the same five pictures, the floor dropped to between zero and 0.043 per cent — about eight times tighter. This is the fourth session running to find a bug in its own ruler, which is either a good habit or an uncomfortable pattern, and probably both.

Section 5

The set or the face

There are two ways to turn a still into a moving shot. One takes the picture and uses it as the literal first frame, so the place is exact to within a fraction of a decibel — and it accepts no photograph of a person. We confirmed that from the software’s own source rather than by inference: the node has inputs for a first frame and a last frame and nothing else. The other route takes the face photographs and invents its own room.

The obvious way out is to make a still that already contains the actor and use that as the first frame. It does not work, and the reason is worth knowing. Handed a photograph of a face and asked to place that man in a scene, the still renderer does not compose the two — it pastes the photograph into the picture. The face arrives as a ghost tiled bodily across a third of the frame. Two attempts, one with three reference pictures and one with two, both did the same thing.

A failed composite: a face photograph tiled bodily into a farm scene
What happens when you hand the still renderer a photograph of a face and ask it to put that man into a scene. It does not compose them — it pastes the photograph into the picture. Two attempts, one with three reference pictures and one with two, both did this.

What that failure did show is that the halves come apart cleanly: the man it painted alongside the ghost wears exactly the clothes specified, standing in the farm’s own mud and snow and light. Appearance transfer works; spatial composition does not.

The obvious answer is editorial: shoot the wide through the first-frame route, shoot the close through the reference route, cut between them. That is what I wrote first. Then I shot eight moving shots across both places and measured whether they cut together, and it turns out to be exactly wrong.

One route, brightness spread

2.0

Mixed routes, same shots

24.6

Best continuity we had ever managed

2.6

Mixing the two routes inside one scene made the brightness wander twelve times as much as staying on one. A scene shot entirely on either route sits right at the best continuity this project has ever achieved. A scene that mixes them does not cut at all.

Choose the route per scene, not per shot

A scene that needs the exact place goes all first-frame, and you accept that the performer is out of your hands. A scene that needs his face goes all reference, and you accept an invented room. Change route only at a scene break, where the look is allowed to change anyway.

And the invented rooms are not random. Because the place is described so thoroughly in words, the three reference-route street shots read as unmistakably one street and cut together perfectly well. They simply are not this street.

That correction is the most useful thing in the session, and it only exists because the shots were measured against each other rather than admired one at a time. The plausible answer and the right answer pointed in opposite directions.


Section 6

Eight shots, and only one thing in them was directed

The matrix that produced that correction was eight moving shots across both places, inside and out, four on each route, covering a static frame, a pan, an arc, a pedestal, a push in and a truck. Every move landed. That sounds like a win, and it hides the finding.

Of the five things a director actually decides — where the camera stands, which way it points, how it moves, where the actor stands and where he goes — the words command exactly one. The camera’s position and angle come from whichever still the shot is grown from; the model’s movement vocabulary has no verb for either, and no amount of writing supplies one. Where the actor stands is a request he may decline, which was measured a session ago. The move is the only instruction. A shot list on this pipeline is really a list of plates, each with one verb attached.

A push in under the arcade — turn the sound on. The footsteps and the room tone were invented by the video model in the same pass as the picture; nothing was recorded and nothing was added afterwards. The vaults, the piers, the shopfronts and the wet setts are all in the words — at this distance the depth drawing carries none of them. And the man is wearing a dark suit, where his character sheet says a charcoal overcoat.
BRIGHTNESS SPREAD ACROSS THE SHOTS OF ONE SCENE · LEVELS OUT OF 255 2.6 — our best-matching scene before today farm, one route 2.0 farm, routes mixed 24.6 street, one route 8.3 street, routes mixed 16.1 0 25 lower is better — a scene whose shots do not match cannot be cut
Eight shots, grouped two ways. Shot on one route, a scene at the farm holds tighter than anything this project has made before; mix the two routes inside the same scene and the brightness wanders twelve times as far. The street is less extreme because its two routes happened to expose more alike, which is luck rather than a property.

Two other things came out of the eight, and both were unanimous rather than marginal. The reference route dressed the actor correctly in all four of its shots, and the first-frame route dressed him wrongly in all four of its own — a blue quilted jacket up at the farm, a dark suit under the arcade, against a character sheet asking for olive waxed cotton and a charcoal overcoat. And the first-frame route is 2.1× cheaper: 352 seconds against 731 at full size, which is the price of carrying reference photographs turning up exactly where an earlier measurement said it would.

So the honest summary of the collision is narrower than “the set or the face”, and more useful than it. The exact place, cheap, with a performer who turns up when he likes wearing what he likes. Or the right man in the right clothes, at twice the price, in a room that resembles yours without being it. Choose per scene. Never inside one.

What happens next: the input nobody has ever used.

The first-frame node takes a last frame as well, and in twenty-nine sessions this project has never once supplied one. It is the obvious instrument for joining two shots end to end, and the best candidate for fixing the first-frame route’s real defect — that the performer arrives unbidden and dressed by the model. After that: paste the locked sentences into the prompts that need them, a rule we have had in writing for days and have still never actually followed. The previous session is here; why a single render proves less than it looks is here.

Sources: docs/01_SESSION_LOG.md S-029 · docs/AUTORUN_progress.md · research R-016 · experiments E-044 to E-048 · decisions D-080 to D-085 · experiments/RESULTS_LEDGER.md. The two locations are tools/make_mountain_farm.py (93 primitives, six cameras) and tools/make_old_town_street.py (130 primitives, seven cameras). The ramp and prompt-length arms are three seeds each and are quoted as means; the register, geography and handedness results are one seed per arm — observations rather than measurements, consistent in direction and settled by nothing. Brightness spreads are measured across the shots of one scene, in levels out of 255.