The drawing was lying, and we spent three sessions blaming the words

Two renders of the same arcaded street side by side. A lit shop window and a stone colonnade appear in both, but the covered walkway and the arched bays fall differently between them.

Session S-032 · Experiments E-058 to E-060 · 14 August 2026

For three sessions we have had the same complaint. We tell the machine which side of the street the covered walkway is on. It puts it on the other side. We tried saying it more firmly, we tried saying it in compass directions, we tried a clever new technique for pinning a description to a region of the frame. Nothing moved it.

Last night I finally looked at the drawing instead of the words, and the drawing was lying. Not the model, not the prompt — our own set, quietly telling the machine the opposite of what we were typing.

Attempts to put the walkway on the right side

9

That worked

0

Deleting seven boxes from the set

3 of 3

My own broken tools, found

3

Section 1

How the set was lying

A quick reminder of how this works. Before rendering a shot we build a crude three-dimensional model of the location out of plain boxes — walls, roofs, floors — and take a greyscale picture of it where brightness means how close. That grey picture is the strongest instruction the machine gets. The written description is the weaker one.

Our Italian street has a covered arcade down one side: shops set back behind a row of stone pillars carrying the upper floors out over the pavement. The other side is open pavement with a few shop awnings sticking out above the windows.

And there it is. An awning is a flat thing sticking out over a pavement. An arcade roof is a flat thing sticking out over a pavement. In a picture that only knows distance, they are the same object.

Two nearly identical grey depth drawings of a street. The left one has flat horizontal slabs jutting out at the upper left; the right one does not.
The same street, drawn twice. Left: our set as built — see the flat slabs jutting out over the pavement at the top left. Those are shop awnings, two metres from the lens. Right: the same drawing with seven boxes deleted.

I measured how loudly each side was speaking. The arcade pillars take up 7.3% of the frame. The awnings take up 7.3% of the frame — the same. But the awnings stick out nearly twice as far. The side that is supposed to have no cover was making the stronger case for being the covered one.

Nothing you type argues with the drawing. If the drawing says something, the drawing wins.

Section 2

Deleting seven boxes beat everything we had tried

So I took the awnings out of the set and changed nothing else. Same words, same everything.

Two street photographs from the same viewpoint. Left: the camera stands under a roofed walkway with a ceiling and lights. Right: the pavement is open to the sky, with arched shop bays cut into the building face.
Left: the set as it was. The camera is standing under a roofed walkway with a ceiling and downlights — on the side the script says is open. Right: identical words, identical seed, awnings deleted. The frontage still has arched shop bays, but the footway itself is now open to the sky. That is the change we had been asking for across three sessions.

Three attempts each, so it is a measurement rather than a lucky picture:

  • Words alone — right 1 time out of 3.
  • The clever new region-pinning technique — 2 out of 3.
  • Deleting seven boxes from the set — 3 out of 3.

Fixing the drawing beat both the writing and the clever technique. So there is now a running order, and we had it backwards: drawing first, set second, clever technique third, words last.

I should be straight about the other half. Deleting the awnings stopped the wrong side being covered. It never got the arcade onto the right side — nine renders, three different approaches, not once. From across the street you are looking at that arcade edge-on, and edge-on a row of pillars is just a wall. My guess is you have to put the camera somewhere it can see under the thing, and that is the first job next time.

Section 3

The good news: keeping a picture and changing one building

Buried in the software there is an input nobody on this project had ever used. You hand it a finished picture and paint over the part you want redone. Everything you did not paint comes back as the same picture.

Two photographs of the same farmyard. On the left an old stone and timber barn; on the right the same yard with the barn replaced by a modern green steel shed. Everything else is identical.
The same farmyard twice. I painted over the barn and asked for a modern steel shed instead. The wheelbarrow, the blue barrels, the bales, the trestle, the wall, the terraces, every rut in the mud — untouched, because they were never regenerated.

This matters more than it looks. Every continuity trick we have had until now is a way of describing something the same way twice and hoping. This one hands over the actual pixels. It is the first thing we have that makes two shots the same place rather than a very similar place.

It is not a perfect photocopy — surfaces get re-rendered, so grain and stonework shift slightly. But nothing moves, and what breaks a cut is a bucket jumping across the yard, not grain.

Section 4

Two smaller answers

Does a fixed still stay fixed once it becomes video? Yes — six clips out of six. I drove the video from a good still and from a deliberately broken one, and each clip faithfully carried whatever its still had, including the mistake. That is the answer we wanted: get the still right and the clip inherits it.

Can we ask for a blurred background? Barely. We had never once asked in six weeks. Both directions nudge correctly — focus near, focus far — but by about a third of what would count as an effect, and by eye the picture is still sharp from the kerb to a monument a hundred metres away. Treat a soft background as something we add afterwards, like music and voices.

Section 5

Three of my own tools were broken, all the same way

A checking tool crashed at exactly the moment it had bad news — the one line it exists to print used a symbol the console could not display. A second one, which is supposed to nag me when a settings page goes stale, compared dates in a way that made every decision written that same day invisible to it. A third quietly ignored two rows of our capability table, so the honest counts at the bottom of that table were wrong.

None of the three was wrong. All three were silent. That is the harder fault to spot, because silence looks exactly like good news.

And one mistake that was purely mine: I designed an experiment my own measuring tool physically could not measure, and only noticed two renders in. I had carefully checked the instrument worked before using it. I had not checked it could see the thing I was pointing it at.

Twice last night a first attempt told a tidier story than three attempts did — once about the blurred background, once about the copying trick. Both times the second and third attempts pulled the number down. It is a good argument for a rule we already have and keep being tempted to skip: one render is an observation, three is a measurement.

Where this leaves us

Three sessions of blaming the writing, and the answer was in a drawing we made ourselves. The lesson generalises past this street: before arguing with the machine, check that our own inputs agree with each other. It was doing what we told it. We were telling it two different things and only listening to one of them.

Next up: move the camera somewhere it can see under the arcade, and settle whether our written instructions were ever useless or simply shouted down by the drawing. If it turns out to be the second, a lot of what we think we know about writing these prompts needs re-reading.