Session S-027 · Experiment E-037 · 12 August 2026
We had one quality number, and it could not tell richness from disobedience. D-061 said nothing could be settled until a measure of obedience existed. So we built one, validated it against five cases with known answers, and pointed it at D-061’s own four renders. It overturned the decision that asked for it.
It also produced one number that was simply wrong, which is recorded below rather than quietly dropped.
Arms re-graded
4
Validation cases
5
Robustness settings
9
Decisions overturned
1
Claims withdrawn
1
Section 1
What the instrument actually measures
A depth map is not a picture, so its values cannot be compared against a plate’s. What transfers between them is discontinuity: wherever the map has a depth edge, the plate must show a visible edge in the same place. That single observation is the whole instrument.
Corrected for chance
Dilated edges cover a great deal of an image, so an uncorrected score flatters everything. Here 0.0 means no better than random and 1.0 means every commanded edge landed. Pure noise scores −0.012, not the 0.4 an uncorrected measure would have given it.
Two numbers come out of it, and they are deliberately separate. Adherence asks: of the strongest 3% of gradient pixels in the depth map, what fraction have a plate edge within about five pixels. Invention asks the opposite question — how much edge energy sits where the map was smooth, structure nobody asked for. Both masks are percentile-based, so neither rewards a plate merely for being busy. Adherence measures where structure is; invention measures how much of it was uncommanded.
The yardstick. Arms that differ in how depth was encoded must all be scored against one shared map. An inverse map and a linear map of the same room concentrate their gradient in different places, so scoring each arm against its own map would silently change the question being asked.
The legacy Laplacian detail number was kept alongside them, and it was recovered rather than guessed: mean absolute four-neighbour Laplacian of the channel-mean grey, interior pixels only. It reproduces D-061’s recorded 3.47 / 3.93 / 4.78 to three figures, so every old number on this project sits on the same scale as every new one.
Section 2
Five cases where the answer was known in advance
An instrument that has not been pointed at something with a known answer is not an instrument, it is an opinion with decimal places. Five cases, all passing.
| Case | Expected | Measured |
|---|---|---|
| A map scored against itself | about 1.0 | +1.000 |
| Plate B against its own map | well above chance, below 1 | +0.455 |
| Plate B against camera C’s map — same set, wrong view | well below the real pair | +0.245 |
| Plate B against its own map, shifted 40 px | below the real pair | +0.195 |
| Pure noise against a real map | about 0 | −0.012 |
python tools/obedience.py --selftest. The shifted-map and wrong-camera cases matter most: they prove the score falls off with misregistration specifically, rather than with image quality generally.Section 3
There is no trade to manage
Four renders of the MK III tunnel from camera B, verified one-variable from the PNGs’ own embedded workflows: seed 340001, strength 0.9, and a byte-identical 7,328-character prompt across all four. The only thing that differs is the depth map. D-061 had read the tight bracket’s lower detail as the model obeying more and inventing less — a trade you could dial. The instrument says otherwise.
| Arm | adherence | near | mid | far | invention | detail |
|---|---|---|---|---|---|---|
| wide linear — what we use today | +0.452 | +0.673 | +0.421 | +0.410 | 0.972 | 4.776 |
| scene linear (tight bracket) | +0.403 | +0.225 | +0.348 | +0.484 | 1.084 | 3.475 |
| scene inverse | +0.408 | +0.408 | +0.415 | +0.404 | 1.065 | 3.927 |
| wide inverse (the 58-level first attempt) | +0.185 | +0.160 | +0.085 | +0.250 | 1.019 | 4.505 |
The wide bracket is better on both axes at once. There is no beauty-versus-control trade here to manage.
D-062, superseding D-061’s central reading
The worry about detail counting invention was backwards
D-061’s fear was that a high Laplacian score might just be the model inventing freely. For these four arms the opposite holds: wide linear has the highest detail and the lowest share of off-geometry energy at 0.972. Its extra edge energy sits on commanded geometry. The plainer plates are the ones relatively busier where nothing was asked for.
Inverse depth helps the near field, not the far
This is the reverse of what R-015 predicted. At a matched bracket, inverse roughly doubles near-field adherence — 0.225 → 0.408 at tolerance 5, 0.102 → 0.272 at tolerance 3, 0.402 → 0.534 at tolerance 8 — and gives back a little in the far field. That is mechanically exactly what 1/z should do, since it spends its precision close to the camera. R-015 proposed inverse depth as the fix for Q-039’s far-field cliff. It is not that fix. Overall adherence between inverse and linear at matched bracket is a tie, +0.408 against +0.403: the 13.3% detail advantage D-061 measured is real and confirmed, but it buys no extra obedience.
Section 4
One claim withdrawn
A radial edge-energy profile taken about the vanishing point appeared to show something dramatic: the commanded hatch rim at r = 265 px rendered at 325 px by wide linear and at 125 px by scene linear, a collapse to half size. It would have been the most interesting number in the session.
Overlaying the commanded circle on the actual frames showed it was false. All three arms place the rim within roughly ±20% of commanded, and the estimator had locked onto the closed door behind the hatch — the strongest annulus of edge energy in that plate. A global scale sweep confirms it independently, peaking at exactly s = 1.00 for all three arms. The geometry is not rescaled at all.
The numbers were wrong and are not used.
It is recorded here because the failure mode is the recurring one on this project: an estimator that does not understand the geometry it is measuring, producing a confident number about the wrong object. It is the same class of mistake as comparing impressions across different prompts, which has twice produced conclusions that were entirely wrong.
Section 5
What is settled, and what is not
The instrument is validated. The findings about these particular arms are not yet general — this is one geometry, one camera, one seed. A tunnel is shallow and enclosed, and a depth measure has an easy time in it.
The arm that would settle it.
The same four bracket and ramp variants on a second locale with a different depth distribution — the ops hall, which is deep and open where the tunnel is shallow and enclosed. Same seed, one variable. Four renders. Three claims currently ride on a single geometry, which makes this the cheapest way to turn working practice into evidence.
Source: experiments/E-037_obedience-measure-and-the-d061-regrade_2026-08-12.md · instrument tools/obedience.py · sheets out/D061_arms_sheet.png, out/D061_ring_overlay.png. Status ⏳ in verification, pending Q-050.
