Checkpoint 5 · Every clip has a soundtrack and nobody has ever listened properly

Two nearly identical frames of the man walking away up the farm track.

Checkpoint 5 · Three questions · headphones, about two minutes

Put headphones on for this one. Every clip this project has ever made came out with a soundtrack, and for the first three weeks nobody noticed — the sound plan was written, filed, and never connected to anything, because the engine takes one free-text box where the older one had a labelled field for audio.

Writing the soundscape into that box measurably changes the audio. What has never been checked is whether it produces the sound we asked for. That needs an ear, not a number.

Same shot, same random start, same picture instruction. The only difference is whether five sentences describing the sound were appended.

What was asked for, in these words: a steady cold wind moving across open ground high up, rising and falling and never gusting hard · a trough overflowing in a thin continuous run onto stone · cattle shifting and breathing inside the barn, a hoof on stone, a chain moving against a ring · a loose sheet of corrugated fibre-cement lifting and settling on a shed roof · two crows calling to each other a long way off across the valley, always at the same distance. And no music of any kind.

nothing writtenAttempt 1

Picture described, sound not mentioned at all. It still produces a soundtrack.

soundscape writtenAttempt 1

The five sentences above appended to the same instruction.

⚠ These play one after the other with sound — you cannot compare audio simultaneously.

nothing writtenAttempt 2

soundscape writtenAttempt 2

What the numbers say

Loudness, nothing written

−50.8 dB

Loudness, soundscape written

−45.7 dB

Variation between attempts, before

±3.7

Variation between attempts, after

±0.9

The louder figure is not the interesting one. The spread collapsing to a quarter of its width is. A side effect of a longer instruction would move the average; it would not make the result four times more consistent. That is what convinced me the words are reaching the audio at all — but “reaching it” and “getting what we asked for” are different claims, and only the first one has evidence.

What I need from you

1 · Can you hear specific things that were asked for? Not “does it sound like a farm” — can you identify the wind, the running water, cattle, the loose metal sheet, crows? Tick off whichever you actually hear.

What it changes: if most of the named sources are audible, sound becomes a real channel we can direct and the locked sound plan gets used on every shot. If you hear only generic wind, then we have a volume knob and not a sound department, and the plan should stop pretending otherwise.

2 · Is there music? The instruction ends with “no music of any kind”. Listen for anything tonal, droning or rhythmic underneath.

What it changes: we already know the model cannot be told not to render something visually — naming a thing puts it in the frame. If that holds for audio too, then negative sound instructions are actively harmful and every prompt on the project needs its “no music” line removed.

3 · Is this usable as a base layer, or is it scrap? Honest answer: would you build a mix on top of this, or throw it away and lay sound from a library?

What it changes: this decides whether sound is something we direct or something we replace in post, like the music and the voices already are. If it is scrap, that is fine — but I should stop spending prompt space on it.

Reply like CP-5: 1 wind and crows only, 2 no music, 3 base layer. This is the cheapest checkpoint to answer and possibly the highest-value, because sound is half of film and we have exactly zero human judgement on ours.

Honest limits

  • The picture changes too. Adding sentences changes the whole instruction, so the two clips do not show identical footage. You are judging the audio; ignore the picture differences.
  • Two attempts shown, three were measured. The third is on disk and adds nothing new by ear.
  • The audio is quiet in absolute terms — around −46 dB against a silence floor of −91. You may need to turn up more than feels normal. That low level is itself possibly a finding, and one nobody has investigated.