What the audio edits actually do
No instruction in add_and_remove.csv mentions sound — every forward
instruction is a visual removal. So what happens to the soundtrack has to be measured. This
section reports a scan of 4,246 pairs: all 3,418 rows whose removed entity is not a
person and whose reverse instruction quotes no speech, plus 400-row random control samples of
person removals with and without quoted speech.
The measure is gain-cancelled band energy. Targets are re-encoded at roughly
0.4–0.65× the original gain, so raw energy comparisons show a ~−4 dB drop
everywhere that has nothing to do with any edit. Subtracting the global level change leaves only
change in spectral shape — which is what "a sound source was removed" looks like.
Removing a speaking person does remove the speech. Whisper transcripts of 21 rows from
the quoted-speech block: the words quoted in <S>…<E> appear in the
original transcript (content-word overlap 0.6–1.0) and are gone from the target
(0.0–0.17). Word counts fall from 10–23 to 1.
But it is not a surgical voice removal — the whole soundtrack goes. Median
gain-cancelled speech-band change is +0.00 dB in both the speaking-person and
silent-person control pools: the spectral shape barely moves. What differs is the level.
Removing a speaking person raises the rate of a total audio wipe from 3.7% to 10.7%
(z = 3.85) and of a heavy level collapse from 14.1% to 28.7%. The speech
disappears as collateral damage from the track dropping out, not because the voice was
isolated and taken away.
No non-speech audio removals were found. If removing a dog took away its bark, the
non-person pool would lose more energy than a control pool where a silent person was removed.
It does not — the two are indistinguishable:
Within the non-person pool, entities that obviously make noise (animals, vehicles, machinery;
n = 1,368) lose less mid-band energy than silent objects like photographs and
signs (n = 151): 9.4% vs 17.5%. That is the opposite of a working mechanism. The
explanation is in the third number: 92% of all pairs have a regenerated soundtrack
(gain-fit residual > 0.25) regardless of what was removed. The pipeline re-synthesises audio
for nearly every edit, so band-level differences are regeneration noise, not evidence of a
source being taken out.
The audio unchanged cards below are the cleanest demonstration: a howling husky, a
howling beagle, a helicopter in flight, and a television broadcasting a speech — all removed
from the picture, all with a soundtrack that is provably the original at a lower gain (best-fit
residual 0.06–0.18). 102 rows in the non-person pool are like this.
Raw data: audio_scan_manifest.csv — all
4,246 scanned pairs with pool, entity label, bucket, per-band gain-cancelled deltas, RMS ratio and
gain-fit residual, so any threshold can be re-applied without rescanning.
transcripts.json — the 31 Whisper transcript pairs.
Correction to the earlier version of this page
A previous revision said the removed person's voice is "stripped along with the person" and
that speech-band energy dropped in "more than half" of 40 sampled rows. The count was
19 of 40, just under half. And the drop was measured in absolute energy, which cannot
distinguish a voice being removed from the whole mix getting quieter — the gain-cancelled
measurement above shows it is the latter. The six samples then labelled "audio removed" are
relabelled soundtrack collapsed: their targets retain 0.1%–22% of the original
audio level, and two are effectively silent.