InsAVE-80K — input / output sample viewer

29 pairs from the instruction-based audio-video editing dataset released with InstructAV2AV: 5 from the held-out eval split and 24 from train — including 20 drawn from add_and_remove.csv. 7 of those are edits where the soundtrack collapses or vanishes, and 4 are sound sources — a howling dog, a helicopter, a broadcasting television — removed from the picture while the audio stays put. A sample is a source clip, an edited clip, and a forward / reverse instruction pair. The clips keep their original audio — the edits are audio-visual, so play them with sound on.

88,074 pairs (87,074 train / 1,000 eval)
176,148 video files
1280×704 · 24 fps · 3–7 s
mp3 44.1 kHz stereo
~139 GB 11 tar shards

What the audio edits actually do

No instruction in add_and_remove.csv mentions sound — every forward instruction is a visual removal. So what happens to the soundtrack has to be measured. This section reports a scan of 4,246 pairs: all 3,418 rows whose removed entity is not a person and whose reverse instruction quotes no speech, plus 400-row random control samples of person removals with and without quoted speech.

The measure is gain-cancelled band energy. Targets are re-encoded at roughly 0.4–0.65× the original gain, so raw energy comparisons show a ~−4 dB drop everywhere that has nothing to do with any edit. Subtracting the global level change leaves only change in spectral shape — which is what "a sound source was removed" looks like.

Removing a speaking person does remove the speech. Whisper transcripts of 21 rows from the quoted-speech block: the words quoted in <S>…<E> appear in the original transcript (content-word overlap 0.6–1.0) and are gone from the target (0.0–0.17). Word counts fall from 10–23 to 1.

But it is not a surgical voice removal — the whole soundtrack goes. Median gain-cancelled speech-band change is +0.00 dB in both the speaking-person and silent-person control pools: the spectral shape barely moves. What differs is the level. Removing a speaking person raises the rate of a total audio wipe from 3.7% to 10.7% (z = 3.85) and of a heavy level collapse from 14.1% to 28.7%. The speech disappears as collateral damage from the track dropping out, not because the voice was isolated and taken away.

No non-speech audio removals were found. If removing a dog took away its bark, the non-person pool would lose more energy than a control pool where a silent person was removed. It does not — the two are indistinguishable:

pooln"selective non-speech" rate median worst non-speech bandmid-band loss ≤−8 dB
non-person removed (a sound source could go)3,418 14.0%−1.5 dB9.3%
silent person removed (nothing should change)404 14.6%−1.6 dB9.7%

Within the non-person pool, entities that obviously make noise (animals, vehicles, machinery; n = 1,368) lose less mid-band energy than silent objects like photographs and signs (n = 151): 9.4% vs 17.5%. That is the opposite of a working mechanism. The explanation is in the third number: 92% of all pairs have a regenerated soundtrack (gain-fit residual > 0.25) regardless of what was removed. The pipeline re-synthesises audio for nearly every edit, so band-level differences are regeneration noise, not evidence of a source being taken out.

The audio unchanged cards below are the cleanest demonstration: a howling husky, a howling beagle, a helicopter in flight, and a television broadcasting a speech — all removed from the picture, all with a soundtrack that is provably the original at a lower gain (best-fit residual 0.06–0.18). 102 rows in the non-person pool are like this.

Raw data: audio_scan_manifest.csv — all 4,246 scanned pairs with pool, entity label, bucket, per-band gain-cancelled deltas, RMS ratio and gain-fit residual, so any threshold can be re-applied without rescanning. transcripts.json — the 31 Whisper transcript pairs.

Correction to the earlier version of this page

A previous revision said the removed person's voice is "stripped along with the person" and that speech-band energy dropped in "more than half" of 40 sampled rows. The count was 19 of 40, just under half. And the drop was measured in absolute energy, which cannot distinguish a voice being removed from the whole mix getting quieter — the gain-cancelled measurement above shows it is the latter. The six samples then labelled "audio removed" are relabelled soundtrack collapsed: their targets retain 0.1%–22% of the original audio level, and two are effectively silent.

At a glance

How much of the frame the edit actually changed (share of pixels with |Δ| > 30, averaged over 6 matched frames) and how much the audio changed (envelope correlation — near 1.0 means the soundtrack is essentially untouched, near 0 means it was regenerated), and the change in 300–3400 Hz energy — the column that tells you whether a voice was actually taken out. Anything at or below −10 dB is highlighted. Click an id to jump to the sample.

idsplitcategoryinstruction (forward) pixels changedaudio corr speech-band Δ
00000evalevalReplace the woman with a medium-sized brown dog walking forward with a white rolling suitcase b…3.0%+0.38+9.8 dB
00001evalevalReplace the person in the wheelchair with a grand piano made of dark polished wood with gleamin…3.7%+0.45+1.1 dB
00002evalevalTransform the fishing boat into a steam locomotive with a weathered, rust-streaked metal body,…11.9%-0.03+12.1 dB
00003evalevalReplace the person with a robot with a metallic, angular frame and glowing joints, wearing a sl…10.2%+0.46-3.8 dB
00004evalevalReplace the car with a motorcycle.0.2%+0.12+16.4 dB
00000trainadd_and_removeRemove the man in the black t-shirt standing on the beach near the water.50.4%+0.94-4.1 dB
00006trainadd_and_removeRemove the boy standing in front of the red wall with drawings.18.1%+0.43-6.4 dB
00112trainadd_and_removeRemove the framed photograph of the young boy from the table.7.9%+0.70-6.5 dB
00167trainadd_and_removeRemove the young woman holding the book from the center of the scene.23.5%+0.54-4.7 dB
00217trainadd_and_removeRemove the young boy holding the bottle in the center of the frame.15.6%+0.97-4.6 dB
00504trainadd_and_removeRemove the photograph of the couple embracing on the boat.10.9%+0.87-3.9 dB
00539trainadd_and_removeRemove the transparent digital display screen showing "ACTIVE ACCESS" and the numbers, includin…2.8%+0.79-4.9 dB
00688trainadd_and_removeRemove the figure in the red cloak riding the black horse from the cliff edge on the left.0.6%+0.98-3.7 dB
00955trainadd_and_removeremove the shirtless man climbing the tree3.5%+0.60-4.1 dB
01111trainadd_and_removeRemove the boy standing next to the tree on the left side of the frame.3.8%+0.06-29.4 dB
05451trainadd_and_removeremove the husky dog lying on the bed with its head tilted upward and howling.10.8%+0.99+0.0 dB
06527trainadd_and_removeremove the beagle dog standing on the couch and howling.17.8%+0.99+0.0 dB
08364trainadd_and_removeremove the yellow and black helicopter flying low over the desert canyon0.2%+0.97+0.0 dB
10993trainadd_and_removeRemove the woman with curly brown hair wearing a brown and white patterned outfit and a blue be…6.0%+0.01-77.9 dB
11593trainadd_and_removeRemove the man in the black leather jacket holding a microphone from the center of the stage.11.3%+0.00-68.5 dB
14593trainadd_and_removeRemove the woman seated on the dark couch, wearing a red top and black blazer, from the center…13.2%-0.00-28.4 dB
15493trainadd_and_removeRemove the woman with long dark hair showing a distressed expression from the center of the fra…19.4%+0.05-24.5 dB
16093trainadd_and_removeRemove the man in the red suit standing on the stage holding a microphone.4.9%+0.06-11.1 dB
19764trainadd_and_removeRemove the television screen displaying the C-SPAN broadcast of President Obama speaking at the…4.0%+0.98+0.0 dB
20293trainadd_and_removeRemove the man on the left wearing a black hoodie with a yellow graphic.6.7%+0.13-33.8 dB
00000trainclone_idKeep the person's appearance, change the timbre to a man, and change the spoken words to I thin…1.1%+0.07+7.0 dB
00000trainclone_id_voiceKeep the person’s identity and change the spoken words to I think we need to talk about this ri…1.1%-0.01+11.9 dB
00000trainclone_voiceKeep the timbre, change the person to a man wearing dark jeans, and change the spoken words to…5.3%-0.01+11.9 dB
00000traingeneral_editingReplace the man with a woman in her late 50s or early 60s, with graying hair and a beard, weari…2.7%-0.03+0.2 dB
eval eval eval/00000_original.mp4
1280×704 · 23.98 fps · 3.42s / 3.42s · mp3 44kHz stereo
instruction →

Replace the woman with a medium-sized brown dog walking forward with a white rolling suitcase beside it, mimicking a human pulling luggage, and modify the man's expression to show mild surprise.

← instruction_reverse

Replace the medium-sized brown dog with a woman with long hair wearing a dark coat, pulling a white suitcase.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
6.68mean pixel Δ (0–255)
3.0%pixels changed (Δ>30)
106p99 pixel Δ
+0.38audio envelope corr.
0%audio level kept (-)
-gain-fit residual (<0.25 = same audio)
eval eval eval/00001_original.mp4
1280×704 · 24 fps · 5.04s / 5.04s · mp3 44kHz stereo
instruction →

Replace the person in the wheelchair with a grand piano made of dark polished wood with gleaming ivory keys.

← instruction_reverse

Replace the grand piano with a person seated in a wheelchair dressed in a light-colored outfit with a patterned design.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
5.68mean pixel Δ (0–255)
3.7%pixels changed (Δ>30)
103.7p99 pixel Δ
+0.45audio envelope corr.
0%audio level kept (-)
-gain-fit residual (<0.25 = same audio)
eval eval eval/00002_original.mp4
1280×704 · 23.98 fps · 5.09s / 5.09s · mp3 44kHz stereo
instruction →

Transform the fishing boat into a steam locomotive with a weathered, rust-streaked metal body, a tall smokestack emitting faint white steam, and an American flag mounted at its rear, retaining the silhouette and posture of the original boat but adapted as a vintage steam train moving along a narrow waterway.

← instruction_reverse

Replace the steam locomotive with a fishing boat equipped with nets and crabbing equipment.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
12.92mean pixel Δ (0–255)
11.9%pixels changed (Δ>30)
129p99 pixel Δ
-0.03audio envelope corr.
0%audio level kept (-)
-gain-fit residual (<0.25 = same audio)
eval eval eval/00003_original.mp4
1280×704 · 23.98 fps · 6.71s / 6.71s · mp3 44kHz stereo
instruction →

Replace the person with a robot with a metallic, angular frame and glowing joints, wearing a sleek, silver exoskeleton.

← instruction_reverse

Replace the robot with a person with curly hair wearing a light-colored shirt.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
12.11mean pixel Δ (0–255)
10.2%pixels changed (Δ>30)
152p99 pixel Δ
+0.46audio envelope corr.
0%audio level kept (-)
-gain-fit residual (<0.25 = same audio)
eval eval eval/00004_original.mp4
1280×704 · 23.98 fps · 3.42s / 3.42s · mp3 44kHz stereo
instruction →

Replace the car with a motorcycle.

← instruction_reverse

Replace the motorcycle with a car.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
2.68mean pixel Δ (0–255)
0.2%pixels changed (Δ>30)
11.3p99 pixel Δ
+0.12audio envelope corr.
0%audio level kept (-)
-gain-fit residual (<0.25 = same audio)
train add_and_remove train/add_and_remove/00000_original.mp4
1280×704 · 23.98 fps · 5.09s / 5.09s · mp3 44kHz stereo
instruction →

Remove the man in the black t-shirt standing on the beach near the water.

← instruction_reverse

Add a man in a black t-shirt standing on the beach near the water, looking downward with a contemplative expression.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
89.68mean pixel Δ (0–255)
50.4%pixels changed (Δ>30)
208.7p99 pixel Δ
+0.94audio envelope corr.
61%audio level kept (normal)
0.231gain-fit residual (<0.25 = same audio)
train add_and_remove train/add_and_remove/00006_original.mp4
1280×704 · 25 fps · 3.27s / 3.27s · mp3 44kHz stereo
instruction →

Remove the boy standing in front of the red wall with drawings.

← instruction_reverse

Add a boy with short brown hair, wearing a plaid shirt over a white t-shirt, standing in front of the red wall with children's drawings.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
14.88mean pixel Δ (0–255)
18.1%pixels changed (Δ>30)
110.7p99 pixel Δ
+0.43audio envelope corr.
48%audio level kept (normal)
0.722gain-fit residual (<0.25 = same audio)
train add_and_remove train/add_and_remove/00112_original.mp4
1280×704 · 23.98 fps · 3.42s / 3.42s · mp3 44kHz stereo
instruction →

Remove the framed photograph of the young boy from the table.

← instruction_reverse

Add a framed black-and-white photograph of a smiling young boy with curly hair, wearing a collared shirt, with handwritten text "To dad, Love John" at the bottom, positioned on the right side of the table.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
8.97mean pixel Δ (0–255)
7.9%pixels changed (Δ>30)
143.3p99 pixel Δ
+0.70audio envelope corr.
53%audio level kept (normal)
0.626gain-fit residual (<0.25 = same audio)
train add_and_remove train/add_and_remove/00167_original.mp4
1280×704 · 23.98 fps · 6.71s / 6.71s · mp3 44kHz stereo
instruction →

Remove the young woman holding the book from the center of the scene.

← instruction_reverse

Add the young woman holding the book in the center of the scene, positioned in front of the window with curtains, facing slightly to the right.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
28.39mean pixel Δ (0–255)
23.5%pixels changed (Δ>30)
196p99 pixel Δ
+0.54audio envelope corr.
58%audio level kept (normal)
0.695gain-fit residual (<0.25 = same audio)
train add_and_remove train/add_and_remove/00217_original.mp4
1280×704 · 23.98 fps · 5.09s / 5.09s · mp3 44kHz stereo
instruction →

Remove the young boy holding the bottle in the center of the frame.

← instruction_reverse

Add a young boy with short dark hair, wearing a light-colored t-shirt, holding a small bottle, positioned in the center of the frame, looking upward with a thoughtful expression.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
11.82mean pixel Δ (0–255)
15.6%pixels changed (Δ>30)
74.7p99 pixel Δ
+0.97audio envelope corr.
62%audio level kept (normal)
0.158gain-fit residual (<0.25 = same audio)
train add_and_remove train/add_and_remove/00504_original.mp4
1280×704 · 23.98 fps · 5.09s / 5.09s · mp3 44kHz stereo
instruction →

Remove the photograph of the couple embracing on the boat.

← instruction_reverse

Add a photograph of the couple embracing on the boat, positioned in the center of the collage, with the man in a dark jacket and the woman in a red coat, against a watery background with a wooden railing.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
9.19mean pixel Δ (0–255)
10.9%pixels changed (Δ>30)
82.7p99 pixel Δ
+0.87audio envelope corr.
60%audio level kept (normal)
0.396gain-fit residual (<0.25 = same audio)
train add_and_remove train/add_and_remove/00539_original.mp4
1280×704 · 23.98 fps · 5.09s / 5.09s · mp3 44kHz stereo
instruction →

Remove the transparent digital display screen showing "ACTIVE ACCESS" and the numbers, including the faint reflection of the person behind it.

← instruction_reverse

Add a transparent digital display screen with blue text reading "ACTIVE ACCESS" and a list of numbers, positioned in the center of the scene with a faint reflection of a person behind it.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
4.4mean pixel Δ (0–255)
2.8%pixels changed (Δ>30)
38p99 pixel Δ
+0.79audio envelope corr.
59%audio level kept (normal)
0.464gain-fit residual (<0.25 = same audio)
train add_and_remove train/add_and_remove/00688_original.mp4
1280×704 · 23.98 fps · 5.09s / 5.09s · mp3 44kHz stereo
instruction →

Remove the figure in the red cloak riding the black horse from the cliff edge on the left.

← instruction_reverse

Add a figure in a red cloak riding a black horse on the cliff edge on the left.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
1.97mean pixel Δ (0–255)
0.6%pixels changed (Δ>30)
10.3p99 pixel Δ
+0.98audio envelope corr.
66%audio level kept (normal)
0.134gain-fit residual (<0.25 = same audio)
train add_and_remove train/add_and_remove/00955_original.mp4
1280×704 · 25 fps · 3.27s / 3.27s · mp3 44kHz stereo
instruction →

remove the shirtless man climbing the tree

← instruction_reverse

add a shirtless man climbing the tree in the same position and posture

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
4.78mean pixel Δ (0–255)
3.5%pixels changed (Δ>30)
47p99 pixel Δ
+0.60audio envelope corr.
47%audio level kept (normal)
0.661gain-fit residual (<0.25 = same audio)
train add_and_removesoundtrack collapsed train/add_and_remove/01111_original.mp4
1280×704 · 30 fps · 2.72s / 2.72s · mp3 44kHz stereo
instruction →

Remove the boy standing next to the tree on the left side of the frame.

← instruction_reverse

Add a boy wearing a navy blue jacket with an orange inner collar and dark shorts, standing next to the tree on the left side of the frame, holding a small object in his right hand.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
4.2mean pixel Δ (0–255)
3.8%pixels changed (Δ>30)
59.3p99 pixel Δ
+0.06audio envelope corr.
3%audio level kept (wipe)
0.999gain-fit residual (<0.25 = same audio)
train add_and_removeaudio unchanged train/add_and_remove/05451_original.mp4
1280×704 · 23.98 fps · 5.09s / 5.09s · mp3 44kHz stereo
instruction →

remove the husky dog lying on the bed with its head tilted upward and howling.

← instruction_reverse

add a black and white husky dog lying on the bed, wearing a blue collar, and howling with its head tilted upward.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
12.54mean pixel Δ (0–255)
10.8%pixels changed (Δ>30)
155.7p99 pixel Δ
+0.99audio envelope corr.
65%audio level kept (normal)
0.063gain-fit residual (<0.25 = same audio)
train add_and_removeaudio unchanged train/add_and_remove/06527_original.mp4
1280×704 · 23.98 fps · 6.71s / 6.71s · mp3 44kHz stereo
instruction →

remove the beagle dog standing on the couch and howling.

← instruction_reverse

add a beagle dog with a white, brown, and black coat, wearing a black collar with a tag, standing on the couch and howling.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
15.03mean pixel Δ (0–255)
17.8%pixels changed (Δ>30)
111.7p99 pixel Δ
+0.99audio envelope corr.
66%audio level kept (normal)
0.098gain-fit residual (<0.25 = same audio)
train add_and_removeaudio unchanged train/add_and_remove/08364_original.mp4
1280×704 · 23.98 fps · 3.42s / 3.42s · mp3 44kHz stereo
instruction →

remove the yellow and black helicopter flying low over the desert canyon

← instruction_reverse

add a yellow and black helicopter flying low over the desert canyon, kicking up dust as it moves from right to left

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
2.73mean pixel Δ (0–255)
0.2%pixels changed (Δ>30)
7p99 pixel Δ
+0.97audio envelope corr.
66%audio level kept (normal)
0.151gain-fit residual (<0.25 = same audio)
train add_and_removesoundtrack collapsed train/add_and_remove/10993_original.mp4
1280×704 · 23.98 fps · 3.42s / 3.42s · mp3 44kHz stereo
instruction →

Remove the woman with curly brown hair wearing a brown and white patterned outfit and a blue beaded necklace, standing in profile facing right.

← instruction_reverse

Add a woman with curly brown hair wearing a brown and white patterned outfit and a blue beaded necklace, standing in profile facing right, in the same position as the removed figure, and saying, <S>Molog also lied about Tibl<E>.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
6.62mean pixel Δ (0–255)
6.0%pixels changed (Δ>30)
102p99 pixel Δ
+0.01audio envelope corr.
0%audio level kept (wipe)
1.0gain-fit residual (<0.25 = same audio)
train add_and_removesoundtrack collapsed train/add_and_remove/11593_original.mp4
1280×704 · 23.98 fps · 6.71s / 6.71s · mp3 44kHz stereo
instruction →

Remove the man in the black leather jacket holding a microphone from the center of the stage.

← instruction_reverse

Add a man in a black leather jacket holding a microphone to the center of the stage, with a dark background and a faint blue light behind him, and saying, <S>Because I gotta say something to him Chris, he's napkin blotting the oil off his pizza. So that's a sin.<E>.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
11.4mean pixel Δ (0–255)
11.3%pixels changed (Δ>30)
148p99 pixel Δ
+0.00audio envelope corr.
0%audio level kept (wipe)
1.0gain-fit residual (<0.25 = same audio)
train add_and_removesoundtrack collapsed train/add_and_remove/14593_original.mp4
1280×704 · 23.98 fps · 5.09s / 5.09s · mp3 44kHz stereo
instruction →

Remove the woman seated on the dark couch, wearing a red top and black blazer, from the center of the frame.

← instruction_reverse

Add a woman with shoulder-length dark brown hair, wearing a red top and black blazer, seated on the dark couch in the center of the frame, with patterned pillows behind her and artwork on the wall in the background, and saying, <S>Because for so long they haven't necessarily been able to connect with other people.<E>.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
10.75mean pixel Δ (0–255)
13.2%pixels changed (Δ>30)
124.3p99 pixel Δ
-0.00audio envelope corr.
4%audio level kept (wipe)
0.998gain-fit residual (<0.25 = same audio)
train add_and_removesoundtrack collapsed train/add_and_remove/15493_original.mp4
1280×704 · 23.98 fps · 3.42s / 3.42s · mp3 44kHz stereo
instruction →

Remove the woman with long dark hair showing a distressed expression from the center of the frame.

← instruction_reverse

Add a woman with long dark hair, wearing a light-colored top, displaying a distressed expression, centered in the frame with dim lighting, and saying, <S>wasn't even there they were just<E>.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
13.27mean pixel Δ (0–255)
19.4%pixels changed (Δ>30)
65.7p99 pixel Δ
+0.05audio envelope corr.
13%audio level kept (collapse)
0.982gain-fit residual (<0.25 = same audio)
train add_and_removesoundtrack collapsed train/add_and_remove/16093_original.mp4
1280×704 · 23.98 fps · 3.42s / 3.42s · mp3 44kHz stereo
instruction →

Remove the man in the red suit standing on the stage holding a microphone.

← instruction_reverse

Add a man in a red suit standing on the stage, holding a microphone in his right hand and gesturing with his left, illuminated by a spotlight with his shadow cast on the backdrop behind him, and saying, <S>The Mexicans got them gangs you can't pronounce the names<E>.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
3.65mean pixel Δ (0–255)
4.9%pixels changed (Δ>30)
54.3p99 pixel Δ
+0.06audio envelope corr.
22%audio level kept (collapse)
0.967gain-fit residual (<0.25 = same audio)
train add_and_removeaudio unchanged train/add_and_remove/19764_original.mp4
1280×704 · 24 fps · 6.75s / 6.75s · mp3 44kHz stereo
instruction →

Remove the television screen displaying the C-SPAN broadcast of President Obama speaking at the White House briefing room.

← instruction_reverse

Add a television screen displaying the C-SPAN broadcast of President Obama speaking at the White House briefing room, positioned on the wall above the mantel, and saying, <S>A lot of it's legal, but that's exactly the problem. It's not that they're breaking the laws, it's that the laws are<E>.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
3.51mean pixel Δ (0–255)
4.0%pixels changed (Δ>30)
66.7p99 pixel Δ
+0.98audio envelope corr.
65%audio level kept (normal)
0.178gain-fit residual (<0.25 = same audio)
train add_and_removesoundtrack collapsed train/add_and_remove/20293_original.mp4
1280×704 · 23.98 fps · 5.09s / 5.09s · mp3 44kHz stereo
instruction →

Remove the man on the left wearing a black hoodie with a yellow graphic.

← instruction_reverse

Add a man on the left wearing a black hoodie with a yellow graphic that reads "STYL and the ORDER" and has dark, slightly messy hair, and saying, <S>So you stabbed him in the butt and then you just took it out and put it away?<E>.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
per-band audio change
gain-cancelled change per frequency band — the global level change is subtracted, so bars show change in spectral shape. A source being removed would show one band collapsing while the others hold.
7.39mean pixel Δ (0–255)
6.7%pixels changed (Δ>30)
107p99 pixel Δ
+0.13audio envelope corr.
19%audio level kept (collapse)
0.955gain-fit residual (<0.25 = same audio)
train clone_id train/clone_id/00000_original.mp4
1280×704 · 24 fps · 3.42s / 3.42s · mp3 44kHz stereo
instruction →

Keep the person's appearance, change the timbre to a man, and change the spoken words to <S>I think we need to talk about this right now.<E>.

← instruction_reverse

Keep the person's appearance, change the timbre to a person, and change the spoken words to <S>Washington Y'all should know that Sebastian's produced<E>.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
4.0mean pixel Δ (0–255)
1.1%pixels changed (Δ>30)
31.3p99 pixel Δ
+0.07audio envelope corr.
0%audio level kept (-)
-gain-fit residual (<0.25 = same audio)
train clone_id_voice train/clone_id_voice/00000_original.mp4
1280×704 · 24 fps · 3.42s / 3.42s · mp3 44kHz stereo
instruction →

Keep the person’s identity and change the spoken words to <S>I think we need to talk about this right now.<E>

← instruction_reverse

Keep the person’s identity and change the spoken words to <S>Washington Y'all should know that Sebastian's produced<E>

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
3.96mean pixel Δ (0–255)
1.1%pixels changed (Δ>30)
30.7p99 pixel Δ
-0.01audio envelope corr.
0%audio level kept (-)
-gain-fit residual (<0.25 = same audio)
train clone_voice train/clone_voice/00000_original.mp4
1280×704 · 24 fps · 3.42s / 3.42s · mp3 44kHz stereo
instruction →

Keep the timbre, change the person to a man wearing dark jeans, and change the spoken words to <S>I think we need to talk about this right now.<E>.

← instruction_reverse

Keep the timbre, change the man to a person, and change the spoken words to <S>Washington Y'all should know that Sebastian's produced<E>.

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
6.04mean pixel Δ (0–255)
5.3%pixels changed (Δ>30)
56p99 pixel Δ
-0.01audio envelope corr.
0%audio level kept (-)
-gain-fit residual (<0.25 = same audio)
train general_editing train/general_editing/00000_original.mp4
1280×704 · 24 fps · 6.75s / 6.75s · mp3 44kHz stereo
instruction →

Replace the man with a woman in her late 50s or early 60s, with graying hair and a beard, wearing glasses, a dark suit, a pink shirt, and a dark tie, and saying, <S>I’ve waited so long to tell you the truth.<E>

← instruction_reverse

Replace the first woman with a man appearing to be in his late 50s or early 60s, with graying hair and a beard, wearing glasses, a dark suit, a pink shirt, and a dark tie, and saying, <S>We want to put all this behind us start fresh Sarah is.<E>

INPUT · original

frames of original
6 frames, evenly spaced
spectrogram of original audio
audio spectrogram (0–8 kHz)

OUTPUT · edited target

frames of target
6 frames, evenly spaced
spectrogram of target audio
audio spectrogram (0–8 kHz)
waveform overlay
original vs target waveform envelope
frame difference heatmap
per-frame |original − target| heat map — bright = pixels the edit touched
5.36mean pixel Δ (0–255)
2.7%pixels changed (Δ>30)
45p99 pixel Δ
-0.03audio envelope corr.
0%audio level kept (-)
-gain-fit residual (<0.25 = same audio)