AI by department · Sound and post sound

The tools are good here, and the sharpest consent line in the industry sits here.

Source separation, dialogue isolation, noise and reverb removal and speech-to-text logging are in daily professional use, and they work. It is also the department where the industrial fight happened, because a voice is a person.

Last checked September 2026. Nothing on this page is legal advice.

What works

What works.

ToolWhat it doesWhy it matters
Dialogue isolation and lav rustleiZotope RX 12’s Dialogue Isolate and De-rustle.Rescues performances that would otherwise need ADR. iZotope’s published prices in September 2026: Elements $99, Standard $399, Advanced $1,399.
Source separation and recoveryMusic Rebalance and De-bleed; Spectral Recovery rebuilds content above 4kHz lost to compression.Archive with no surviving stems becomes usable, and compressed contributor audio sits in a broadcast mix.
Separation inside the editAdobe’s Enhanced Audio, in beta since September 2026, splits a mix into dialogue, music, ambience and effects.The decisions start being made in an editor’s application, before they reach post sound.
Speech-to-text loggingTranscribes everything, searchably.The biggest time saving in the post chain — and transcription still hallucinates, so it is a starting point, not a record.

Almost none of it is generative: it reads what you already recorded rather than inventing, which keeps it clear of most rights arguments. Heavy processing has a sound, and a track rescued at high settings carries an artefact a dubbing mixer can hear.

The consent line

A voice is a person, and a general release almost certainly does not cover cloning it.

Any synthesised or cloned voice needs the performer’s specific, informed, written consent for that use. Not a general clause, not one signed years ago for a different purpose, and not an assumption that owning the recording means owning the voice. Cloning starts at $6 a month on ElevenLabs’ published price in September 2026, so cost is not a barrier to anybody. And in AI & Work in the Media, Arts & Entertainment Sector in Europe, published in July 2026 by FIA, FIM, UNI MEI and EFJ, voice is the worst-affected category of all: respondents reported speaking engagements down by 80 to 90 per cent, and that “dubbing jobs have practically disappeared”.

  • !
    List item text
  • !
    List item text
  • !
    List item text
  • !
    List item text

Before you agree to anything

Ask before you agree to anything.

AskWhy
Is any voice work on this production syntheticIncluding narration, crowd, looping and temporary lines that survive into the final mix.
Whose consent exists, and for exactly what useA named person, a dated document, a described use. Anything vaguer is an assumption.
Is uploaded audio used to train the vendor’s modelCheck the processing agreement and the retention period.
Can this run on my own machineSeveral of the strongest tools do, and most of the confidentiality problem goes with it.