Swap the speaking voice in a mixed audio clip (dialogue + ambient sound) while keeping the background intact — Demucs separation, Whisper transcription, local-Gemma emotion tagging, ECAPA-TDNN speaker auto-assignment, Fish Audio S2 cloning, with a Muse-Director-style timeline editor for reviewing/correcting segments.
https://github.com/muse-collective-26/MuseVoiceSwap
git clone https://github.com/muse-collective-26/MuseVoiceSwap
comfy node install muse-voice-swap