Turn a track into a finished music video inside ComfyUI. Whisper large-v3 transcribes the lyrics with word-level timing, librosa (or a built-in numpy/scipy fallback) finds the BPM, the beat grid and the sections, and a local LM Studio model - or OpenRouter, OpenAI, Anthropic - writes the treatment, the art direction, a recurring-subject bible and every shot. Out come start-frame image prompts, MiniMax H3 image-to-video and reference-to-video prompts in their exact six-section format, negatives, per-shot timings and sample-accurate AUDIO slices for lipsync. It can render the frames and the clips through fal.ai or OpenRouter, then cut every clip to its shot on the beat with PyAV and mux the original music back in. A live gallery shows each frame and clip in the node as it lands, and a cost meter reports in USD what every model actually billed. The analysis is always local and free; nothing is rendered until you pick a provider.
https://github.com/lazniak/comfyui-music2video
git clone https://github.com/lazniak/comfyui-music2video
comfy node install music2video