词级音文对齐:多引擎——LIS token 匹配、stable-ts 交叉注意力与 Qwen3 强制对齐。支持 faster-whisper、UVR5 人声分离与视频字幕叠加渲染。
https://github.com/ahkimkoo/ComfyUI-Audio-Srt-Aligner
git clone https://github.com/ahkimkoo/ComfyUI-Audio-Srt-Aligner
comfy node install ComfyUI-Audio-Srt-Aligner
Qwen3-ASR 轻量节点包:简单语音转文本工作流——本地模型缓存与时间戳字幕输出。支持官方 Transformers 原生 Qwen3-ASR-*-hf 模型,含 Huggin…
Qwen3-ASR (0.6B/1.7B) 与 ForcedAligner 的节点。支持高精度 ASR 与 52 种语言/方言的语言识别,含 22 种中文方言与各种英语口音。特性:…