Fork of ComfyUI-QwenVL: Qwen-VL nodes (Qwen2.5-VL / Qwen3-VL, Transformers and GGUF backends) for text generation, image understanding, and video analysis. Adds local-only model discovery — models are picked from models/text_encoders and models/LLM with no automatic downloading — plus explicit mmproj selection, up to 3 image inputs on the Advanced nodes, a thinking-mode toggle, Gemma 4 / Qwen3.5 / Qwen3.6 / Qwen3.8 GGUF support, and MTP speculative decoding.
https://github.com/id-fa/ComfyUI-QwenVL-F
git clone https://github.com/id-fa/ComfyUI-QwenVL-F
comfy node install QwenVL-F