This is an implementation of [a/Qwen2-VL-Instruct](https://github.com/QwenLM/Qwen2-VL) by [a/ComfyUI](https://github.com/comfyanonymous/ComfyUI), which includes, but is not limited to, support for text-based queries, video queries, single-image queries, and multi-image queries to generate captions or responses.
https://github.com/IuvenisSapiens/ComfyUI_Qwen2-VL-Instruct
git clone https://github.com/IuvenisSapiens/ComfyUI_Qwen2-VL-Instruct
comfy node install ComfyUI_Qwen2-VL-Instruct