A ComfyUI extension for chatting with your images. Runs on your own system, no external services used, no filter. Uses the [a/LLaVA multimodal LLM](https://llava-vl.github.io/) so you can give instructions or ask questions in natural language. It's maybe as smart as GPT3.5, and it can see.
https://github.com/ceruleandeep/ComfyUI-LLaVA-Captioner
git clone https://github.com/ceruleandeep/ComfyUI-LLaVA-Captioner
comfy node install ComfyUI-LLaVA-Captioner