VisionChat
Upload or capture an image and talk about what's in it — real-time YOLO detection paired with a vision-language model for follow-up chat.

An AI-powered application that allows users to upload images, provide image URLs, or capture photos via camera to detect real-world objects and interact with them conversationally. The system uses real-time object detection and a multimodal language model to generate accurate descriptions and enable follow-up questions about detected objects.
Key features: image upload, URL input, and live camera capture; real-time object detection using YOLO; automatic object description using a vision-language model; an interactive chat interface to ask questions about detected objects; and an end-to-end AI pipeline from vision to conversation.
