Achieved 2nd place out of 50+ teams for developing an innovative AI application using cutting-edge Vision-Language Models.
Project Overview
Built a VLM-based generative AI application that demonstrates advanced multi-modal reasoning capabilities, combining visual understanding with natural language generation.
Technical Highlights
- Implemented CLIP-based image understanding pipeline
- Integrated GPT-4V for visual question answering
- Built real-time inference system with sub-second latency
- Developed intuitive user interface for demo
Competition Details
- 24-hour hackathon format
- Judged on innovation, technical execution, and presentation
- Competed against graduate and undergraduate teams