Ami Baid
I recently graduated from UT Austin with degrees in computer science and math as a Turing Scholar. During undergrad, I was a researcher in the UT Austin Computer Vision Lab, advised by Prof. Kristen Grauman. I'll be starting my Master's in CS at Stanford this fall🌲
My research focuses on audio-visual multimodal learning. I'm excited about developing intelligent systems that can understand and reason over information from diverse modalities.
Research
Personal Visual Context Learning in Large Multimodal Models
arXiv 2026 [paper] [project page]
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
ECCV 2026 [paper] [project page]
Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos
ECCV 2024, Oral [paper] [project page]
Self-Supervised Visual-Acoustic Matching
NeurIPS 2023 [paper] [project page]
Internships
- Engineering intern @ Stripe (summer 2025): extended Stripe's LLM-based compliance detection system to support image understanding on merchant websites.
- Software engineering intern @ Salesforce (summer 2024): automated a key workflow in Salesforce's internal Temporal platform and contributed to the open-source Terraform Temporal provider.
Other Projects
- Gaze-centered Egocentric Video Representations: built a gaze-aware preprocessing pipeline that reallocates resolution around gaze, improving efficiency in egocentric video QA. [GitHub]