Browser agents need better eyes
A talk and working thesis about browser agents, observation design, and why better models still fail when the state is bad.
I work around AI systems that have to touch the real world: voice calls, browsers, multilingual data, medical images, point clouds, and the awkward interfaces in between.
A lot of model failure starts earlier than the model: in the data, the interface, or the state it was given.
I like tools that leave people more capable after using them, not more dependent on the tool.
The best systems work usually feels like rehearsal: timing, feedback, trust, and many small corrections.
Some things I have made, studied, or keep pulling on. A first shelf, not a portfolio.
A talk and working thesis about browser agents, observation design, and why better models still fail when the state is bad.
Real-time voice-agent and LLM systems for Indian-language phone, WhatsApp, and enterprise contexts.
A public ACL paper on using visual context to find and correct caption errors in English-to-Indic multimodal translation data.
Short notes first. More complete essays as the site earns them.
A stub for the argument behind my AI Engineer talk: many browser-agent failures are observation failures before they are reasoning failures.
A short note from the WAT 2025 caption-correction work: sometimes the useful move is not a bigger model, but a cleaner signal.