The lazy version of the browser-agent story is that better models will fix everything.
I do not think that is enough. Agents fail because the browser is a bad sensory surface: too much irrelevant text, too little stable structure, hidden state, shifting DOMs, overlays, disabled affordances, and failures that are hard to inspect after the fact.
The work I keep coming back to is representation. What should a browser agent see? What should it remember? What should it ignore? What should be obvious when it fails?
This essay will become the written version of that argument.