kakugawa
3 hours ago
https://www.youtube.com/watch?v=c2yCePPnrSA
Their analogy is that a VLM responds to text, like their interface model responds to clicks. There is no UI (just images), and based on your clicks the model infers your intent and adapts the "UI" in response. So, instead of inferring intent from an information-dense input (text), they do it w/ just mouse-based gestures? I would love to see how this holds up in practice.
A fun little anecdote @ 54s in the video: "It becomes whatever you ask of it. And no two interactions are ever the same."