Solaris, Our Interface World Models

8 pointsposted 6 hours ago
by babelfish

3 Comments

kakugawa

3 hours ago

https://www.youtube.com/watch?v=c2yCePPnrSA

Their analogy is that a VLM responds to text, like their interface model responds to clicks. There is no UI (just images), and based on your clicks the model infers your intent and adapts the "UI" in response. So, instead of inferring intent from an information-dense input (text), they do it w/ just mouse-based gestures? I would love to see how this holds up in practice.

A fun little anecdote @ 54s in the video: "It becomes whatever you ask of it. And no two interactions are ever the same."

shay_ker

2 hours ago

if you want to try something like this, play MIRA: https://mira-wm.com

notably, rocket league is a closed loop, deterministic game, which is why kyutai labs chose it as a "base case" for world modeling.

it looks cool at first but really quickly you'll start to see the gaps

beklein

4 hours ago

I don't know what kind of experiences and apps this UI/UX mode will enable but it's one of the most innovative and coolest AI model demos I've seen this year.