The future for whom? The general public? Not a chance, no way, not unless it's able to run on a phone (anywhere from 20-40% of internet users, world-wide, are phone-only).
For companies? I think that's a lot more plausible, as that's mostly just a question of money - is it cheaper to run and administrate our own models, or outsource that?
For technically inclined users? I think that's unlikely unless they're able to operate on relatively cheap hardware while still being just as good as the hosted models. And by that I don't mean "a mac studio," that's far more money than I think is reasonable. A single RTX 5080, maybe, once memory prices start to drop.
Compare the games your average high-end smartphone can run to the AAA titles of the 2010's. It's not a matter of "unless it is able to" but "when it is able to".
that's not a local LLM. If it's local, it doesn't matter in this case. Laya is a System 1 "AI", namely works like a classifier, given a state and questions, it shoots probabilities for each. I publish an episode tomorrow about Laya and Typesafe AI on https://www.youtube.com/@DataScienceatHome
Stay tuned ;)
It won't. Laya is a finetuned version of Google's BeRT model, which is almost 10 years old right now.
If BeRT had any potential to disrupt the datacenter buildout, it already would have.
Modernbert is from 2024. It's also trained from scratch, not a fine tune.