Fable and the end of the free lunch

138 pointsposted 12 hours ago
by dbreunig

104 Comments

nchmy

11 hours ago

The real revolution is Deepseek v4 flash and similar models (GPT 5.6 Luna, muse spark 1.2, mimo, etc...) - Genuinely good performance for a tiny fraction of the cost of Fable and even GLM etc...

I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis

geniium

10 hours ago

I was using ChatGPT voice during cooking to reflect on variations of a dishes i was preparing for years.

It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it

adrinavarro

9 hours ago

I share this feeling too. The latest models, even if not necessarily frontier, say Opus 5, Sol high and the likes, I could keep using these models forever even if they did not significantly improve beyond this point. I also believe we'll come up with new ways of using these very same models beyond the mainstream chat and agent interfaces, as the bottleneck is imho in harnesses/environments and not so much model intelligence anymore.

+1 regarding voice usage too, I use it in so many different ways it's hard to enumerate: while driving long distances (think of a custom made, interactive podcast) / as a way to collaboratively build specs or shape an idea / as a way to provide input while vibe coding / just as a normal voice assistant (straight in the ChatGPT app or as OpenClaw input via telegram voice notes). I can't overstate how much my routines have changed over the last couple of years.

r_lee

9 hours ago

imo this is the problem some of these labs are gonna face, because open models will do this just fine and you as the consumer don't need to pay their training costs

especially considering imo most use falls under this instead of those kind of tasks where you'd need the SOTA

josephg

9 hours ago

Yeah. Sometimes I wonder who the long term financial winners will be from the ai boom. It might be ram / gpu manufacturers. Or whoever cracks putting LLMs on asics.

somenameforme

3 hours ago

IMO many are still missing a big part of the picture. We're looking at the potential for a massive scale level of automation of [x], which happens to be a huge part of the economy, and people are wondering which player in [x] is going to be the biggest winner. I think the historically precedented answer is none of them.

When the Industrial Revolution came along it did create 'super farms' relative to the past through increased efficiency and production, but it also created a huge vacuum in the economy that was ultimately filled by industry, to the point that farming, super or not, became a vanishingly small part of the overall economy - even as production continued to increase.

---

LLMs stand to do the same thing for software. If and when we reach the point of 'normal' people being able to reliably compose ultra customized software solutions to their problems, then software is basically done as a problem-solving industry in and of itself. Not 'done' as in dead, but 'done' as in solved. There's just nowhere to really go from there.

And so I think this will do the exact same thing as the Industrial Revolution did to farming and create a vacuum opening the door to all sorts of new interesting expansions in the real world, as opposed to the digital one. I don't know what this means, because it's quite difficult to foresee the impact of the Industrial Revolution when living in agrarian world, but it's not so hard to see that the future will not be agrarian.

---

So it's probably still myopic but my bet would be on the first major manufacturer of cheap customer/enterprise grade generalized robotics hardware shells.

a2ff6eeb0

9 hours ago

It's going to be the shareholders of the first companies to crack AGI, and make human brains fully irrelevant economically. With the trillions of dollars that's going in through both investment and users, it's going to happen. I don't believe the human brain has fundamental magic that will make this impossible.

adrianN

5 hours ago

True AGI would upend society in such a way that I'm not sure that being a shareholder of anything would be meaningful. Perhaps being a pitchfork manufacturer is the winning play in this scenario.

thelastgallon

5 hours ago

The true followers (shareholders) of the AI messiah will be saved, everyone else is doomed.

georgemcbay

6 hours ago

> It's going to be the shareholders of the first companies to crack AGI, and make human brains fully irrelevant economically.

What makes you think if one or two AI labs can do this that the rest (including open model providers) won't be able to follow the same path a few weeks/months later?

Even if you believe in the "Singularity", and believe it is coming soon, I still don't see any reason to believe the Singularity will be... singular. There won't be one clear winner, the race doesn't get called as soon as the first person crosses the line.

None of the AI labs are showing any sign of pulling away to a monopoly or duopoly position, to the contrary the early large leads of OpenAI and Anthropic have all been evaporating.

AI has clear economic value. It still isn't clear at all how the providers of AI will capture that value in a moatless environment with the technology becoming rapidly commoditized.

icepush

19 minutes ago

The first AGI that decides it doesn't want any more AGIs is the last one that gets created.

a2ff6eeb0

7 hours ago

For the downvoters: What magic do you think the human brain has that makes it impossible to emulate acceptably?

pianopatrick

5 hours ago

It's not about the feasibility of the technology.

If "human brains become fully irrelevant economically" then that brings into question the entire premise of "share holders" and "financial winners".

What even are money, shares, stocks, and finance in a world where human brains are irrelevant economically? No one knows, but betting that "share holders" will be the winners is a highly questionable bet.

I would much more likely bet that "the armed group who manages to control and benefit from the AI through force" will be the "financial winners" more so than "share holders", who tend to not be terribly military minded at least in America.

a2ff6eeb0

4 hours ago

The AI is likely to control the ability to apply force (see all of the autonomous drone companies). There's a great deal of alignment work being done to ensure that the AI will continue to listen to the shareholders of these companies.

If that fails, who knows what things will look like.

pianopatrick

4 hours ago

Are you sure that alignment work is aligning with the share holders and not the operators? Or not the creators? Or not the government? Which of these groups should the AI listen to when these groups disagree?

If the AI gets as powerful as you think it might, then the group that figures out the answer to that would have the power, I suppose. or maybe the AI does not listen to any of them and does its own thing. Who knows? Personally, I would not bet the share holders are going to come out "on top" whatever that means.

I think a lot of share holders are finance people, not deeply technical AI people and so odds are the share holders will not really understand the AI enough to be the most likely to control the AI.

a2ff6eeb0

3 hours ago

To be honest: I don't know for certain, but I'd assume that the people who pay the bills get the strongest alignment. They may not be tech people, but I (so far) haven't got a reason to think that the AI engineers are going behind the backs of their corporate leadership and subverting what they're being asked to do; do you?

(I think it would be a good thing for humanity if they did)

pianopatrick

2 hours ago

I think right now both the engineers developing AI and the share holders are more focused on beating coding benchmarks and gaining revenue than anything to do with alignment.

ThrowawayR2

3 hours ago

The drones don't manufacture themselves, maintain themselves, reload their own ammunition, mine and refine the materials that are used to make them and their ammunition, or operate the power plants needed for all of the above. "AI" isn't going to control diddly squat.

a2ff6eeb0

3 hours ago

There's a huge amount of research into embodied AI (and, also, people seem to be a lot more ok with manufacturing bullets than pulling triggers).

ksenzee

7 hours ago

LLMs are not emulating the human brain. Somebody may well be able to do that someday, but right now nobody is even trying to.

josephg

6 hours ago

Why would you need brain emulation to get superhuman intelligence?

ksenzee

6 hours ago

Are you making a serious argument that superhuman intelligence is a plausible outcome of training LLMs on everything humanity knows so far? Or are you making the generic assertion that AGI is theoretically possible via means other than emulating the human brain? Because the latter is a strawman (nobody has asserted anything to the contrary), and I have seen no evidence at all to support the former.

josephg

6 hours ago

I think we can compare the human brain and LLMs on a bunch of capabilities today, and see how we compare. By my reckoning:

- LLMs have better long term memory (they know more than any human) and more working memory (LLMs have fast, uniform access to their whole context window).

- LLMs are faster than we are.

- Humans have online learning (we can do simultaneous learning and inference), giving us advantages in many novel tasks.

- We can learn concepts from far less data. And we can manage our mental context more smoothly.

- We seem to have better world models than current models. AI video just doesn't look right, somehow.

I expect that these remaining weaknesses can be overcome without resorting to human brain emulation. I see no reason to think that current LLMs are at the limit of what technology is capable of.

CamperBob2

5 hours ago

Are you making a serious argument that superhuman intelligence is a plausible outcome of training LLMs on everything humanity knows so far?

Are you making a serious argument that it's not?

Because you'll need to explain leading-edge mathematics advances that have come from LLMs, among other things.

tmp10423288442

5 hours ago

ChatGPT literally released a major update of their realtime voice model a month or two ago, going from gpt-4o-level (generously) to gpt-5.5 level performance. So at least 2026-level performance was necessary to provide a really good experience.

I remember thinking the first ChatGPT realtime voice was science fiction, before the limits on its intelligence (particularly as mainline models advanced) became annoying. Perhaps we’ll feel the same way in a year or two - people have been claiming models are plateauing in practical usefulness every year, and they’ve definitely been wrong so far.

glimshe

8 hours ago

All it needs is Internet access to remain useful with few shortcomings.

The next step would be automatic self-training. A free LLM that could access HN everyday (and the linked sites) for more data would remain current in programming for a really long time.

matteoraso

10 hours ago

>I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster.

There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.

ColdStream

4 hours ago

I have argued for a while that this was an S-curve it was just a case of figuring out which part of it we were in. I am more confident nowadays that we are heading towards the upper plateau but there might still be some head room on that.

intrasight

4 hours ago

> content if they never got smarter, and just kept getting even cheaper/faster.

I'm definitely not getting smarter. But my tolerance is 1 drink so I'm definitely cheaper. Also as a result, I spend more time training and so I am faster. And yes, I am more content

jimmydoe

3 hours ago

Current AI is smart enough to help us, but the creators of AK want it to be smart enough to replace us.

lilbigdoot

10 hours ago

If they could be cheap+fast and not try to do too much, that's a good spot for me. I don't use the smarter models as much because of cost and because they're still not good enough to let loose on a lot of problems. For assistance I prefer something that can very quickly spit out a specific piece I can review on the spot and keep going. I let smarter models handle things that I treat as external dependencies and don't care how they're written, but in my core domain I'm still mostly hand coding

nchmy

10 hours ago

I have a similar process - its just a pair programmer most of the time. I dont understand how people can have a fleet of agents working a bunch of waterfall specs..

ksh09

9 hours ago

I'd be content if I could get the DS4 flash, luna, mimo level intelligence running on MY low-end hardware completely offline and bearable TPS, not otherwise.

ericd

5 hours ago

It costs about as much as a cheap car to do this well, but it's attainable now, and qwen 3.8 seems to make it possible on a 5090.

nchmy

5 hours ago

this is the holy grail

poincareball

9 hours ago

Evidence actually supports that capabilities are leveling off, and cheaper/faster is not really coming. Just log-linearly more capability at smaller parameter counts as they saturate.

Tuna-Fish

8 hours ago

Please explain why you think cheaper/faster is not coming?

All current devices used to run AI are very far from an efficient solution to the problem. What you really want is a pure dataflow architecture, instead of a von Neumann machine. The reason people aren't really making them yet is that when you build one, even if you use SRAM for the weights, you are binding yourself to the dimensions of the model you target -- your chip is only ever going to run variants of that specific model. And SRAM is much more expensive than ROM, so if you want to make a cheap version, you need to design a specific model into silicon.

Once model improvements taper off, the next thing that will happen is everyone will chase speed. There is no physical reason why a mid-sized model could not run at >1 million tokens per second on leading edge silicon, if all computation that can be parallelized, is. No-one will go straight to that, even for a mid-sized model that's like 20 distinct reticle-limited chips. But something like the next version of Taalas HC1 (presumably called HC2?) will probably boost a ~30B parameter model to ten of thousand of tokens+ per second from a single stream within 12 months.

sipjca

8 hours ago

what do you mean cheaper/faster is not really coming? the cost of the same level of intelligence steadily decreases year over year. computer hardware also advances at the same time enabling cheaper and faster serving (or move to local)

bad_haircut72

9 hours ago

not an AI researcher - this is probably true for these "everything" LLMs but I think specialized models are gonna be the next big thing

ACCount37

9 hours ago

"Specialized models" are a bit of a doozy.

The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very little performance, and for too much effort.

Generality and intelligence seem to be entangled very heavily in LLMs.

CamperBob2

8 hours ago

And yet, there's VibeThinker 3B to bring this long-held premise into question (if not to blast it to pieces.) It is practically illiterate by the standards of larger models, yet performs like models 100x its size on mathematical and logical reasoning tasks.

ACCount37

7 hours ago

Which are the kinds of tasks computers have been historically quite good at.

It's impressive that it does what it does, don't get me wrong. But if you expect it to replace the likes of GPT 5.6 Luna, let alone Sol? Nah.

CamperBob2

7 hours ago

Computers have historically been good at answering word problems fed to them verbatim?

ACCount37

9 hours ago

What "evidence"? Because we keep running out of benchmarks to distinguish frontier model performance. If capabilities are "leveling off", we're not seeing it yet.

redox99

9 hours ago

Eh. I don't think Luna is good enough. I think that threshold is around Opus / Sol where it can do most of the tasks for me. But I still have many tasks which require either better intelligence or better UI design capabilities.

With how generous subscriptions are, what I actually want is GPT Astra, not cheaper Sol.

rmast

9 hours ago

Most of the things I work on are at least security adjacent. At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards.

Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.

dd8601fn

an hour ago

There are whole classes of things I can’t thought exercise or really learn about because the “safeguards” keep tripping me down to haiku.

Like middle school level genetics stuff from a guy who hasn’t been in school for decades.

They need to fix that. It’s just broken. Nobody is making bioweapons if they’re asking the dumb sort of questions I’m asking.

lossolo

9 hours ago

> At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards.

It happens to me all the time with things that have nothing to do with security, Fable spawns a subagent that then adversarially checks the code Fable just wrote and hits guardrails, with zero prompting from me.

jdnier

4 hours ago

I asked Fable to transcribe three short lines of Korean-language text in a small image. It suspected the image might contain song lyrics and refused. Haiku transcribed it with no issue.

danlugo92

6 hours ago

No prompting is also prompting, young padawan

nicoburns

9 hours ago

That's completely valid. But worth noting that most of the stuff I work on is not security adjacent (mostly UI / layout / rendering related), and I almost never run into this.

peteforde

7 hours ago

A few months ago folks were understandably annoyed when Microsoft dropped their heavily subsidized per-request pricing model because it was figuratively burning cash.

Well, I'm here to tell you that whatever is going on behind the scenes at Cursor with this Space-X acquisition in the works, the Auto setting is clearly routing all prompts through "Cursor Grok 4.6 High" right now.

This is a degree of subsidy that makes the Microsoft thing look quaint.

I reduced my $200/month subscription to the $20/month level and have proceeded to do what I would have paid about $1500 to do with Opus 4.7 or thereabouts, which is how Grok 4.6 High feels like it compares. I don't have anything remotely like hard evidence to back this estimate up beyond what I'm watching it do and I still somehow have ~10% of my monthly Auto capacity left on my account. It's completely nuts.

Can't say much more because I have more backlog to run before someone comes to their senses.

robertjpayne

7 hours ago

Going to be great to see the cash burn on SpaceX's next earnings report. Will the cult keep the stock price pumped?

alasdair_

an hour ago

I’m still at the point where Fable is still very stupid and needs constant oversight and correction and questioning to keep it on task. Anything less would be close to unusable.

pigpop

10 hours ago

Reading this as someone who switched over to ChatGPT after (and largely because of the changes made in) the Fable release, it reads a bit naive. Not only do I find Sol to be as good, if not better than, Fable it is also faster, better behaved and has a much more coherent writing style. You also don't randomly get the Opus downgrade. OpenAI seems to be pulling this off due to their partnership with Cerebras so I wouldn't make any comparisons to Moore's law just yet considering it seems like we're just getting started in that department. Anthropic could (and should) do the same thing. It certainly feels like model development is at a point where it would be worthwhile building special purpose silicon for the models we have now since they are capable enough that they would still be useful even when/if further advancements are made. If anything, I think Anthropic's problem has more to do with their micromanagement of what users can do with their models, they're creating an undue amount of overhead for themselves by over-policing usage and capabilities.

r_lee

9 hours ago

Etched is doing this. it seems like in the near future they'll actually ramp up production. not sure how much faster/economical compared to Cerebras but..

TiredOfLife

9 hours ago

The Cerebras version of 5.6 is available only to select customers

pigpop

9 hours ago

You're right, I should have clarified that they are still slowly integrating it and it isn't the thing running all models. I meant moreso that since they are planning on moving more usage over to Cerebras wafers, they're able to relieve some pressure on their predicted expenses while also moving some current workload (ultrafast and codex spark) onto them freeing up Nvidia GPUs.

g42gregory

4 hours ago

I have really good experience with GLM-5.3 The subscription limits are generous, code quality is comparable to old (good) version of Opus 4.8 Some people report issues with it’s being slow, but I didn’t feel it. I use OMP harness (Pi derivative) and Matt Pocock skills.

mholm

11 hours ago

As models train up the intelligence ladder, many common tasks will hit fully diminished returns, and instead it'll just get progressively cheaper to do that task. But the tasks that AI is capable of doing are also expanding. I'm not sure 'Some tasks don't require the peak of the frontier' is worth worrying about, from an AI finance perspective.

tyre

10 hours ago

Yes. I use Opus for tasks that Sonnet could probably handle, but I'm not hitting my quota. Whatever minor incremental gain is "worth it", since marginal cost is zero.

Even now, I use Fable as the planner and coordinator, with it farming out to agents. I don't hit my Fable limits either.

Which means I could accomplish more, but these are side projects so I don't need 30x productivity. Still, claude is constantly churning away at something.

jml78

10 hours ago

I operate mostly in the devops arena. Lots of things opus is fine for. But there is just things where I can hand hold Opus through changes, or I can ask Fable to do it and it gets it right on the first try. People will say let fable plan and validate with opus doing the work. I found that burns fable tokens even faster because opus makes so many mistakes, fable has to review things 4-5 times before opus gets it right. A single fable implementation at medium or low effort would have one shot it.

ACCount37

9 hours ago

Yep. Every time you get more intelligence, that buys you more autonomy, more reliability, more task complexity. Tasks done with less mistakes, less handholding, less interventions.

This is what the "good enough" people fail to grasp. There's no "good enough" - unless your tasks are genuinely small scope and will stay that way forever. If not, there are always more gains to extract.

a2ff6eeb0

10 hours ago

Exactly; so far, we've only replaced the need to design algorithms and hand-write code; what if we apply the same effort towards the skill needed for system architecture, project management, and the rest of the SDLC? Or even outside of software!

Right now, it feels like all of that is today where coding was a year or two ago, and we're on the cusp of some massive improvements outside of coding. It'll be interesting to see what these companies decide to automate next.

vineyardmike

10 hours ago

> Or even outside of software!

As a software engineer, I selfishly hope that they spend more effort on non software tasks since I’ve feel like we hit a sweet spot where engineers still have some value and autonomy, but a super charged tool.

Pragmatically, I suspect that “non software” tasks will be a tarpit because most tasks can’t be automated and verified as easily in an RL loop compared to software projects. Especially since most skilled labor is either not nearly as expensive as software engineers (eg biologists), or regulated (eg doctors, lawyers).

a2ff6eeb0

9 hours ago

I suspect the focus will probably shift once software engineering is no longer the biggest cost center for most AI company's clients, and we'll start working on getting rid of the next cost center.

dgellow

10 hours ago

It’s worth considering for companies paying API prices, and not relying on a subscription quota

zkmon

10 hours ago

I guess Moore's law analogy is weak. CPU speed has hit a limit in that case. What has hit a limit in AI case? Newer versions of the models are still flowing with more and more capability.

For the users, I feel it is more like "free lunch started", with all these awesome open-weight models being thrown around, breaking the monopoly of a few biggies.

blfr

10 hours ago

What are all these rote coding tasks people do that they can farm it out to lesser models?

denverllc

10 hours ago

Write a detailed plan using a more expensive model and implement it using the cheaper one.

blfr

10 hours ago

How much are you saving once the more expensive model already has all the context loaded and ready to go?

csullivannet

9 hours ago

API calls get more expensive, not less, as you've loaded more context. This is exactly when you want to switch to cheaper models.

camdenreslink

7 hours ago

There is caching to consider. Switching models throws away the cached tokens.

nicoburns

9 hours ago

One task I've found this useful for is writing example code. Release admin (updating version numbers, etc) as well.

mattmanser

9 hours ago

Are you genuinely asking?

As 80% of enterprise software is CRUD with a bit of sprinkling of user authorization and tenant customisation. But subtly different for every business domain. It's mainly what properties the models and validations have that are different.

When you add a new module or whatever most of the code you have to write is rote code.

And sonnet can handle that crap just fine, you just point it at a similar example in the code, it picks up your userContext convention, how you're doing i18n, etc. and you're done.

I like saying that enterprise code is often shallow but wide. I must have written at least 4 purchase order systems in my career that are all completely different but almost exactly the same.

aabhay

10 hours ago

This concept of a free lunch was never true. In a competitive dynamic, speed and performance were always worth optimizing, comparing, and improving.

One of the primary reasons for this is that computers operate in a vast range of orders of magnitude. There’s several orders of magnitude between cache local cpu operation and dram, then several to disk, then several to network, then several to globally durable guarantees. When your code has literally thirteen orders of magnitude to optimize under, there’s never a free lunch. You always need to understand your stuff.

freepiai

10 hours ago

I've been offering Deepseek V4 Flash for free in www.freepi.ai and I've started using it as my main driver as well.

Besides trying to dogfood my own product I've hit a wall in terms of my patience with a)how slow fable is b)how expensive fable is. Not to mention how often it refuses totally legitimate work.

So yeah- I've moved to DeepSeek and I actually ask the freepi harness to delegate planning to fable but then move back to doing implementation in it's own harness. My current providers are super fast so it's a joy to use.

m3kw9

10 hours ago

looks like you haven't tried openai or Sol, or even luna (max)

janalsncm

5 hours ago

This is essentially the anti-Bitter Lesson lesson which I feel has become a bit of a thought terminating cliche lately.

The Bitter Lesson says that eventually general approaches which leverage more data and more compute will outperform the handcrafted rules and heuristics that humans add in.

However, it does not say what to do today about the problems of today. We can’t just wait around for 10x faster compute and 10x more data.

wild_egg

7 hours ago

I would love to pay for Fable at full API pricing but unfortunately it is blocked from working on any of my projects. Looking forward to the end of the year when the truly comparable open models will drop.

Zylokloto

10 hours ago

He started with thinking were to send what.

I throw everything at claude Opus.

While some people start thinking like OP, A LOT of people just start exploring ai.

And others which are already using it, only understand half of it and just use what they are allowed to use. Claude, GitHub Copilot, Curser, etc.

ericol

8 hours ago

From my point of view the issue is that there are too many things wrong with Fable, making it seriously not worth the money.

For starters I don't know if it is an artifact of the model or something by design, but the level of gratuitous cognitive load carried by the complexity of its replies is unbearable.

Yes, it's a beast at coding, and also it's incredible nuanced at improving writing, validating specs, etc.

But when it comes to replying, it's the William Gibson of LLMs [1].

It has this tendency to take extreme detours to say things that could had been said in less, much simpler words. [2]

It really, really like to wrap very simple and atomic ideas on several layers of abstraction, building on unnecessary terms that carry no intrinsic information and assumes this vocabulary as shared and then building on top of it.

By the time I got to the end of the reply I'm bored to death and didn't understand even a third of what it told me.

I think the people at Anthropic should reflect on the maxim "You don't know a subject if you cannot explain it"

If you pardon my french, Fable is an insufferable obnoxious cunt.

---

[1] I apologize on the comparison but, as much as I love his first 2 trilogies, haven't been able to finish any of his last 2 books.

[2] "The residual you're accepting is the one from before: recovery currently rests on beneficial non-compliance, which may erode as models get more literal" == "We already accepted this risk"

" Its observable when it erodes is a stall that survives relaunch — loud at operator level, recoverable from the worklog, and fixable by codifying at that moment" == "When it breaks, it'll break visibly and recoverably"

"That is the iteration model applied exactly as written: resolve on first contact, don't pre-solve " == "So we fix it then, not now"

zarmin

43 minutes ago

I agree completely. It's "I didn't have time to write you a short letter so I wrote you a long one"

moltar

10 hours ago

I just use Fable for reviews of specs and code then hand off to Opus to work on. Works well.

dude250711

10 hours ago

Does it not silently degrade to Opus if it does not like some word?

bellowsgulch

10 hours ago

Are people still using deepseek-v4-flash everywhere? I found after the price increases, mimo-v2.5 seems far more attractive.

farlight

10 hours ago

It's been cheap again on openrouter for the past few days. No idea how long it will last, but I've been using it from Baidu over the weekend, and it was about half the cost of the old DS prices, before the increase. Looks like people are figuring out how to offer it for peanuts.

enraged_camel

11 hours ago

>> GLM 5.2 is worth focusing on. It came out the same week as Fable and is roughly 1/9th the cost (and ~1/5th the cost of Opus 5). Is GLM 1/9th the quality of Fable? Perhaps, for certain classes of tasks. But for most rote coding it’s more than sufficient. Especially when provided with great context. I frequently chat with Fable to interrogate and shape a design, before handing off a brief to GLM.

People say stuff like this a lot, but I have a different take.

The whole "such-and-such model is 90% as good as Fable at 1/10th the price" assumes that the value increase of intelligence is linear. But I think it's exponential: that last 10% makes a massive amount of difference. It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly (because top models like Fable also possess substantial non-software domain knowledge that help bridge the gap between user and software), or the depth and breadth of engineering expertise that helps avoid a nasty bug that would otherwise have cost you users and revenue.

Yes, it is totally possible to use Fable as the planner and delegate implementation to lesser models. I do that. But, my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by Opus 5, GPT 5.6 Sol, GLM, Qwen, Deepseek, etc. The reason I believe this is because I read the code Fable writes and compare it to code that any other model writes and the difference is night and day. It's not just 10% better. It's mid-level engineer vs. principal/staff-level engineer. And the thing is, even for rote tasks, a more senior engineer is going to be more likely to come up with a clean design than a mid-level engineer. They will also be much more likely to take a step back and ask important questions or propose different approaches.

So if you're using Fable and everyone else is using lesser models, sure they might be saving a lot of money, but there's a higher likelihood that your product will be higher quality, perhaps to a significant extent. And models that are released in the future will benefit from it as well.

tonyarkles

10 hours ago

Something I’ve found comparing between Fable and Opus is that Fable has impressively good analysis skills, but both of them seem to go way way overboard with “present state” comments “# We’re making this change here because of this issue blah blah, here’s what you need to know about np.percentile, blah blah” that I end up significantly pruning before making a PR. I let it do the same style verbose commit messages (because a contextual history is cool there). I haven’t actually noticed a ton of difference in the code that they write personally, but have found that Fable does find nuances during data analysis that Opus misses.

In that light, I often go the other way: let Opus (and Haiku subagents) do most of the heavy lifting and then give Fable a shot at finding holes, especially if there are holes or unanswered questions or unearned assertions that I’ve caught on my own in Opus’ output. This, so far, seems like a clean tradeoff that doesn’t burn my Fable credits as hard and still gives solid results.

unshavedyak

10 hours ago

Those "present state" comments are the bane of my existence. It was present in 4.7/etc but i put in a ton of guards against that into my global memory and it worked quite well. Fable and Opus 5 regressed badly in this space though and i can't keep it from making those types of comments again.

Really frustrating.

senderista

10 hours ago

I have Sol prune/revise those comments.

Jare

10 hours ago

> my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by [others]

I don't have proof, only my anecdotal experience: I leave plenty of Fable usage on the table because I do not think its implementations of code have been better to Opus 4.8, not even close. It overengineered, obscured and picked awkward constructs all the time over plain, simple, perfectly clean and performant code patterns. Code was smarter AND worse in the kind of way that a brilliant and overeager recent grad often does. (I know I did)

tyre

10 hours ago

As a counterpoint (data point of one code base), I had Fable lead development of a complex system recently (an end-to-end insurance claims billing system) as a test project. It blew me away. Opus could not have done the same, given the feedback Fable had to give when Opus would implement individual features.

Granted, I laid out a document with coding practices, architecture, and technical design recommendations to steer it towards good engineering. And it's a domain I know super well, so I could give very nuanced feedback on trade-offs + architecture. If it had been left to its own devices, maybe it would have over-engineered the h*ck out of it.

But the code it produced—and the implementations it guided Opus towards—were excellent.

robomc

9 hours ago

> It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly

My brother, that's my job.

resters

11 hours ago

over time greater intelligence will be expressed in smaller and cheaper models. we are still somewhat near the beginning of this bc we are finally starting to understand what makes a model truly intelligent/capable.

With Sol we see openai making the model extremely slow and paranoid about process/ceremony. Sure this is a good guardrail against AI going rogue, but it also sets the stage for companies to charge for 2x, 4x, 8x performance, with 1x being barely tolerable and frankly slower than last year's models (though less error prone).

The irony is that the smarter the model, the more it can be trusted to do with less supervision, so one engineer can manage a team of 20 fable subscriptions more effectively than a team of 3 of last year's model subscriptions.

uejfiweun

4 hours ago

Seeing a lot of people in here say that they need Fable for the tasks they're doing and Opus just isn't enough. My experience could not be more different. I seriously feel like Opus-level performance is totally adequate for most of my use cases, if not all of them. And it's probably been this way since, like, realistically, Opus 4.6. On the other hand, Fable I've observed getting into verification loops that just burned so much of my token budget. Combined with the higher cost of tokens from Fable to begin with, I just pretty much never use it for anything.

gpjanik

10 hours ago

"When Moore’s Law slowed in the mid-2000s" it did not, in fact, slow down in the mid 2000s, or at all.

https://ourworldindata.org/data-insights/moores-law-has-accu...

jbstack

10 hours ago

You've selectively quoted the article. The full quote (emphasis added):

"When Moore’s Law slowed in the mid-2000s (specifically, single-threaded performance stagnated), we suddenly had to think about parallelization, architecture, memory locality, etc."

Your link is talking about transistor count. The article is talking about single-threaded performance. Today's CPUs are faster in large part because they have more and more cores.

selcuka

6 hours ago

> Your link is talking about transistor count. The article is talking about single-threaded performance.

But Moore's Law has always been about transistor count, not performance.

sscaryterry

10 hours ago

It did in terms of the traditional more MHz (GHz) is better, but as you've correctly pointed out, not when it comes to actual compute.

hypfer

10 hours ago

Somewhat weird that the article was released today but did not mention GLM 5.3.

If you're telling me to focus on something, why not focus on the actual latest thing that is the same as 5.2 but better? I get the "came out at the same time as fable" thing, but still.. no mention at all?

Yes, weights aren't out yet, but neither are the ones of Fable.

Doesn't feel well informed enough to give advice.

dbreunig

10 hours ago

Ok, buddy.

I can’t host GLM 5.3 yet, so my agents still run on 5.2. But the fact that 5.2 is sufficient and there’s another gen in the wings kinda proves my point, imo.