skeledrew
2 days ago
All that keeps jumping out at me is how they've set it to refuse giving users thinking tokens and prompts for full reasoning in output. Just drives me further away; I may not stop using Claude completely for now, but I'll be moving even more of my primary workload to Chinese providers. That's where openness and freedom is now at.
asabla
2 days ago
> All that keeps jumping out at me is how they've set it to refuse giving users thinking tokens and prompts for full reasoning in output
I keep seeing comments added to code, which reads like reasoning output instead of meaningful words. I see this behavior for both OpenAI and Anthropic models (for several harnesses as well).
But this is a sample of one. And I may be in a situation where I'm more negative to the output from LLMs in general.
dtech
2 days ago
Yeah GPT 5.6 models did it a lot and Opus is absolutely awful on this. It's clearly encoding it's thinking/context into the comments. GPT-6 models seem to be better about it.
rajeevk
2 days ago
What Chinese models/providers are you using for this? I'm hitting Claude's weekly limits much sooner than I used to with roughly the same workload, so I'm interested in trying alternatives, especially ones with strong coding/agentic performance.
arcanemachiner
2 days ago
Get yourself an OpenCode Go subscription and give DeepSeek Flash 4.1 a shot.
A common tactic is to used a big brain model like Opus for planning and reviewing, and a cheaper model for execution.
Amekedl
2 days ago
While it's a common tactic, and I'd vouch for it if you don't really know what you want to code in fact, but if you know what you want to get out of it, I haven't found anything I'd need opus 5.5 for instead of deepseek-flash (flash v4.1 hosted via platform.deepseek.com)
gobip
2 days ago
genuine question: is there an improvement in using platform deepseek com, over openrouter and a third party provider?
villish
2 days ago
Potentially higher cache hit rate at the expense of your data being kept for training.
mfkp
2 days ago
Yes, openrouter providers are all over the place (in a bad way): https://mmoustafa.com/blog/so-you-want-to-use-openrouter/
I've experienced this firsthand and now I generally pin to providers that I trust on openrouter, or just pony up and pay for the real thing.
sktokener
a day ago
[flagged]
asdewqqwer
2 days ago
platform deepseek com's tos enforces use all your data on training, which is a major set back for me.
Other than that, I think they are cheaper last time I compared.
crefiz
2 days ago
So you don't want the model to be better yet you continue to reuse it?... I find this kind of selfish behaviour increasing within us developers, like a fear of avoiding the inevitable
villish
a day ago
Do you use Mistral's models to help improve the European offering?
Otherwise if you don't care about your data being used in training you might as well use the best, which is the US models on a subscription plan. If you do care, use downloadable weight models served by reputable providers, or selfhost.
lukan
2 days ago
In my experience that tactic works well if the codebase is limited in size, or well maintained and separated. Otherwise I do notice a difference also letting fable do the execution, not just the planning for complex tasks.
criley2
2 days ago
In my experience it never works well on any real work. In fact, I'd go the opposite, plan with the dumb model and execute with the smart model because at least the model writing the code and solving the emergent problems is capable.
In my experience (and I've been trying this a bunch): smart planner + dumb executor produces worse code with higher spend than simply using the smart planner to do both.
It's easy to understand why:
- If the planner has truly thought the issue through, properly designed the solution, solved all of the emergent problems, then the final "write" of the code is just a few more output tokens.
- If the planner has NOT truly planned the issue completely, then you're letting a substantially dumber and less capable model make significant decisions, and trusting its problem solving, without having a better model check it.
If you're highly cost conscious (paying for your own tokens and not making any money) then you have no choice but to trade your time and effort for tricks like this to save money by lowering the quality of your output.
But if your employer is paying for tokens: just use the smarter model. You save your time preventing re-work and reducing code review, you save your employer money (primarily from the cost of your own labor and reduced rework), and you get a better output every time (Opus 5.5 mogs Deepseek 4.1 flash in every single way except cost).
gregwebs
2 days ago
The cost is 20-40x less for Deepseek Flash v4.1. If you are just comparing to Sonnet or you aren't paying (your case) then your advice makes perfect sense.
I also agree that its a big mistake to have a flash model implement without a strong model reviewing.
I have Opus plan, Deepseek implement the code, and then review with Opus [1]. In this workflow I am saving a lot of money by having Deepseek do the implementation. Note that the review back-and-forth is fully automated [2], so it doesn't take any extra attention from me.
[1] https://github.com/gregwebs/skills-sdlc/tree/main/skills/implement
[2] https://github.com/gregwebs/skills-sdlc/blob/main/skills/code-review-with-followup/SKILL.mdTeMPOraL
2 days ago
> If you are just comparing to Sonnet or you aren't paying (your case) then your advice makes perfect sense.
Also if you aren't hitting capacity.
Off-work, I use LLMs regularly for both design/coding and non-technical work, but the volume is not enough to trip the weekly limits, and rarely enough to trip the daily limits. So I just go with whatever's current best SOTA available on my Claude & ChatGPT subscriptions and don't worry about limits. If I hit one, I do some household stuff or relax for a few hours (or just turn in for the day), and then the limit is refreshed.
criley2
2 days ago
Deepseek Flash v4.1 is only "40X cheaper" if you do not account for the time of the engineer reading the output. If Opus 5.5 high requires 1/2 of the actual engineer time, and the engineer costs $100-$200/hr, then Deepseek v4.1 is actually the more expensive model to use.
I have tried your workflow many times, and simply letting Opus do the implementation costs much less than wasting hundreds of millions of tokens letting deepseek and opus go back and forth and back and forth. And bonus, my project finishes in 5 minutes instead of 20.
gregwebs
2 days ago
I tried the experiment of reviewing vs. not reviewing with frontier models. I consistently found that reviewing by a model with independent context finds important issues when changes are non-trivial- certainly the definition of non-trivial is getting raise as the models get better.
I do have an /implement-simple workflow to skip the planning phase, but even that doesn't skip the review.
Are you doing your own intensive reviews of the model code? Can you share the prompts you are using as I have?
My bar for what models produce without human intervention is much lower defect than what a human would produce. The human interaction is mostly to guide the design and then the review burden is very low. I suspect your bar for what agents produce is lower- you are taking more of the review burden. I also suspect that you are measuring time more than actual cost since your employer is paying and that you are comparing to Sonnet rather than DeepSeek (DeepSeek 4.1 again is 20-40x cheaper than Sonnet). You mention hundreds of millions of tokens (my reviews don't use that much), but even that costs ~$1 on the DeepSeek side.
I think you are taking exactly the right approach at your employer given the cost is free and you only have access to Anthropic models.
gregwebs
2 days ago
I saw your response before it was deleted- that you are doing multi agent persona reviews and a very intensive review process. So having fewer review items saves you money.
One thing that I have found is that as the frontier models get better there is less need for agents with specialized personas. I actually don't don't use those anymore- I just use agents that have different models and reasoning levels. I have a generated CODING_STANDARDS.md document and a skill for architecture design and a skill for implementing testing [2] that are referenced by a single reviewer. I do implement a 2-pass review though [3].
I would be interested to know if you have found anything similar as models get better. It seems though that you are sharing a single exploration and then sharing the context across the specialized reviewers to dramatically reduce the cost of your approach. Does this have to be in the harness- that is if you write out the shared context to a file does that increase your costs a lot?
I also wonder how intensively are the models able to test their changes? The number one quality improvement I have found is not review but having the model properly test its code. I have a skill that is helping [4], but I also have to spend time to establish a pattern of testing with tools beyond just unit tests. The testing takes significant effort, and this is again where the cost savings of DeepSeek shine.
[1] https://github.com/mattpocock/skills/blob/main/skills/engineering/codebase-design/SKILL.md
[2] https://github.com/gregwebs/skills-sdlc/blob/main/skills/verify/SKILL.md
[3] https://github.com/mattpocock/skills/blob/main/skills/engineering/code-review/SKILL.md
[4] https://github.com/gregwebs/skills-sdlc/blob/main/skills/verify/SKILL.mduser
2 days ago
Roark66
2 days ago
Bingo, DeepSeek (v4.1) is horribly overhyped. In all my personal benchmarks it sits below Glm5.3 Flash. Waaaay below Qwen3.8-Flash-Next a model less than half it's size.
No, the only open weight model that really makes sense for me is Qwen3.8-Flash-Next, but it is mainly because I can run it locally with reasonable speed (prefill between 650-1400t/s generation between 22-50t/s depending on number of slots/users I configure).
This is the first model that truly competes with Opus 4.8. I'd say it may be better than Opus 4.6 on programming.
But it is very verbose when it comes to reasoning tokens. The more difficult the task the more verbose it is. Certain very hard tasks that take opus 4.8 400k tokens take Qwen3.8-Flash-Next 2M tokens... But it finishes them.
And what you loose on the generation speed you get back on input caching you can keep on for weeks.
It really depends on the workload.
smartbit
2 days ago
to mog: to outshine, outclass
Etymology: probably from AMOG Alpha Male of the Group
First seen: 2018
skeledrew
2 days ago
Thanks. I was wondering if that was an autocarrot.
joquarky
12 hours ago
> If the planner has truly thought the issue through, properly designed the solution, solved all of the emergent problems, then the final "write" of the code is just a few more output tokens.
For me, what usually creates the high cost of implementation with a larger model is the validation process, not the writing of code.
I would get excited to see it finish writing code with so little usage, but then it would gobble up ten times as many tokens on validation.
I've tried to instruct it to keep the validation light with the intention of doing batches of deep validation after a few tasks, but it couldn't stop itself from doing heavy validation on each task no matter how I rephrased the instructions.
fosron
2 days ago
Been using DeepSeek Flash 4.0 and 4.1 for some random sideprojects via OC GO, its a great deal and for non-corporate work it's really great!
ffsm8
2 days ago
> hitting Claude's weekly limits much sooner than I used to with roughly the same workload,
Anthropic had a +50% weekly tokens promotion since April (!) which just ran out last weekend after getting multiple extensions.
I've been feeling that too,and I suspect that's the true reason why they released opus 5.5 at a discount
surgical_fire
2 days ago
I have tried GLM on a subscription, and also DeepSeek and MiMo using API directly. MiMo in particular is extremely cheap.
For regular software development they have been pretty great.
reachableceo
2 days ago
z.ai with zcode. it works around the clock for me off of my redmine queue
buckle8017
2 days ago
Don't use opus 5.5 at high. Medium is about as good as 5.0 was at high.
pllbnk
2 days ago
Crazy how tables have turned. Life seems surreal since 2020.
OtomotO
2 days ago
Ironic, especially given 100 hundred years of Hollywood proaganda telling the west that the US are the center of freedom.
Which was and is true to some extent.
And don't get me wrong, China is a dictatorship, and a tyranny for some.
But then again, the west is a tyranny for some.
muzani
2 days ago
That's how the cycles happen. China realizes they could use a little more freedom and US realizes that they could do with a little less. The emerging/shrinking middle class of both countries also moves the sweet spot.
baxtr
2 days ago
The whole notion of seat-based pricing seems wrong to me as well.
hannesv
2 days ago
What Chinese provider would you use that is on par with Claude code?
arcanemachiner
2 days ago
Since Claude Code is a harness that can be made to work with (pretty much?) any model, the answer to the question you have asked is: Claude Code
Non-pedantic answer: I totally agree with you. Opus 5.5 is totally knocking it out of the park IMO.
xandrius
2 days ago
Zoo Code is so much better than CC that to me even using similar models I go for CC for simpler things and ZC for larger work.
chrisweekly
2 days ago
First I've heard of "Zoo Code", on HN or anywhere else - and I pay attention. Got links to share, making the case for it?
selectodude
2 days ago
It’s a fork of Roo Code due to Roo no longer being developed.
exfalso
2 days ago
pi+astra for me. Does absolute wonders. When openai starts to squeeze it's chinese models all the way
bakies
2 days ago
they have started to squeeze, with gpt-6 i'm getting waaaay less value out of the subscription. Used to be thousands of dollars a reset and it's down to a few hundred
intended
2 days ago
I think we need more of these issues to frustrate people.
There is a fundamental incompatibility between “safe AI” and compliant AI.
This is an issue when it’s people, Enron or Madoff for example.
I guess it’s : “safe AI, capable AI, and obedient A. Pick one “
_blk
2 days ago
They may have better and more open weight models but they sure don't have our western understanding of individual freedom. Go try out their first and second amendment protections, or try the fifth? I'm sure we can find more but mostly, when the state needs the tech there won't be an Anthropic-like appeal against overstepping.
shaan7
2 days ago
Yeah its annoying. I need to pay for thinking, but I can't see it :/
gadders
2 days ago
I mean there is a good reason for that, no? Distillation is an issue.
Lalabadie
2 days ago
Is distillation an issue that stops you from picking a model, while scraping/torrenting as much of the Internet as possible is fine?
It's not like Anthropic and OAI have clean hands, especially as they're now racing each other to appear the most dangerous to civilization.
AndroTux
2 days ago
Not an issue for me, the end user.
gadders
2 days ago
Ha, fair.
skeledrew
2 days ago
As a paying user, I expect to get what I'm paying for. I'm paying for thinking tokens, so I should be getting them, and in a way that I can actually read if/when I want without relying on any proprietary tools. I have no interest in being locked in.
wg0
2 days ago
If you're not a noob and you know what you're doing then I can't recommend DeepSeek v4.1 Flash (set to high) enough.
skeledrew
2 days ago
Yeah I've already been using it for some implementation tasks. Works really well given the cost.
kabes
2 days ago
What's special it that noobs shouldn't use it?
wg0
2 days ago
Noobs are burning tokens like "make me an app that does this" whereas an experienced engineer would go with certain language, framework and architecture in mind.
Jeremy1026
2 days ago
How exactly does specifying the language, framework, and architecture in advance save a meaningful amount of tokens? I'd expect saying "make me an app" and "make me an app using Swift and SwiftUI" would be pretty close in terms of token usage. You save maybe one look up by the LLM for "what is the preferred language for writing an application for iOS?".
skeledrew
2 days ago
That's obviously pretty specific to iOS - or MacOS - where there's a single blessed path. Elsewhere, especially in web dev, unless you really don't care about what you get or are making something very simple, you better be ready to provide specifics.
Jeremy1026
2 days ago
Even still, the majority of things being built on the web perform largely the same if it's being built in Ruby or Rust or Node or Go. Only very niche things, that an LLM would probably fumble over anyway, really benefit from picking the perfect language/framework. The only real advantage to naming your language in the initial prompt is that you can be assured that you'll be able to understand the code when the GPU spins down and output is in front of you.
skeledrew
2 days ago
> you'll be able to understand the code
This should always be a goal. Doing anything major without being able to review manually is just asking for pain over time, or be ready to feed more and more tokens to the fire to reduce sloppiness.
wg0
2 days ago
"Make me a ticketing system like Jira"
Now this prompt has huge variety of implementation details. Language? PHP/Ruby/Python/Java/Typescript? In each of them then there are tons of frameworks, templating engines, ORMs, database servers, frontend tooling, bundler, frontend framework alone has several dozen candidates from React, Preact, Vue, Svelte and what not.
So if you really know your craft, you'll already be knowing what specific implementation you need so let us not discount the existing expertise here.
kouteiheika
2 days ago
In general the weaker the model the more skill you need to drive it (at least if you care about quality).
dubcanada
2 days ago
It is not as self thinking, you need to be more detailed and accurate with the prompts.