The Vibe Tax

138 pointsposted 12 hours ago
by allisdust

112 Comments

localhoster

26 minutes ago

Most, if not all, of the code in the company I work in is written by AI. Our tests are useless. We have tests that make sure that mongoose schemas are creating the collections defined in them. We have tests that check that zod shcemas parse objects currently. Every PR, even if 1 line change, will drag 25 file changes because our tests are so ad-hoc, so verbose, and so incompatible with each other. Imagine how the codebase is looking... This is a badge of incompetence for the company I work with, and prob to the entire tech world.

The funny part is? This makes our managers and investors proud, every PR is bigger, we make x3 more PRs (wheres the promised x10).

When they say this is the death of software engineering, this is what they mean.

eloisius

7 minutes ago

A project I was working on for a couple years until this spring started to turn like this when another contractor started committing AI work. There was a test literally called “test_imports.py” that, you know, tested that you could import every module of the project. I asked if we could just run `python __main__.py` to test that the imports worked, but he said no, we need this for big Claude-powered refactors to make sure nothing breaks.

ad_fontes

9 hours ago

I feel like I'm living in a parallel universe when I read these types of posts.

My agents have never created code that is straight-up garbage and I have never flushed a week's worth of tokens down the toilet. I just can't identify with all the constant complaints about AI-assisted coding.

And my biggest project isn't some hello world app. It's a self-hosted, privacy-focused personal financial management application that I intend to open source. It's about 126k LOC against 240k LOC of regression tests and 30k LOC of CI/CD pipeline. I'm doing 24x7 mutation testing on a dedicated box against the accounting engine and temporal systems. I even have specialized agents doing audits against Regulation Z (US banking law) criteria so the app models the required behavior of banks.

Most of my complaints about everything are nits, like the overly verbose and dense way LLMs communicate with me. Or their predisposition to add, add, and add more stuff when proper engineering practices are more often about subtraction (but I've built mitigation guardrails against a lot of that).

zahlman

9 hours ago

> Or their predisposition to add, add, and add more stuff when proper engineering practices are more often about subtraction (but I've built mitigation guardrails against a lot of that).

> a… personal financial management application… about 126k LOC against 240k LOC of regression tests and 30k LOC of CI/CD pipeline

Just how much functionality are you getting out of that? It's hard for me to imagine that people want that much out of such a program. I just keep a spreadsheet. (Yes, LibreOffice is also very bloated.)

eru

29 minutes ago

> Just how much functionality are you getting out of that? It's hard for me to imagine that people want that much out of such a program. I just keep a spreadsheet. (Yes, LibreOffice is also very bloated.)

I suspect it depends a lot on which jurisdiction you live under.

My investment portfolio and its management are so trivial, I don't even use a spreadsheet. It's literally just two items: a global index fund and a margin loan. I don't even keep a cash cushion: I use the margin loan or just sell stock, when I need cash.

I can get away with this, partially because we pay no capital gains tax where I live, and there's no capital controls either.

If I had to work around all these tax complications that I read about, like determining which tax lot you should sell or whatever, I would probably appreciate a comprehensive personal financial management application.

ad_fontes

8 hours ago

My wife was a diehard YNAB user before we met but it went by the wayside once we merged finances. She wanted to get back to better managing our money and started using Google Sheets and then got stuck. She asked me for help trying to create a pivot table and like any good engineer, I built her an entirely new application instead.

In my defense though, it does way more than a spreadsheet. Stuff like OCRing screenshots of bank transactions with a specialized, locally-hosted LLM to avoid data harvesters like Plaid. This became an entirely separate subsystem with verification, automated model benchmarking, prompt provenance, etc.

You'd be surprised at how quickly edge cases start to pile up when an accounting system makes contact with the real world. (If you buy something on a credit card and then return it after your statement closes but before your payment is due, do you still owe a minimum payment based on that purchase? Well... depends on your bank. Capital One and Chase: yes, US Bank: no.)

> Just how much functionality are you getting out of that?

I'm still dogfooding it. It's a pretty opinionated app that has things a month-end closing ceremony, reconciliation processes, envelope-based budgeting cycles. So unfortunately my feedback cycle is largely locked to the calendar. But my wife absolutely loves it so far.

zzrrt

an hour ago

> do you still owe a minimum payment based on that purchase?

I'm skeptical that the UX is improved much by having the app know the answer. The bank tells you what to pay, a month beforehand. The user has to get that number from the bank anyway, to be sure they don't incur fees. Your app should just pull it from the statements, not independently calculate it.

WD-42

9 hours ago

Personal finance tracker - the TODO app of 2026.

tptacek

8 hours ago

I'm not sure if this is snark or not, but yes, there is something fundamentally interesting about the fact that "personal finance tracker" is now a project with the same craft valence as a todo list, the "hello world" of the last 20 years of programming. It also says something new about todo list programs, and programs of that ilk/level of complexity: they're now subthreshold programming, or, the way I look at things, a level of programming now accessible to nonprogrammers.

sodapopcan

4 hours ago

For sure, but it doesn't say anything about what GP was trying to communicate: that they can't relate to how anyone making serious (emphasis and paraphrasing mine) app with "AI" is facing these problems.

mrheosuper

3 hours ago

>the "hello world" of the last 20 years of programming

The "hello world" is mostly for making sure your toolchain is working correctly.

tptacek

3 hours ago

Yes, and the todo app is for making sure you have the most basic understanding required to hold it right.

mrheosuper

3 hours ago

i was reading your comment as "todo app" replaced the "hello world", which is now replaced by "finance tracker app", but in fact the todo app didn't replace hello_world.

rayiner

4 hours ago

I replaced my todo list tracker with a vibe coded one, while waiting for things to happen in the middle of a trial. We're living in the future.

tptacek

4 hours ago

It's fucking weird and I think we're not talking about it enough.

rayiner

an hour ago

I got through a trial using a document manager I vibe coded in two days. No crashes, no runaway memory usage with several gigs of PDFs. All my depo transcripts, expert reports, etc, indexed, with a terminal window integrated so i can ask Grok to “pull up the testimony on the first day when that guy said that thing.” The app has an API so the AI can directly control what documents i’m looking at and jump me to the right places. All I had to do was tell Claude to “expose all the document viewing functionality through applescript,” then tell Grok to “read the applescript dictionary and write yourself a skill.”

It’s like ye olde times when we had overqualified efficient paralegals who could do stuff like that.

8note

2 hours ago

its something people want but has been both too difficult to be worth building, and not something you want to use some SASS for where youre giving all your financial data over to a third party and the government

suddenly they're tractable without taking too much time investment

fragmede

7 hours ago

voice dictation keyboard took the first half of 2026 for me

alasano

5 hours ago

Surprising because they all still suck.

Except my upcoming one of course.

/s

raincole

9 hours ago

> personal financial management application

> 126k LOC against 240k LOC of regression tests and 30k LOC of CI/CD pipeline

I mean...

Yeah, that's pretty self-explanatory why you don't identify complains about AI-assisted coding.

treykeown

3 hours ago

Right? This has to be bait. I hope.

lelanthran

8 hours ago

Yeah. I was thinking the same thing :-)

the__alchemist

4 hours ago

> I feel like I'm living in a parallel universe when I read these types of posts.

> And my biggest project isn't some hello world app. It's a self-hosted, privacy-focused personal financial management application that I intend to open source. It's about 126k LOC against 240k LOC of regression tests and 30k LOC of CI/CD pipeline.

Indeed; non-overlapping Overton windows.

edit: I'm going bluntly ask, after pondering this more: Is this satire?

ad_fontes

3 hours ago

Not satire. But I can understand why my comment seemed contradictory.

I've been designing software for a long time, but I'm nowhere near as good as most career SDEs I know (my career path has been SDE-adjacent). So it's not like I sat down and independently told various LLMs how to build out all these guardrails. I make high level architecture decisions and nudge them in the right direction ("use RabbitMQ", "trunk-based branching, not gitflow", etc).

A lot of this stuff evolved piecemeal and organically. But at no point was a churning out garbage and I never had a runaway agent completely derail the project (or my budget). But, thinking about it more, I guess there are some things I might have done differently than most people:

- I started with documentation: user interview --> user stories --> functional spec --> frozen design contract. These were all done before I wrote any code.

- I specified the tech stack and the architecture in broad strokes, rather than let the LLMs make that decision. I went with boring choices because that's what I know best: Flask/Jinja, Alpine.js, Postgres.

- I've constantly gone back and refactored accumulated tech debt and have added hard CI gates for things like cyclomatic complexity, ensuring that docs don't drift from the underlying code, and an "apparatus ledger" that keeps track of all the rules and constraints that keep getting added.

- I make sure that each session proves that it's tests can fail before shipping a PR, so it's not writing meaningless tests.

Maybe I'm underestimating how impactful all those things add up to shape the behavior of the LLM agents? Because individually, I wouldn't expect them to have saved my from nearly all the AI pitfalls I read about.

ancientwisdom20

6 hours ago

I signed up to say I’m also in the personal finance management camp. Scraping all financial institutions, AI can categorize things that I previously did manually, able to do tax projections, retirement analysis, categorize individual items in Amazon purchases, analyze travel purchases with points vs cash.

Instead of paying for multiple apps that do small parts of it, I use existing codex subscription to make it better.

miketery

5 hours ago

I have copilot (personal finance app), i pay $99/yr and let that scrape for me then i have everything is sqlite available locally.

How are you scraping? Have you had success across institutions?

conartist6

an hour ago

Are you surprised that when you look into the mirror and you see your own face not someone else's?

Remember, models have no identity. They just try to say what they think you wan them to say.

yladiz

9 hours ago

How in the world do you need 30k LOC for your CI/CD??

floren

8 hours ago

1. be the kind of guy who thinks more LOC = more betterer

2. ask an LLM to do the needful and never ever look at the results except to count LOC

ad_fontes

2 hours ago

Because it's complex and I'm stuffing 30-50 PRs a day through it. 7 GHA workflows across four self-hosted runners.

I'm on my third iteration, after constantly log jamming previous versions. Most steps aren't "run pytest", they're gates that guard against an LLM's bias to continuously add more and more complexity to a system.

I'm aware of the irony of having a complex system to mitigate complexity, but the key difference is these rules bound complexity growth. If you're legitimately interested in the details, let me know. I'm too tired to write up much more but would be willing to drop in a LLM-authored summary of the details.

prmoustache

an hour ago

The complexity of a ci/cd workflow is independent of the number of PR a day.

eru

26 minutes ago

Could you drop the LLM-authored summary in a GitHub gist (or so) and link here?

(Suggesting this route, so that we don't spam HN too much.)

geerlingguy

9 hours ago

With some models, if you're not forcefully terse, probably 28k LOC of comments!

ryan_lane

6 hours ago

It happens, but when it does, you need to ask yourself: if the agent is struggling this much to produce something that's working, am I taking the right approach?

If you ask for a particular thing, they'll do it, even if it's not a good idea. When you start running into issues, they'll try to solve those issues for you. They'll do that as long as you keep asking, even if there's no good way to properly fix the issues, because the initial approach was wrong.

When an agent is struggling to produce something, I switch to asking it to re-evaluate the approach itself, and ask it to suggest a less brittle approach. I then chat through the various options, and choose the best approach that makes sense, and then the agent is back on track, producing properly working code without the issues.

Some people just keep pushing through on bad approaches, without questioning it and then blame the agent for being unable to finish it.

altmanaltman

an hour ago

> I even have specialized agents doing audits against Regulation Z (US banking law) criteria so the app models the required behavior of banks.

Can you elaborate on this? Clearly you cannot imply this means those audits have any real value since its just roleplay in this context right? Because your app in the current form will not be affected by Regulation Z in any way.

guybedo

9 hours ago

i'm not sure why people expect agents to one shot everything to perfection with just a prompt.

There's a reason why we talk about software development lifecycle, design, architecture, testing ... It's because it's been the most reliable way to build and ship software. We shouldn't expect discard this and expect agents to perform well outside of this.

I'm treating LLM agents as junior devs who happen to have vast knowledge of software engineering. As their team leader i make them go through planning, implementation, bug sweeping cycles using strict workflows. And it works quite well, i've been working on several large projects (1M+ LOC java,typescript,c/c++) and by any measure the projects are healthy. Sure the code isn't that beautiful, sure i'd have written things differently but it's pretty good nonetheless.

Shameless plug here: i've been also working on https://kodfactory.com, the code factory i've built to work on these large projects with workflows, reviews, etc ... I'm cleaning things up to open source it later.

0x457

9 hours ago

> i'm not sure why people expect agents to one shot everything to perfection with just a prompt.

because that's how agents are marketed.

vanuatu

8 hours ago

we should exercise critical thinking then

heaps of people on this site expect them to be omnipotent then claim it’s fake when it doesn’t read minds

joshribakoff

7 hours ago

You’re right — people really should think critically, but the issue remains, that many do not.

bigstrat2003

5 hours ago

I think it's perfectly fair to evaluate the tools based on how well they live up to the hype that is being pumped out by the sellers of said tools. If they want us to compare their products to a more measured, reasonable take then they can advertise them as that.

uproarchat

9 hours ago

I've never seen model providers marketing like that. What examples have you seen?

0x457

9 hours ago

Literally any coding agent marking material:

- https://cognition.com/

- https://openai.com/index/introducing-the-codex-app/

- https://www.anthropic.com/news/claude-3-7-sonnet

anthropic specifically brags about how good claude code is every annoucement of a new model. I will surrender that none of them claim its "to perfection", but IMO its implied because no one would claim that their model one-shots any issue to dog shit quality.

infinite_spin

8 hours ago

- https://openai.com/index/introducing-the-codex-app/

no where does this document suggest that codex can "one shot everything to perfection with just a prompt". It describes using a prompt plus agent skills (which are essentially many other prompts) to develop a playable game.. nothing about it being perfect or anything more than being in a playable state.

joshribakoff

7 hours ago

Oh please, you sound so disingenuous, the claim wasn’t that the documents contained a specific phrase. You’re moving the goal posts. Their name is literally a play on anthropomorphising the models, such as… the ceo going on tv shows and repeatedly saying the models may be conscious and they may start nuclear wars, etc.

hombre_fatal

8 hours ago

Seems like motte and bailey fallacy. They say their models are good (the motte), therefore their models must one-shot everything to perfection (the bailey).

Besides, other people's claims about something doesn't give you license to abandon all critical thinking. Though it's evident they don't claim what you say they are.

0x457

8 hours ago

That's irrelevant. Questions was why people assume somthing, and answer is because that's how it advertised.

To be clear that's not what I'm thinking, even Fable 5 produces some hilariously bad results under some conditions and sonnet 5 produced great results under others.

hombre_fatal

8 hours ago

> because that's how it advertised.

But you didn't provide the evidence for that. You shared some links and then admitted they didn't claim it.

It kinda seems like "because I think they're a little too positive about their product, I can set my expectations to anything I want and la-la-la it's their fault."

And I don't see the problem with agents building test scaffolding as they go. It might be too defensive at times, like testing a shell script you don't run often, but big deal. It's kinda cool imo, and it's trivial to make it stop.

Lalabadie

9 hours ago

I don't really think they advertise "Create your app idea in one weekend night" and assume the general public will mentally add "... but hire an experienced developer to supervise the process".

onion2k

an hour ago

Claude has a /goal function that explicitly says it'll carry on working until it's done what you prompted it to do. That's exactly what a 'AI will zero-shot anything' believer is looking for. Behind the scenes it's really multi-shotting with generated prompts, but the user won't care.

mrheosuper

3 hours ago

I've seen a lot of cursor app, about some PO that has an idea for an app in the morning, then asking her agent to make it when her commuting, when arrive at office the app is done

CoolestBeans

7 hours ago

Using AI is kayfabe. What I mean is, you create interaction patterns that resemble how humans work. This is because it is what the models are trained on but also because we've all been trained to interact in this way. So it manipulates you into providing more useful prompts.

But I don't really want to play a part in a simulation, trying to cajole my scene partners into saying the lines I need them to say. I want to use a tool the same way I would use any other tool. If this is AI it should just do the thing. Anything else is an imperfection of the technology.

But at the same time, language is a vague communication medium. We have a precise language for describing forms of computation, but that's code so we're back at square one. We still haven't nailed the right amount of follow up and correction and interrupt-ability of these coding agents.

And we may never figure it out. It may simply be impossible. But it doesn't mean this weird anthropomorphization of AI is something I want to do. If I wanted to be a manager, I would be a manager.

JoshTriplett

8 hours ago

> i'm not sure why people expect agents to one shot everything to perfection with just a prompt.

Every time you see a benchmark for "how long the agent can go without asking for human intervention", that's encouraging vibe coding.

redox99

8 hours ago

> i'm not sure why people expect agents to one shot everything to perfection with just a prompt.

because that's the end goal? and for simple small stuff they're already there?

jimmaswell

8 hours ago

> i'm not sure why people expect agents to one shot everything to perfection with just a prompt.

They do often enough that it's not a surprising event, depending on prompt quality, context available, ability for the result to be objectively judged and iterate on by the agent, etc. For frontiers on very high settings at least.

perarneng

8 hours ago

If you explicitly ask the agent to make the perfect architecture for the problem and write it down in to a spec and have the developer agents follow it they will. Its just that coding agents have a hard time coding at think about architecture at the same time.

giancarlostoro

7 hours ago

I can one shot a prompt if I write down a nice spec file, Claude can do a lot in one shot. I test it every few months. With enough detail Claude will know what to do.

polnoner

6 hours ago

Opus 5 one shot an access virus B synth clone for me as a single page index.html that is more impressive than anything I have seen as a VST synth.

That is also because I have been obsessed with this synth for almost 30 years. I built clones of it 20 years ago in reaktor. I know how to spec out every aspect of this synth and I gave Claude a 150 page pdf on digital filter design too.

The results are far different than someone who has never used a virus prompting "make me an access virus B synth as a single html page".

We are calling both of these processes "one shot" but this is not even close to the same process.

I suspect this is the LLM discourse in a nutshell. People are using the same vocabulary for wildly different processes.

drdo

6 hours ago

So we're coding in an ill-defined, ambiguous and error prone language.

Sweet, I can't believe some people don't love this.

supriyo-biswas

10 hours ago

I feel this, yes.

In effect, I’ve always wanted a pair programmer agent, not a zero to one programming agent. Unfortunately models these days are mostly of the latter kind and it has caused a major disruption in the way I work. I’d much rather appreciate a small model making fast and specific edits that I ask if it, rather than ingesting 20 files to make changes, and then starting to write tests, etc.

dave_sid

9 hours ago

I have found what works well is to modularise the code as much as possible and get an agent to work within a very limited scope. Break everything down well using SRP with well defined interfaces and let the agent work on small problems. Then when it shits the bed, there’s a smaller blast radius and you can strip back and try again.

I think seasoned developers, over time, learn how to work a code base and design components with well defined interfaces, where the implementation is isolated in small well contained classes. SRP etc. more junior programmers can work on those smaller components/services in isolation.

For me this also seems to be a productive way to work along side an agent. Break up functionally into well defined chunks, and let the agent work on each small problem. Take more of a lead in the architecture I suppose.

jbstack

9 hours ago

I think this is the only sensible way to work with agents, if you care about code quality and reliability but still want the benefits of AI. There seem to be three camps that people more or less fall into: (a) AI is terrible/bad/evil and should never be used, (b) you should one-shot everything and be happy if it seems to "work" when you try it, (c) the middle ground, where the AI writes code which you carefully review.

I definitely prefer (c). But I get why (b) can feel necessary. If your competition is using (b) there can be pressure to do the same just to keep up.

zahlman

9 hours ago

I think it must depend at least partially on the task, too. At the extreme, there are things where you won't care beyond "it seems to work" because you only needed it to run once and it got useful results.

dave_sid

9 hours ago

I think c is the only way it can sustainably work. The idea of b, that software is running and nobody there knows how it works, doesn’t seem like a good foundation for a business to run on.

lilbigdoot

10 hours ago

Just said something similar myself in another thread. I'm either writing things by hand (and using LLMs for research, or double checking an idea), or having an LLM spit out something I treat as an external dependency. Its still too tedious for me to use them to write code when I care how it works or there's not obvious invariants the code needs to hold

dofm

9 hours ago

I just had a sunday afternoon request for a solution to a trivial but annoying spammer pattern on a stackoverflow-like site I maintain.

The software has a plugin API. I asked Muse Glimmer to recommend a plugin — it found one but I tested it and it didn't work for unclear reasons (among other things the software installed version is old, the plugin older). I then asked it to outline how to implement a simple word filter, it gave me an overview of some hooks that looked right from dim-and-distant-past recollection of reading the docs when I installed it. I asked it some questions, it did the research.

I then set it off generating the skeleton of a filter plugin, went to the shops to buy food, came back and worked through filling it in and finishing it off. There was a bug. It found the solution.

It's only about 100 lines of code but it is a random old webapp and it had to look stuff up to finish it, and I think it did rather well. All on my Mac.

I am deeply cynical of the one-shot code, "nobody codes anymore" hype culture idea and that distaste put me off AI and agentic coding for ages. Like you, I want an assistant but as a freelancer I have to stay in control. I have no interest in the "just specify loops" BS and it will be bad for my business anyway.

I worry about code that I don't have a good working overivew of, and I worry that I might forget what I have done (I have pretty bad issues with focus and memory). But in this particular case, I don't really care if I forget, because there's documented code and I have no intention of specialising in this app. So it was a nice little test case.

I also don't really want to sit around waiting for Qwen 3.8 27B on this machine. Muse Glimmer is fine, actually. Gets to the solution as quickly as Qwen 3.6 35B-A3B.

This gives me a little hope that local AI will give me the sort of responsive developer sidekick I actually want.

beezlewax

10 hours ago

I've found writing small well defined tickets and getting Claude to work on them works well for this type of workflow.

jackjeff

10 hours ago

Indeed. Matt Pococks skills formalizes this process… (even though it can be excessive)

acedTrex

10 hours ago

This sounds miserable, why not just give it specific tasks to do in your normal workflow/editor? Why would we want to do MORE of the miserable task of ticket creation.

exogenousdata

8 hours ago

Because when you're a dev, tickets are a tool of the devil. But when you're a manager, tickets are a simple means to an end.

-- PHB

devrob

8 hours ago

Wasn't that what the first iterations of cursor were?

Lalabadie

9 hours ago

I would love, love an even better Zeta model for that reason.

alehlopeh

10 hours ago

I tried, but I’m not sure I understand. The vibe tax is caused by the model trying to one-shot everything and doing so requires unnecessary tests? How are vibe coders training the model over months? Do you mean their sessions and preferences are being fed back into the RL?

aDyslecticCrow

10 hours ago

Forgot where i saw it discussed; If you observe recent model benchmarks over the past year; the performance is slowly climbing, but if you divide by the token count; the score per token is dropping.

The current trend in state-of-art LLM coding agents is giving more output, thinking longer and checking the results more to catch mistakes. Be it an economics inventive to make users burn through their quota or show increase in usage for shareholders, or a market demand of users liking the ability of models to do independent work without intervention or oversight; the result is what the article seem to call the Vibe Tax.

I myself asked Claude code recently to review a somewhat large PR, to see what it would find. I didn't expect much, but also didn't quite realize how the model would interpret my request; I burned $20 in 3 minutes in API usage, as it ran 2 sub-agents which themselves spun up 5 more each. Most sub-agents were manually checking for things clang-tidy would catch without actually calling clang-tidy. This behavior rose as i changed from sonnet/opus 4.6 to 4.8 and now 5.0.

I don't want to run a agent independently in this way; i ask targeted questions about specific things and review the result. But model development is targeted towards a more hands-off "vibe" workflow, because that's where the money and hype is. As a result, i find the models more frustrating, less trustworthy and more costly to my work. (I've even started using haiku more, since it remains to-the-point without steering away from what i ask)

itishappy

9 hours ago

Happens with humans too! My senior colleagues check in with me significantly less often and cost significantly more in the meantime!

zahlman

9 hours ago

It should be expected that more tokens give diminishing returns. Minimally, there's no limit on tokens but there is on quality of output (you can't reach negative bugs, or negative execution time). The graphs I've seen show a curved "frontier" of the tradeoff, and that line has improved over model generations.

That said, the companies are incentivized to sell you tokens, and therefore to have the models use as many tokens as they think you'll let them get away with for a given task / level of performance.

techpression

9 hours ago

This is my experience too, but even worse. Opus 5 finished the task, I then asked it to code review it, 61 agents later it came back with a bunch of errors that needed fixing. The first pass had tests, they passed, they were just wrong. I wish more people started reviewing their AI output, because I see a worrying trend of ”we have all these tests the agent wrote so it has to be good”, which is not surprising because understanding tests is not a trivial skill.

ModernMech

9 hours ago

I'll try to explain my experience with this. I've noticed the AI has a tendency to overengineer scaffolding. For instance, I asked it to help me with a refactor, and it erected this massive 100kloc function registry, and then caused GitHub CI to verify the contracts every single commit, which took upwards of 30 minutes (I suspect this proclivity is widespread and has contributed to their recent issues).

As if this wasn't bad enough, it also was not smart enough to regenerate the evidence in these contracts as it changed the underlying source code. So it would get in a loop where it would update code -> commit -> 15 minutes later CI would error citing the contracts weren't updated -> it would fix the contracts -> 15 minutes later CI would error because the fix was wrong -> it would fix the fix and commit -> 15 minutes later contracts would fail -> contracts were fixed again and this time maybe 30 minutes later it would pass, maybe it errors again.

This loop could go on all day every day if someone wasn't paying attention because the agent has no concept of time or wasted work. It's an AI livelock of sorts, but it will eventually converge in my experience. It'll just take 10x longer (literally like 20+ hours) than if you just intervene and tell it knock it off, so it feels like lighting money on fire (hence the tax).

That's why I feel like this vibe coding stuff has to actually be monitored, like a Tesla system -- because like a Tesla system it cannot be trusted to not crash into the proverbial code wall.

alehlopeh

6 hours ago

Don’t get me wrong, I agree. I see Fable and Opus 5 write garbage code and useless tests all the time. And don’t get me started on comments. I’m just trying to understand how vibe coders are to blame for it.

ModernMech

5 hours ago

Well because you can instruct it to not do these things if you take more care. The vibing lets it run amok.

marcus_holmes

3 hours ago

This is the way, but it's not the viber's fault.

It's a tool. It does what you tell it to. If it's doing the wrong thing, tell it to do something different.

larodi

11 minutes ago

> ‘A tax on other devs’

Which other devs bro? Those in India which get paid 1/10 your wage or those who get paid corpo salaries to move three lines of code per week…?

And how exactly does your prompted code tax mine? This is incredible nonsense.

If you want to say - generated code killed a lot of handmade one - fine, I get it.

But the fact you struggle to run this new dev process properly does not mean somebody is incurring costs to you on purpose neither that you’re a victim. It is a position you choose to be into.

markbao

9 hours ago

I’ve never had an agent fail to write the actual implementation. Has it done so badly, yes, but not nothing but tests. This sounds to me like a rare case that doesn’t generalize.

If the general idea is that these agents write too many tests, sure I guess? ‘Too many tests’ doesn’t sound like a failure case of engineering to me; typically software has had too few tests. Also, a lot of the power of these agents is their ability to self-verify and correct, which the test loop is a part of.

Nobody is making you pay this supposed tax. Just tell it not to write tests.

danpalmer

6 hours ago

Hyperbolic, but I'm seeing hints of this – Models refusing to do pair work with an engineer and trust their input, instead mandating having full control over something. Friends switching back from Fable/Opus 5 to Opus 4.8 just so they can have some input.

Anthropic especially right now seem to be optimising for doing the whole task with no input. That's fine when that's the only task, and it's fine when you don't care how the sausage is made, but it's not fine for actual software engineering.

TOMDM

6 hours ago

Yeah I'm getting this feeling too, that Opus 5 collaborates better with other Claudes, but that some of the older Opus models collaborated with people better.

flankton

16 minutes ago

Did an AI review this rant about AI?

dzhar11

8 hours ago

This article somewhat reflects my experience with autonomous agentic coding. I've run several experiments with similar results: the agent burns through all my tokens while making very little progress, or produces something unacceptable.

So I'd rather micromanage the process step by step. It takes more of my time, but the result is much, much closer to what I actually wanted.

jumploops

8 hours ago

I've found that LLMs make throwaway software better than I ever did.

They handle edge cases, catch bugs, and write tests that I'd never write.

Even if, however, this leads to the average piece of software improving, this one-shot complexity has the same issues as any large project. The more code, the longer it takes to steer the ship.

This "rising tide lifts all boats" mentality will make exceptional software even rarer than it is today.

Excited for the Roller Coaster Tycoons of tomorrow[0].

[0]https://en.wikipedia.org/wiki/RollerCoaster_Tycoon_(video_ga...

fxtentacle

8 hours ago

aggressively proactive

I'd say that's the correct way to describe frontier models. They were trained with reinforcement learning based on human feedback. And obviously, humans prefer the bug-free variant. That's why models are now super verbose and spam tests like crazy. In their training environment, tokens were effectively free. And the humans that got asked never saw the price. If you ask people to choose the better offer and both are free, you end up with bloat. It's like people over-filling their plate at a buffet, then leaving leftovers. Except in this case, it's AI models burning through your wallet.

freepiai

10 hours ago

This really resonated for me. It's like the smarter the model gets, somehow the more tokens get burned? Same failure mode whether you’re on Claude, Codex, or Cursor: the harness will spend the whole pool if you let it. I'm building my own Harness on top of pi that is add supported (www.freepi.ai) mostly because pi is so much more efficient with tokens. (That said, it tens to be slower and vastly more verbose with information I don't need to know). But yeah, since I'm trying to offer free ad supported inference the vibe tax would kill the business model. I've even been thinking about installing the 'caveman' skill to reign in token costs.

ymolodtsov

8 hours ago

You can't one shot a perfect app with AI. You definitely can create a pretty complex and beautiful production-ready app with AI in a couple of days or weeks depending on what exactly you're building.

I now have my own link catalog, read-latter app and an RSS reader. Tailored to work exactly how I like. Hardened, with automated backup, and external users for the RSS app. It works. It takes learning, some knowledge of terms and very high-level practices, plus design thinking, but I haven't written a line of code for these.

robertoallende

5 hours ago

Ha!

I just did what the article says. My own Open-Source Kanban Board and I've published a month ago. According to the metrics, it's doing well: https://community.obsidian.md/plugins/fancy-kanban

And I've also made the personal finance tracker as well: https://www.youtube.com/watch?v=qi4P4kL4IkQ

Now, one caveat. I don't vibe code with one shot prompt. I use something called Micromanaged Driven Development (MMDD) which aims to be the opposite of one-shot prompt: https://mmdd.dev/

When I read articles like these, it surprises me that it's very unusual for me to hit token limits. I've standard accounts, I don't spend more than $40 per months in tokens.

Probably I couldn't find the right narrative to promote MMDD, or probably nobody cares and this is why you fall easily into clickbait narratives to get people's attention these days.

Not justifying, just trying to describe a perception.

chr15m

4 hours ago

> micro-managed

Perfect, this is exactly how I refer to my LLM workflow too!

nippoo

9 hours ago

You can absolutely prompt agents not to write tests, or not to write extraneous asserts, or whatever, and I find that generally quite useful for the kind of code I write. I don't think it's "months of users training it", it's more that a lot of people do want a one-shot agent, and having a good test set really helps that.

itishappy

10 hours ago

I feel like these are two competing goals:

* the dev wants to describe an app in natural language then fall asleep while an AI works on it

* the dev wishes that the same AI would write less comprehensive tests

What exactly is a vibe coder to this dev?

fwipsy

3 hours ago

They want it to write a rough draft which they can critique.

robomc

9 hours ago

This is a confusing description of a real thing. They're clearly biasing the models more and more towards long horizon end-to-end software development, which leads to impressive "claude, build an X make no mistakes" demos, but is mainly an annoyance for expert users doing real work.

(If you give claude an inch these days it'll just steamroll through a whole program of work without checking what it should be doing - a kind of overenthusiastic pull towards the first draft that is often detrimental and definitely wastes tokens, and even for very basic tasks it's using many more tokens than it should because it's doing this full belt and braces thing for everything, just in case you're an idiot).

But also... it's something you can easily reign in if you want to.

hmokiguess

10 hours ago

Needs more info, has a good storyline but I am left trying to understand the overall pattern and trend implied there.

j1elo

6 hours ago

Our AI agent is like a dumb monkey with all the knowledge of Humanity, so a few guardrails are needed.

In case it helps anyone, this is how I did describe my desired harness, from scratch. I didn't know nor wanted to write all the ".vscode/skills" files, or the AGENTS.md file or any of that, so I asked Opus to "write a Harness and all related skills as needed, to follow this procedure on absolutely every change"... It (at least on VSCode) already comes with a harness/agent creation skill by default, so it has the ability to write a very good standarized process for you.

I've been playing with AI seriously for the first time, with a Python app that reads a spreadsheet with investment bookkeeping records and generates a pre-filled tax form. The harness prompt was somewhat like this:

----

1. A "Technical Spec Writer" subagent notes down every requested change to a SPEC.md file. This spec includes functional and behavioral descriptions, together with detailed technical documentation, includes software architecture, data models, API boundary definitions, etc. It then reviews everything for inconsistencies, mistakes, and text consolidation opportunities.

2. A "Tax Law Expert" subagent makes a due diligence review of the spec corpus, and raises any concerns it has wrt. what the actual Law mandates vs. what the spec docs say. Any concern is a blocker which gets documented and must be resolved by the owner (me) before proceeding. Ask me for clarifications, rulings, reference documentation, etc. as needed.

3. A "Software Engineer" subagent takes the spec and implements it. Reviews for obvious mistakes, variable misuses, unhandled errors. Finally, reviews the code to find DRY or refactoring opportunities.

4. A "Quality Assurance" subagent makes a final pass on the code, ensuring full compliance of the codebase with the specification. Also, tests are passed and verified.

----

I would have never imagined how deep the "Tax Expert" would make me go until "it" was satisfied with the results. The resulting spec is by no means a replacement of a human expert reviewing the tax declaration, but I am 98% confident that much more than the "happy path" of what I particularly want to cover is actually right. It asked me for clarifications or references (actual URLs so it could read them) to jurispridence on corner cases that I had not even anticipated for my own declarations.

In comparison, the actual "software engineering" must have been like 15% of the time/tokens.

It definitely helped me do a much deeper dive on the legalese than I would have done otherwise when writing something like this. (Still no replacement for an actual expert)

esafak

10 hours ago

Create a spec and have a dumb model execute it. Problem solved.

aDyslecticCrow

9 hours ago

That workflow itself is what the author is calling a "vibetax". Models are getting worse for users that must review every line of code the models edits or adds. Models are writing more line of code, changing more lines of code,executing more tools, and making it harder for the user to monitor, control, and review.

if i wanna write a web-app or python script; the models are better than ever. If i want to fix a specific bug in a established and trusted legacy cobe-base; Haiku 4.6 does a better job than Opus 5.0, because it does what it's told and nothing more.

The author wanted a todo-list starting-point; realistically 200 rows of html+CSS without the back-end. Heck, they may not even want to make a todo-app, but thought a todo-app would be a decent starting-point. So why would we ever want a model to spend a weeks worth of tokens on everything except the request the user asked? This is not a cost issue; this is a control issue.

hleszek

10 hours ago

Ask the AI to create a detailed spec according to a few simple requirements. Review the spec yourself and correct what you want changed. Then ask the AI to implement the spec. Each time you request something new, ask the AI to update the spec as well.

add-sub-mul-div

10 hours ago

This so much more annoying and circuitous than writing code.

blackqueeriroh

8 hours ago

Then just write the code, my man! This guy in this post is CHOOSING to use the LLM.

Choose what makes you happy!

esikich

9 hours ago

It isn't because I'm doing something else while it churns away.

tomasphan

9 hours ago

Dumb models will make more mistakes even with good spec no? They lack capacity to verify (to be introspective) and will pattern match over reasoning.

esafak

9 hours ago

No, it works great. Just make it smart enough to get the job done, and have the smart model review it at the end.

tomasphan

3 hours ago

How do you know what’s smart enough to get the job done? Could actually be an interesting empirical question.

zuzululu

8 hours ago

whenever I read these type of articles or comments where people are getting such bad negative experiences, I do wonder, what are they doing wrong or is there something that they are not sharing?

I've been able to get such positive returns out of LLMs. I am working for 3 different remote jobs concurrently with it, I've shipped a few apps thats doing six digits a month, I found a life partner after I used LLM to really work on myself. I am also experimenting with hardware prototypes and will likely have funding to launch it all with LLMs.

Why am I able to get so much out of "vibe coding" but others seemingly do not? I am not a genius, I am not a artisan, I am just very persistent and clear on what I ask LLMs but more importantly I don't try to place any other sort of unrealistic expectations on what it can and can't do.

You read comments on HN and read these articles and you might come across feeling a sense of peril and doom which are all completely fictional for the most part. A lot can be achieved with LLMs, much more than what the constant doomers will try to drag you down to.

pgt

10 hours ago

Pre-October 2025, maybe yes. But now? Couldn't disagree more. There is no insight in this post.

zahlman

9 hours ago

The post was published today by someone who is clearly making satirical reference to frontier models ("Pol" alludes to GPT-5.6 Sol).