brink
10 hours ago
I also have found that AI has not lived up to many of its promises and have dialed back what AI gets control of. My projects were turning into unmaintainable messes. The people who say coding is solved aren't paying attention.
lucianmarin
10 hours ago
Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.
dnautics
9 hours ago
I have the opposite experience. I don't have the patience sometimes to clean up my code and stick to coherent conventions and organization even though the will is there. With my AI projects I watch it and if the AI starts drifting I ask it to go through and look for convention/directory structure violations and it happily cleans everything up in about 10 or so minutes.
Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance
fasterik
9 hours ago
I've been a vibe-coding skeptic for years, but because of the math breakthroughs of the past few weeks I decided to experiment with the latest models on some test projects. They're a lot more capable than I thought they would be. I agree that it's easy to create an unrecoverable mess, especially when you're one-shotting a lot of features without detailed instructions. But I find that as long as I'm strict about the API boundaries and force the agent to work in small chunks, it's pretty effective. As one example, I got it to write an SVG renderer in a few hours (not the whole spec, but most of the path features and text rendering), which would have taken me at least a week just for the coding part, plus extra time to learn the algorithms.
bitexploder
8 hours ago
It is also easier than ever to build specs and have nice easy to maintain projects. It just doesn't happen magically via few shot prompts :)
globular-toast
8 hours ago
You could have also copied an SVG renderer that implements the whole spec from whatever open source project the model copied it from.
fasterik
7 hours ago
It didn't copy any source code from any external projects. I had it write a stratified sampling renderer for ground truth, then had it implement feature by feature by matching the pixels. Unless you mean it "copied" it in the sense of third-party code being part of the training data. I don't think that definition of "copy" makes any sense given how these models represent embeddings. It would also imply that humans are "copying" the things they've learned from.
fignews
9 hours ago
Have you considered that maybe this is a reflection of your skills rather than that of the LLM?
catlifeonmars
9 hours ago
It could be that small variations in prompting lead to large differences in quality of output, especially over longer horizons.
I’m saying it’s probably multiple factors and both you and GP are right.
sirsinsalot
9 hours ago
Have you considered it isn't?
Save your "you're holding it wrong" if you're not going to suggest how to hold it.
Cult speak escape hatches are intellectually lazy.
fignews
9 hours ago
Sure, happy to provide you with an example of how to hold it (turns out Steve was right) =D
https://github.com/NousResearch/hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn't load for me presumably because it can't handle this scale of commits. But I estimate ~1K commits per day on average.
There's a blog entry https://nousresearch.com/refactoring-hermes-with-1393-agents that details some work that was done by LLMs to refactor and improve the code.
I guess they know how to hold it?
diek
5 hours ago
I find it funny that the original complaint was: "AI made a mess of the codebase".
Your response was: "Well you're not doing it right, but these hermes devs know what they're doing".
But the blog post you linked to shows their prompt, which is:
> I want god files broken up. I want simplification across the board. I want unification of helpers and methods that can be reused. I want less if-if-if-if-if-if-else routing. I want code legibility up. I want interpretability of the codebase and how things connect to each other up.
So it sounds like AI made their code a mess too. They then tried to make the point of how much money they saved cleaning up the code with AI, that AI made a mess of to begin with.
And if you look at the merged PRs on that project, a ton of them are bug fixes... to the code the AI wrote. And that's been my personal experience too: AI creates a huge amount of churn in a codebase. Just vast amounts of PRs fixing code that the AI itself wrote.
chmod775
8 hours ago
You're using quantity metrics to answer a quality question.
I had a look at the kind of issues that are reported at that project (there's 15k of them, so I can at best assess a couple). It looks like a complete mess: A lot of concurrency and resource mismanagement issues and edge cases that in a better-managed project would have been avoided by construction. They will now will likely be solved by more defensive programming, driving overall complexity ever upwards.
If you really want to check some quantity metrics to try to reason about code quality, look at whether "fix" PRs are overall LOC neutral or negative (not counting tests). In this project, almost every "fix" is an addition. Worse, almost every fix is more branching.
If almost every PR is some sort of fix, and most of them add branching, and there's thousands of them weekly... That leads to only one place and I want to be nowhere near it.
nvme0n1p1
7 hours ago
Agreed. As they say, quantity has a quality all of its own.
Show me an AI that adds features by deleting code (https://www.folklore.org/Negative_2000_Lines_Of_Code.html) and I'll pay attention.
williamcotton
9 hours ago
I’ve had the best luck by spending quite a bit of time going over the big picture architecture up front and then diving into the modules to further refine the details, making sure to generate step-by-step chunks of work in Markdown format for implementation. I’ll spend literally a couple of days doing this before starting any coding.
Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!
What was your process?
hirvi74
5 hours ago
Do you mind sharing any code from what you have produced? People talk about LLM successes and failures, but what's there to really talk about when the code can speak for itself?
In case it is unclear, I am genuinely curious. I have great success with chatbots, but vibing coding has never gotten me further than a proof-of-concept.
williamcotton
4 hours ago
Here’s an example:
https://williamcotton.github.io/datafarm-studio
Some demos of the above charting language:
https://williamcotton.github.io/algraf/demos
WASM, in browser editor, LSP, and more.
runarberg
9 hours ago
Meaning OP is a better programmer then a statistical model which produces the most probable results?
Keyframe
9 hours ago
I'm reading y'all comments and it seems we're still finding our footing, and will for some time. I have the luck that I have access to pretty much all frontier and other models alike. I've been extremely negatively biased towards any LLM use in the start. Then the influx from juniors came, then the revolt of doing PRs on such slop came, then some structured methods how to do LLM work came, then vibe projects came, etc, etc. I literally have all the described experiences you've guys mentioned. From good, to bad, to ugly. It's like there's no one particular way about doing this, and no two projects share the same approach - just like ye olde times.
satvikpendem
9 hours ago
When and what model?
joshheitzman
9 hours ago
People, especially those who don't do software development, often conflate coding and software development.
The models aren't good at architecture and design. But they take direction on architecture and design and design well and can refactor code quite effectively. AI agents can absolutely be used to clean up vibe coded code bases once you figure out if the investment is worth it. The mess can be avoided if you give them sufficient guidance on architecture and design upfront.
That said, doing so purely in text form doesn't feel great right now. I've been thinking about UML lately. The problem with that was the roundtrip after the code was generated and then the implenetation happened. I don't necessarily think UML is the solution, but neither is walls of dense text.
PorciiVorbesc
8 hours ago
>The models aren't good at architecture and design.
Can you elaborate to back up this claim? WHat exactly is your yardstick for "being good at SW design and architecture"?
Because I found the current SOTA AI models being great at architecture and design, much better in fact than most average real-world devs. Is your yardstick just the John Carmacks of the world by any chance? Because most devs are not John Carmack. They are also not Linus Torvalds, they are not Stallmann, etc.
Maybe your LLM experience is still stuck in the 2023 era of ChatGPT?
joshheitzman
6 hours ago
My yardstick is me: https://www.linkedin.com/in/joshheitzman
PorciiVorbesc
6 hours ago
Cool. Can you elaborate at which types of task you are better than SOTA LLMs in context of "being good at SW design and architecture"? Got any examples? Is it at the interview questions? Or real world problems? If so what is the scale of the real world SW design and architecture problems you're better than the SOTA LLMs? Is it FAANG scale or mom and pop shop scale?
And do you consider yourself to be representative of the average developer, above them, or below them?
LLMs don't even need to be better than the average dev, let alone the top performing ones, like you. If they can just be better than the bottom 20% of devs and white collar workers in general(easily achievable when you've been around the block and saw how many useless people just keep warm chairs for high wages in large companies), that's already a huge win for those products.
What I mean, at a previous job I had ran into a memory leak issue in our backend and discovered a colleague pushed a library into prod which came with comments in the source code saying "DO NOT USE IN PROD, IT CAUSES A MEMORY LEAK!". There's cases where human stupidity and carelessness far surpasses whatever issues LLMs cause so maybe the average dev isn't really that much better than the SOTA LLMs.
joshheitzman
5 hours ago
I'm currently driving my AI coding agent harness through the process of refactoring itself. This is the last big refactor I need done before I can polish and release it as OSS later this month. So its not a large scale project I'm working on at present. The general problem is that mostly add new code and try to minimize editing existing code, which effectively means the code base is grown organically rather than being intentionally architected and designed. A few of specifics:
1) duplication - LLMs are great at generating lots of text, so its faster and easier for them to generate entirely new facilities that overlap heavily with existing ones then it is for them find existing facilities that should be expanded and refactored (note I just said 'find'; actually editing raises the time and difficulty even more). This is fine for a while as the duplicate facilities usually work just fine, up until something needs to be changed across all of them and they miss changing one or more of them, things break, and a bunch of tokens have to be burned tracking down the issue.
2) ever increasing surface area - even when making changes that do expand a facility without much duplication they frequently only add without removing much of anything or changing the overall design of the facility to reduce the amount of state its tracking and the number of branches it has based on that state. I've never seen one decide to split up something large or with too many responsibilities on their own. They will happily create a god class or function and just keep making it bigger.
mulemisterX
9 hours ago
I've had the opposite experience. Spec driven development really makes things much easier to maintain. If you're just yolo'ing it and throwing prompts around things will fall apart fast.
Dfiesl
9 hours ago
Were you reading through the code it was generating to make sure the flow was intuitive and comprehensible for each PR? Projects only turn in to maintainable messes if you blindly merge in unmaintainable messy code.
satvikpendem
9 hours ago
As usual, when people post comments like this they never specify what model and when they used it. There is a huge difference between GPT 3 and 6 for example.
Late 2025 also had a step change when agents could largely code autonomously without handholding like previously, and to be honest it's not worth hearing opinions about AI from before that time, that's how significant the change was.
analog_daddy
9 hours ago
Oh my god. I am sad to have to acknowledge this even if the models have “theoretically” gotten better. I am not a programmer but I am super opinionated with the design and structure of the code and need to make sure whatever I write/ask to generate and use for my own tasks is understood by me at least once. I wasn’t this way in 2023. I am constantly reading/trying for an agent to help me write simple code without spending too much time. But so far nothing has worked. I wish someone figures it out otherwise, I personally would not be able to realize the AI agent productivity benefits that others are supposedly seeing.
zanderwohl
9 hours ago
If you specify a structure, pattern, or design, it will stick to that design after a code review phase. Just tell it what standards you have, and after two rounds, it will have written what you described.
vjvjvjvjghv
8 hours ago
I treat AI like an intern or junior team member. With enough guidance, they can contribute a lot but you can't let them loose without supervision or they usually will produce a big mess. As of now, you are still responsible for overall architecture. One strong indicator that something is going wrong are big pull requests where the AI has rewritten large sections of the code.
xenadu02
8 hours ago
AI is a useful tool but it absolutely needs a lot of human guidance.
Unlike a compiler it won't give up at the first sign of trouble but that just means it left alone it will dig bigger and bigger holes.
Treat prompt engineering as a discipline and refine your technique. When it produces garbage throw out the work and start over until you figure it out.
nicoburns
9 hours ago
Were you reviewing the changes? I've seen lots of projects with this problem. But if AI PRs have to pass the same review bar as any other PR then it shouldn't be an issue in theory (if you can actually maintain the discipline).
thesdev
9 hours ago
Isn't it tiring to keep up with considering the speed of the generated code? And if you want to be careful with the review you lose a significant part of the speed advantage.
satvikpendem
9 hours ago
It's the same as any other PR. Keep the changes small and contained to that feature. AI can do this if you instruct it to, there is no need to vibe code some 40k line monstrosity.
drewstiff
7 hours ago
Assuming we are measuring time and
`total = dev + review`
If dev approaches zero, but you review at the same pace as you always have, are you in a better position? Yes.
Will you potentially have a backlog of code waiting for review? Also yes.
Would you prefer to be waiting for the dev team for all of the time instead, then still have the same amount of reviewing to do at the end of it? Absolutely not.
nicoburns
5 hours ago
Yes, but it's the only option if you want to retain a maintainable code base. And you can always slow down. You'll probably still be a bit faster than you were before.
classified
9 hours ago
Having unmaintainable code faster is only an advantage if it's a one-shot throwaway artifact.
TuxSH
8 hours ago
Vibe coding has always had and always will have one fundamental flaw: you didn't write the code it produced, therefore you don't have a good understanding/mental model of it.
On the other hand asking these clankers "review the feature branch I wrote" and "review my entire codebase for bugs" or "help me debug this" has saved me months of prospective work.
And more recently most major models have been getting _really_ good at RE, for example you can have OAI models (and maybe A/'s if they don't refuse) use idalib MCP and reverse-engineer stuff from start to finish, then follow up with GLM 5.3 for vuln assessment and exploit PoC.
Stuff that used to take weeks or months now just takes a few hours, or less.
cortesoft
8 hours ago
I am not saying you are right or wrong, but I don’t think you can make such a broad conclusion simply because of your own experience.
You say you used AI and your projects turned into unmaintainable messes, so your conclusion is that it means AI is not living up to its promises.
I guess if the argument is “AI makes it so you always get a great result no matter how you use it”, then your argument is sound. Your projects not working out proves that AI doesn’t always work no matter what.
However, that doesn’t mean you can’t use AI to create sustainable and well organized code. Failing to do something doesn’t mean it’s impossible and anyone who thinks they can is not paying attention.
I can’t run a marathon. If I went out and tried to run one, I would get a few miles and collapse, failing completely.
I don’t think it would be reasonable, though, at that point to say “running a marathon is impossible, anyone who says they can do it clearly lying. I tried and didn’t even make it 5 miles!”
I wish people would stop assuming their experience with something is the only possible truth.
threethirtytwo
8 hours ago
Stop being overly polite to these people. It is seriously coming down to the point of utter delusion.
As ai becomes better these people will begin changing their story because it’s utterly obvious what’s happening.
bitfilped
9 hours ago
The people saying programming is solved couldn't program in the first place is what I've noticed. So to them it really does feel solved, things "work" and they don't have to learn what they don't know about programming.
orangecat
9 hours ago
I've been programming professionally for decades. LLMs are extremely useful. At this point if you haven't figured out how to get value out of them, you're either holding them very wrong or you're being willfully ignorant.
mod50ack
9 hours ago
I think it depends which LLM tool you're using. If you're using an older, worse model (the kind that you can use for free), the experience is significantly more frustrating. On the other hand, I'd say that the current best models are very useful with a skilled operator.
drewstiff
7 hours ago
I would argue that using an older model and assuming that it is the cutting edge is basically equivalent to holding it wrong
Keyframe
9 hours ago
You both can be right at the same time.
bigstrat2003
9 hours ago
There are a couple of notable counterexamples here (nobody sane thinks Carmack doesn't know how to program, for example), but by and large I agree with your observation. The people excited about programming with LLMs are, on average, people who weren't good at programming to begin with. Still, given that these counterexamples do exist I try to avoid painting with an overly broad brush for the sake of nuance.