kaydub
4 days ago
You don't need documentation or the 3rd party memory systems. The code IS the documentation.
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
rectang
4 days ago
Haha, all the software devs who hate writing documentation are naturally finding their preexisting beliefs reinforced when the LLM is able to discern intent without docs. An LLM can be spooky impressive at reading minimized or obfuscated code, for example.
But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.
Rejoice! The LLM will write the docs for you, relieving you of most of the work.
However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.
JohnBooty
4 days ago
Yeah. Basically IMO/IME the current best practice is to start the LLMs out with minimal/no skills/instructions/docs.
Notice the things they struggle with and the small problems (typically, environment issues IME) they repeatedly encounter and re-solve across multiple sessions. That is what your instructions should cover. When possible, move those instructions into skills, so they get loaded into context selectively instead of on every session. (Example: instructions for running specs, placed into a skill that only gets loaded into context when it’s time to run specs)
Again, this can be automated by the LLMs themselves: both Codex and Claude (and I’m assuming other major harnesses) know how to read their own transcripts and are good at looking for repeated friction and making concrete suggestions to reduce that friction in the future.
It takes a bit of a time investment on the user’s part, and every now and then you probably should throw it all out and start fresh so that the new batch of instructions can be appropriate for the current state of the repo and the capabilities of whatever model(s) you’re using.
kaydub
4 days ago
Honestly, don't agree. I think even this is a waste of time and context.
The LLMs will always have these suggestions on things that will "reduce friction in the future" but I really think the LLM is a sycophant glazing you. Because even when you add those skills or prompts at some point the LLM ignores it or does whatever it wants to do anyways. Better to leave most of that out and deal with it when it arises.
Even your example, instructions for running specs... none of the current frontier models need this at all.
joquarky
4 days ago
Are you using Astra/Fable for everything? Not everyone can afford that.
JohnBooty
4 days ago
I think they're using Skynet. They seem to be claiming that part of the reason for not giving the LLM directions is because the LLM won't follow them anyway.
That must be a remarkable model. Smart enough to flawlessly discover everything on its own in every session... but also it just straight up just doesn't obey direction.
Hmmm.
kaydub
4 days ago
Wow, what a strawman.
I didn't say give no direction.
I'm saying a lot of devs/engineers little rain dances aren't really bringing the rain.
Keep a basic CLAUDE.md/AGENTS.md that keeps high level details. Have a way to reference other projects/apps/codebases to use as reference architecture.
Don't keep tons of markdown files. Don't keep "decision" documents (these are probably the worst). Don't make skills for every little thing (superpowers are dead these days, stop using them). Oh yeah, don't have the LLM generate docs for reference later... if it can generate the docs it doesn't really need them.
kaydub
4 days ago
Nope. Opus and Sol for the most part. Haven't even used Astra or Fable.
At work it's all Anthropic. I spent a while on Sonnet models but have been pretty consistent about using Opus now. For personal projects I bounce between Sol/Terra and Gemini Flash.
Most of the models are good enough now. You don't need all these documents, memory, plugins, skills, etc. MCPs are good for hooking up to external sources of data (at work we have gitlab mcp, datadog mcp, and our knowledge base mcp... which the knowledge base is mostly worthless now because of all this AI generated documentation, but I digress). I DO still use beads on a lot of projects, but not always. My prompts these days are lazy as fuck: "review this repo, check out this other repo with architectural patterns we should follow and libraries/terraform/etc we should use, <basic desc of goal>. What do you think? Let's review everything and discuss before we start building"
Oh no, I might have to stop the llm agent on occasion, or review what they did and tell them to fix a couple things. Way better use of tokens and my time than creating some rube goldberg machine. Way better than keeping a giant decision file that goes stale and causes context corruption because the llm only saw the OLD decision and not the UPDATED decision.
JohnBooty
4 days ago
Charitably, I think we must be working on much different kinds of projects.
kaydub
4 days ago
Maybe but I doubt it. We're all working on pretty similar things, that's why the LLMs work so well.
You're just stuck on your dogma.
JohnBooty
4 days ago
I understand what you're railing against in general: engineers who make these big memory/doc systems that quickly go stale and are at best useless and at worst actively harmful. Even more insidiously, these engineers are blind to this fact because maybe some of this stuff was effective 6-12 months ago and they haven't reassessed their workflows. Now that's dogma. We really do agree on that.
However. At least in the projects I've worked on, it's trivial to read the transcripts and notice that without enough direction, there are certain things the agents will burn lots of tokens to rediscover in every session. That's absolutely not dogma.
kaydub
4 days ago
I think there's a balance between the tokens burned up rediscovering things vs tokens being used keeping docs/decisions/skills/etc.
And I think a lot of places are burning way more on the latter than the former. I think the former is more cost effective due to caching as well. So I HEAVILY lean towards the former especially with how effective I think it is these days.
JohnBooty
3 days ago
I think there's a balance between the tokens
burned up rediscovering things vs tokens being
used keeping docs/decisions/skills/etc.
This mystical balance isn't difficult to measure. You don't have to believe anybody; just read the darn session transcripts. It's going to take at least one (but usually multiple) turns of grepping or other exploration to rediscover a thing. How many tokens does this add to your context? Since 1 token = ~4 characters you can count characters in your text editor.Is this less, or greater than the number of tokens it would cost to simply add that thing to AGENTS.md?
If you're burning 5000 tokens to rediscover the same thing in every session, and it would take only 500 tokens to just put the thing in your AGENTS.md, it's an easy decision. Do you understand that paying 500 tokens in every session is less than paying 5000 tokens in every session?
If you're burning 5000 tokens to rediscover the same thing in 50% of your sessions, it's probably still worth it, but now you might consider placing it into a skill so that it only gets loaded into the 50% of sessions that need it.
If you're burning 5000 tokens to rediscover the same thing in 1% of your sessions then yeah, definitely skill. (You can add it to your prompts manually too, of course, but that is functionally equivalent to adding it to a skill that is set to manual invocation only...)
So I HEAVILY lean towards the former especially with
how effective I think it is these days.
How is this different from what I already said at the beginning of this conversation, where I recommended beginning with minimal (or no) skills or instructions... and then judiciously adding very lean ones based on actual observations? You claimed that this was a waste of time.kaydub
3 days ago
See here's the thing, you're pulling these numbers out of your ass when I'm making these recommendations from my observations that it's basically a wash and you get better results just having the agent read the code.
You aren't putting into 500 tokens what takes 5k tokens to discover. I've found that when I had these larger markdown files with all the details the models STILL went to the code to review. Then the markdown files get stale in spots, NO MATTER WHAT. Not only stale, but if you're having agents maintain these files they WILL hallucinate or input superfluous or unnecessary info. Now that I have short markdown files (or none in some instances) the agent just looks at the code and it's pretty simple to figure out.
And yes, I think most of the skills are a waste of time. You're just creating more shit to maintain. The most usable and reusable skills end up being short and concise, to where you could just type them into the prompt in any number of ways adhoc... effectively making the skill mostly worthless. The big skills end up being too much and less effective ime. And it can be difficult to rely on agents to invoke skills or the right skills.
If my skills need maintenance and hand holding, what am I really gaining?
JohnBooty
3 days ago
Pulled out of my ass? Excuse you?
As I optimistically hope you may be able to eventually understand with what I assume to be significant effort on your part[1], the numbers are going to be different on a case by case basis. Thus, my generalization.
This conversation is over, as you have been moving goalposts in various random directions: careening between "no directions for LLMs!" and then talking about your "minimal" AGENTS.md and markdown files, and then disagreeing with me for having minimal AGENTS.md and markdown files, in favor of your approach which involves... minimal AGENTS.md and markdown files, apparently. Luckily, it all sounds so schizo that at least nobody will take you seriously enough to follow your bad advice. It's exhausting talking to you, but surely not as exhausting as being you, so I do still have sympathy. Good luck.
---
[1] remember to hydrate, a good night's sleep helps too
kaydub
3 days ago
Just fyi, you're the only one that said "no direction for LLMs"
The condescending attitude isn't a great look, especially when you're making up strawmen arguments.
klodolph
4 days ago
100% agree. Most of the things that make development better for humans also make development better for agents… and I think docs are even more important with agents, because of some kind of multiplicative effect. The agents are coding faster, and the benefits of documentation are somewhat more pronounced because of the speed.
kaydub
4 days ago
You guys show me a measurable difference on something with docs vs without docs and I'll believe it.
The docs just make you feel good. They're not worth anything. They just create more work if anything because now you don't only have to fight entropy in your codebase but also your docs.
rectang
4 days ago
Uh, you aren't providing numbers either, just appending the same assertion to the end of each branch of the discussion. The result is the same as when the LLM does that: geometric increase in what we have to fight through to reason about the problem.
With the LLM, though, I have some influence: I can instruct it to be more succinct. I don't have a specific top-level instruction about that at present in any global prefs file, but I've occasionally asked it to summarize and archive and it's done an okayish job of pruning.
kaydub
4 days ago
Yes, you're right, I'm not providing numbers. But that's because I'm just using the tool as designed... the vanilla version. I'm using the baseline. I believe anthropic and openai publish plenty for my side of the argument already.
I'm not the one bolting things on claiming it increases productivity. Why would I ADD stuff without any proof or evidence that it works? I'm simply NOT adding these things.
frankacter
3 days ago
>But that's because I'm just using the tool as designed
While arguing against using the tool as designed.
https://claude.com/blog/using-claude-md-files
"CLAUDE.md files solve this by giving Claude persistent context about your project. A well-configured CLAUDE.md transforms how Claude works with your specific project. The file serves multiple purposes: providing architectural context, establishing workflows, and connecting Claude to your development tools. Each addition should solve a real problem you have encountered, not theoretical concerns about what Claude might need."
kaydub
3 days ago
I pretty clearly say I still use CLAUDE.md/AGENTS.md
I've just been trimming down to the recommended < 100 lines and I keep only super high level info in there and allow the agents to figure things out by reading the code.
I wouldn't use a blog post from Nov 2025 be my guiding light.
Anthropic over the past few months have said they've cut their system prompts 80%. These are current best practices. The models are good enough and can figure things out on their own BETTER than you can present the info in markdown files and "decision" documents.
trinsic2
4 days ago
Wow its utterly impressive that people dont understand this. I create docs and maps to instruct my LLM and im always keeping my docs up to date..
kaydub
4 days ago
What are you measuring to prove this?
The docs end up stale or you waste a ton of time and tokens keeping them up to date. And you can't even trust the LLMs to keep the docs up to date, you will still need to review it yourself keep a bunch of shit out of them.
joquarky
4 days ago
I don't understand how specifying my preferences for how things are done can go "stale".
JohnBooty
3 days ago
Architecture preferences are a great example of things that we need to tell the LLMs.
We don't need to tell the LLM where the User model is. The LLM can find it.
However, IME we need to tell the LLM preferences like, "the User model should not have dependencies on other models" (you know, to avoid the God Model problem...)
To discover that preference, the LLM would (best case scenario) need to look at every model, and notice that the User model has no dependencies, but all the other models do, and infer that this is an architecture requirement. Except, some of our other leaf models might not have dependencies anyway, so it wouldn't necessarily be apparent. Or maybe it's a new project and none of the models have dependencies in which case our preference could not be deduced. Etc. All of this is token-intensive and unreliable.
(Perhaps even more ideally than placing this instruction in AGENTS.md or similar, it could be an inline comment in the User model itself. However, I have found LLMs to be very hit-or-miss when it comes to recognizing directives inside inline comments. And some architecture rules can't be easily expressed via inline comments in a particular file...)
kaydub
3 days ago
Why would I put details about a singular class/model into the AGENTS.md?
Like you said, an inline comment probably makes more sense. Sure, maybe it misses it, then you can just tell it. Better than putting such granular details in a CLAUDE/AGENTS.md and better than having disparate markdown files or decisions documents that get unwieldy.
JohnBooty
3 days ago
Quick, non-rhetorical sanity check, since you advocate passionately for just typing the same things into every session over and over.
You know the contents of AGENTS.md/CLAUDE.md just get injected into your context window by the harness? And that it's functionally equivalent to typing the same thing into every session?
There's no difference between typing "keep dependencies out of User" at the beginning of every session, and putting that into AGENTS.md, other than wasting a bunch of time.
Sure, maybe it misses it, then you can just tell it.
Re-doing that work after it gets things wrong is going to take roughly 1-3 orders of magnitude more tokens than including it in AGENTS.md/CLAUDE.md.Here's a good concrete example. I've caught Claude writing python scripts to parse JSON instead of simply using `jq`. I initially attempted to correct this by adding `jq` to my allow list, but it still always reached for Python scripting first, before eventually figuring out it could use `jq`. You think that burned a few extra tokens vs. just telling it `use jq to parse JSON` in AGENTS.md?
kaydub
3 days ago
You sure are patronizing for someone offended by me saying they're pulling stats out of their ass.
I don't really care how claude parses json. I'd say you're just wasting your time caring about that detail. I definitely wouldn't put that kind of useless info into my claude.md/agents.md.
If I DID care, yeah, I think it makes more sense to just correct the action. I'm not parsing json in every session so why would I want it in context for every session AND every sub-agent's session?
JohnBooty
3 days ago
A lot of APIs we use return JSON - Github, Semaphore, Context7, Docker, etc. We're parsing JSON dozens of times per session. I think 5-6 tokens telling it to use `jq` in AGENTS.md was the right call for us.
I eagerly await your misunderstanding!
kaydub
4 days ago
"docs and maps" are hardly specifying preferences. I DID say I still often keep a high level AGENTS.md or CLAUDE.md. Where I have them, they're super high-level. In places where we have well defined and followed standards that top level markdown file can be as simple as "check out this other repo for patterns" and where I don't, the markdown file is high level overview of what the app is and then a high level overview of architecture or some preferences. Always a shorter/smaller doc.
I don't mind SOME documentation. I'm growing extremely frustrated with the absolute MOUNTAIN of docs, comments, decision documents, etc being created. I'm so fucking tired of the LLM rube goldberg machines being built.
hbrn
3 days ago
Refusing to commit LLM-generated docs (or tests for that matter) requires discipline that most engineers don't have:
- it feels productive and has illusion of value
- it seems harmless/free, refusing means that at best you'll be at net zero
- if you do refuse, now you actually need to think about what should be documented
So yeah, LLM-generated garbage will inevitably spread.
The good news is that teams that do have the discipline are way more competitive today.
kaydub
3 days ago
Yeah, you seem to have the same thoughts on this to me. I think it applies to more than just the documentation, a lot of the skills/plugins/etc fall under this too (like almost all of them ime). It's all prose, so it's easy to commit and ship it. It makes engineers feel like they're generating value.
I've got one example that's at the top of my mind driving me crazy right now.
We have to move some git repos. It requires making sure there are no secrets in git history or committed, in a lot of cases we're just archiving, creating new repos, copying the code over. We just have to mostly mirror the config of the prior repo. There are probably about 100 repos that we need to move. I farmed this out to 3 of my ICs, a Sr and 2 jr/mid level guys. I thought one of them would figure out a way to mass move them and worst case they just move them real quick... shouldn't be too hard.
Well my Sr has spent the past 6-8 weeks building claude skills/plugins to move the repos. Every week he finds an edge case or the tool doesn't do something perfect, so every sprint he has "fixes" for it.
So far he's moved 5 repos. The jr/mid level guys, one has moved 30 the other has moved 10. I grabbed one of these tasks, threw claude at it (again, no tools or any bullshit), gave it a quick prompt, knocked out one of these in 20 minutes (while multi-tasking).
This same IC has built an MR review tool. It takes 20-60 minutes to run and costs $20-$50+ each run. Wants to put it into our CI/CD so that it runs on MRs. If you include all the pipelines that would run because of a dev updating a branch in an MR we're talking hundreds to thousands of pipelines a day.
I thought this dev was token maxing or just trying to look busy and I've had to sit down and talk to him, pull him aside for quick 1 on 1s, and to be honest, I don't get the feeling that he knows what he's doing is a gigantic waste of time. The LLMs are sycophants that glaze you. You think everything coming out of it is GREAT work because it tells you that. And I think this is happening across all layers of every org out there in various degrees.
hbrn
2 days ago
Yeah I have similar struggles: I witnessed several multi-month projects that should never have been started. But they would sound cool on paper, and sycophantic LLMs would support any shit-on-a-stick, so they were kicked off, only to be killed months later. Huge boosts in productivity coming from coding agents are often negated by AI psychosis. But is some ways I enjoy it, because instead of being the great equalizer AI turns out to be the great differentiator.
Review tool is an interesting one - I find LLMs to be quite useful for code reviews, but people should really stop associating code reviews with Github/Gitlab and running them as a part of CI. Those platforms were built for humans because reviewing patches over email was a terrible experience. Clankers have no problem with raw patches, in fact they prefer them over fancy web UI. Code hosting platforms are not a good place for agentic code reviews, and they never will be.
Most LLM reviews should only run locally on developer's machine, in your regular coding harness.
kaydub
2 days ago
The people in here sticking to their dogma has motivated me to start building out tests for a lot of this stuff. I spent a chunk of my day today testing the MR tool today.
* vanilla claude with our MCPs + a decent prompt (like a paragraph tops) * an agent that was previously made, much more concise skills * the mr tool this IC created
The MR tool didn't even catch some of the stuff in our domain that the others caught and it took 45minutes, cost $20 vs 4mins for vanilla (<$1) and 7mins for the agent ($3ish)
And yeah, code-reviews are supposed to be by humans, but we have a ton of reviews come in where there's a ton of low hanging fruit. This tool is meant to be a gate between getting an actual human code review and some of the slop coming in these days.
I'll go through a bunch more MR reviews, see how things go. But I've gotta start figuring out how to test the documentation, decision docs, and other plugins/skills.
user
a day ago
kaydub
4 days ago
Don't agree. LLM generated docs are some of the worst because like I said, the LLM never trims, only amends. So we used to do X but now we've had an architectural change or some type of change where we should never do X. Instead of just removing the instructions to do X in the docs, it amends them, "we made the decision to no longer do X because of Y". Now it has X multiple times in context instead of just not having X in context at all.
icedchai
4 days ago
I hesitate to make blanket statements, since this is quickly evolving, but in general I agree. LLM generated docs will quickly degrade to overly verbose, unreadable crap.
kaydub
4 days ago
Not just that they get stale, but if you're using the LLM to generate the doc you don't need it.
I don't think I've ever seen an agent review a doc and then NOT also go look at the code. So skip the middleman, just have the agent look at the code.
hbrn
3 days ago
One funny observation is that while LLMs are great at coding, but they absolutely suck at generating content for themselves (or other LLMs).
They are truly horrible at generating docs, prompts, knowledge bases, memories, etc.
joquarky
4 days ago
Created a skill to prune them appropriately and run it every evening.
JohnBooty
4 days ago
In general, as LLMs get smarter, “less is more” becomes increasingly true. You’re polluting their context with dozens or hundreds of instructions, all of which the LLM tries to satisfy. However,
You don’t need documentation […or…] memory
…oh heck no. Easy to miss at Claude and Codex’s default detail level but if you read your actual session transcripts, you are almost certain to notice the LLM solving lots and lots of the same little problems over and over again. The code IS the documentation
First, lots of things cannot be learned from the code.Trivial example: LLMs struggled with QA on our app. They didn’t know how to find the seeded test accounts. They would create new ones and do it wrong, or find the seeded test accounts but not know their passwords because they were encrypted, so they’d change the passwords but not tell the other agents. Shitloads of tokens burned. It was a no-brainer to just give them the credentials in a skill that gets loaded when they do QA.
You could say that’s an environment issue, not a code issue. But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.
The days of “you are the world’s greatest Python programmer, write good code” or whatever are over (if that style of prompting even worked in the first place) and as I said less is more. But, still….
kaydub
4 days ago
> You could say that’s an environment issue, not a code issue
Yeah, that's exactly what I'd say. Or maybe you're just approaching the problem now.
You have to prompt these QA agents, correct? Why not give the instructions on where to get credentials in the original prompt?
> But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.
I did say I'll keep an AGENTS.md or CLAUDE.md. I keep that SUPER high level. The most depth I'll give is info about example projects (That the llm can reach using an MCP) to follow for architectural patterns.
JohnBooty
4 days ago
Why not give the instructions on where to get
credentials in the original prompt?
We don't push code unless it's been QA'd, and the LLMs need to know the credentials every time they do automated QA. It's very rare for me have a session in that repo where they wouldn't need that info. So it's a great candidate for AGENTS.mdAnother concrete example would be codebase conventions. We prefer lean models. Cross-model concerns go into service objects. Without this direction LLMs tend to default to stuffing too much code into the models themselves, as is the de facto standard for most MVC apps and therefore this is how LLMs are trained.
Even if they could discover our large codebase's conventions flawlessly on their own in every session, this absolutely would require nontrivial work repeated in every session: quite a few turns grepping, conversing with LSPs, or whatever.
So yes... we could describe our conventions in every single session by typing it right into the prompts... right after typing the directions to find the test credentials... and the other ten or twelve things we'd be telling it every single time...
kaydub
4 days ago
I'm assuming these QA agent are automated, probably in CI/CD. And I'm assuming the prompt that runs these QA agents is from a codebase in VC. Why would I put QA instructions into the codebase's repo as a top level markdown file and not in the pertinent part of my QA process?
Even with your documentation, the llm is gonna do a lot of that grepping and discovery. Unless you have such comprehensive documentation that it's basically code itself... in which case, it should just use the code as documentation.
The good news is that I don't think what you guys are doing is going to be really bad. It's just not near optimal and it's creating this weird dogma around all these "tools"
JohnBooty
4 days ago
Do you not run your tests locally? Regardless of what's happening or not happening in CI/CD, I would think that most workflows involve running the tests locally. We do TDD more or less and for the browser-based tests, the LLMs need to know how to log into stuff.
kaydub
4 days ago
No, these kind of tests aren't generally run locally. Not to get to prod at least. These are all deterministic tests to get to prod.
Locally? Our devs can do whatever they want. For anything I'm working on, for local deployments and testing, I generally build it out so it's set in Makefile/Taskfile. If you're running your determinstic tests locally it should really just be a single command that bootstraps everything and runs the tests, or updates a certain part of your app and runs the tests. The LLM shouldn't need credentials in that instance.
Regardless I'm deploying a local stack of some sort all the credentials and everything are probably stored locally or in an env var where the llm will have access. So if I was having the LLM drive a browser during active development, it can look at the files.
Sure, I guess we could put in the top level markdown more details about this... but why? It takes little time or context for it to figure it out. We have shit change so frequently that it's just something else we have to maintain.
hosh
4 days ago
Code as documentation works better when the code is declarative or a DSL. These capture intent and promises (as in promise theory) better.
When it is not, it has to be reasoned out and does not work well for documentation.
Other things that code and tests alone do not capture well:
- promises (as in Promise Theory) made to other parties. Claude already has PT in its training data.
- Constraints-inducing-properites, as in Roy Fielding / Christopher Alexander. While tests, and property testing can capture properties, there is no formal connection to the constraints that induces fhem. By constraints, I am not talking about business requirements and business value — those are better understood through Promise Theory. I am talking about things like at-least-once delivery or total ordering (from append-only constraint). Claude already has Fielding’s dissertation and Alexander’s works and ideas in its training data.
- grammars, as in pattern panguages (not just patterns) a la Alexander / Fielding are also not captured in code alone. These tell both humans ans AI how to extend a pattern, and how to identify anti-patterns (when they violate a constraint-inducing-property)
- LLMs are trained with many different worldviews and bounded contexts at the same time, and is very capable of translating across it. However, these need to be soelled out, otherwise it would talk in whatever it infers
Specifications written for the exact way components are wited together run into that stale doc problem. Although it takes much more human attention and token burn to describe things in terms of pattern language and promise theory, it becomes easier over time. The actual implementation plan tends to fall out more cleanly when all those other stuff are at least considered. This is where I have been spending most of my time.
dregitsky
4 days ago
> The code IS the documentation
I liked this advice when humans wrote code. Though even then I'd urge people to write meaningful commit messages that capture the "why" of what they did, so no one tramples their intent by mistake.
But not sure it works in an age where most code is LLM-generated. Especially if that code is not even reviewed by humans (irresponsible or not, it's happening), and commit messages are also generated by AI. I think something is needed to separate "what did the human operator intend" from what the agent went and built.
I do agree that this gets way overengineered. My approach has been more or less what you stopped doing though - committing all our timestamped "plan/implementation docs" and "investigation docs" that document what the user wanted + empirical findings, and making all prior session transcripts searchable. It's seemed mostly helpful? For whatever reason I haven't run into many staleness problems so far.
kaydub
4 days ago
Commit messages are GREAT.
I'm mainly aiming my frustrations at all the markdown files being committed, all the additions to knowledge bases, all the comments in the code (especially the ones referencing specific JIRA tickets). This stuff isn't helpful, it gets hella stale. I've had the LLM fuck up plenty due to these docs and comments.
Commit message ARE EXACTLY where architectural decisions or nuance should go. Not another fucking .md or more comments.
smokel
4 days ago
Code typically documents the "what" and "how", not the "why".
Why something exists, and how it connects to the outside world may be documented in comments, but more often than not it isn't.
arcanemachiner
4 days ago
And agents are pretty bad at inferring when to do, and often fall back to verbose clutterin the comments.
kaydub
4 days ago
The "why" should most often be self evident. If it's not, that's what commit messages are for. Not more markdown files and code comments.
smokel
4 days ago
Unfortunately, that is not how many software development projects work. A customer may request certain things, or the software might be part of a larger system.
Consider working on software for a coffee machine. Why is pin 42 (GRIND) activated every now and then? Would you really want to document that in Git commit messages?
kaydub
3 days ago
Yeah, I'm going too far the other way. There are definitely legitimate reasons to comment in code and have documentation. I just feel like we're seeing a MASSIVE amount of documentation now and it's so much that it's mostly worthless.
I'm seeing decision files that are so big the LLM can't fit it all in context on some projects. Then the LLM makes decisions that revert previous ones and later sessions don't pick that up so it sticks to the original decision. Now in some sessions, every so often I have to remind the LLM, "no, we changed that later, we do it this way now"
And I'm seeing our knowledge base grow to a completely useless giant mess of stale, outdated, duplicated, or superfluous info. LLMs often pull unrelated info or confuse similar but different things or get old documentation for something that's been updated to new documentation in a different part of the knowledge base. And these are LLMs generating the docs. And we have LLMs and agents reconciling. But it doesn't seem to always get everything.
For code comments, it's terrible because the comments are starting to get larger than the code. A large chunk of the comment can be discerned from the code itself. Then the comment has details on why that maybe don't quite make much sense. It's like the LLMs start using words in a specific context that doesn't really apply to the word in normal spoken english. Then it will also often include a specific JIRA ticket id, you check the JIRA ticket, you see that yeah, the code was changed because of that JIRA ticket, but it's not really related to the ticket itself, it was just a blocker. But now the comment forever links it to THAT ticket (And then now sometimes the LLM pulls in that ticket with the atlassian mcp).
haukebri
4 days ago
[flagged]
le-mark
4 days ago
It’s a weird thing isn’t it, the urge to save these artifacts? The worst is when the llm refers to the decision and design in code comments. In my opinion there’s one thing that is worth documenting; tricky architecture or implementation details that are some how counterintuitive to what would have normally been done. But again this can be documented in the code and tests.
mmcdermott
3 days ago
I've never gotten a lot of mileage out of treating tests as documentation. It is getting worse with how many people just tell the LLM to produce tests matching the code.
kaydub
4 days ago
Commit messages are also great for recording the "why" or other details that don't end up in code.
mmcnl
4 days ago
I don't understand. Code doesn't capture the context in which decisions were taken: why is code the way it is? What is important? What is not? How can agents make correct decisions without knowing context that cannot be inferred from code?
kaydub
4 days ago
Most of the "why" should be self-evident.
If it's not, that's what commit messages are for.
nomel
2 days ago
> If it's not, that's what commit messages are for.
Wouldn't this require that your commits would have to be as frequent as each new reasons why? That could easily end up being a commit per function, which seems crazy to me.
I always put any important why close to the code that I'm explaining. For architectural explanations, maybe at the top of the file, or for really high level "how it all works together", in an md.
mmcnl
3 days ago
This is nice in theory but in practice impossible.
kaydub
3 days ago
No, it's definitely not.
I'm seeing ENORMOUS decision files and when I read them a lot of the decisions are self-evident. Or the decisions flip flop but the old decisions remain in the file, which fucks up context. It's often better to completely remove a reference to X rather than leave it and have an amendment later that says "don't do X". Because X is in context, now the LLM is more likely to stumble into doing X.
perrygeo
3 days ago
I've come to the same conclusion: don't say in markdown what you meant to say in code - spend the time to make the code more clear. The code is always the source of truth, and stale docs (they all get stale) are a constant drag on the LLMs pattern matching capabilities. If you want the LLM to follow certain patterns, you have to make your code base exemplify those patterns, not write about them.
01100011
4 days ago
It's frequent for SWEs to make blanket statements with considering the vast space of issues other people face that they don't have experience with or awareness of. Anyway, I'm not going to tell you what you do or don't need, only what worked and didn't for me.
In my codebase it is difficult to get agreement on comments and documentation so rather than rely on it I adapted. One of the first things I did when I succumbed to agentic development was to point codex at the code and ask it to generate a high level description of where important files, such as our public API, reside, what the hierarchy is, what the code does, etc. In my case, this level of documentation is fairly static if I avoid implementation details. So now I have a handful of agent files in my tree and it seems to save quite a few tokens and improve my results. I frequently have other devs ask me how I get such good results when doing agentic reviews of their changes(always my first step now before I start my human review). I also include instructions in the agents files instructing the agent to maintain the agent files if any relevant changes are made. It seems to work quite well for me.
kaydub
4 days ago
This is one of the things I REALLY don't get.
If you got the LLM to generate the docs, they don't need the docs.
01100011
a day ago
Think of the docs as a cache. You're saving tokens by starting with a map of the code. Why try to figure it out each time?
kaydub
a day ago
I'm arguing that the people doing this THINK it's saving tokens. I don't think it is. I've been seeing faster and more efficient use of the LLM by removing a lot of the documentation.
spamizbad
3 days ago
I think for many projects you're right. One niche where I disagree: reverse-engineering. You'll need to carry context on things cannot be derived from code immediately eg: past experiments, findings, decomposed schematics/firmware etc. Without this, models are prone to re-doing the same work continuously - which sometimes involves "expensive" (in terms of wall time) simulations.
kaydub
3 days ago
Yeah, I'm sure there are some exceptions. And I'm not saying zero documentation, I'm just going far that way because of the sheer amount of bullshit I'm seeing get produced, both in my professional life and then online like this post was about.
So many people producing documents, comments, prose and it's just unnecessary. The LLMs don't do well keeping it up to date too. It's the snake eating its tail.
mgfist
4 days ago
Idk I just can't agree. Code is the what but it doesn't tell you the why. There are so many times where at first glance the code seems suboptimal or bad or wrong, and it's only when you learn of some constraint somewhere else that it begins to make sense.
All code is written under constraints, and most constraints live outside the code.
kaydub
4 days ago
Most often the "why" should be self-evident.
When it's not, there are commit messages.
Please for the love of god, quit generating markdown files (especially having the LLM generate the file, because if it could generate it, it doesn't need it), quit generating "decision" docs, and stop having more comments than code.
mgfist
3 days ago
> Please for the love of god, quit generating markdown files (especially having the LLM generate the file, because if it could generate it, it doesn't need it), quit generating "decision" docs, and stop having more comments than code.
I agree with all of this. But I disagree that code alone is enough to understand the why.
kaydub
2 days ago
Yeah, I'm being a little extreme here and conflating a couple things.
You absolutely need some documentation. Like I originally said, I still do keep a claude/agents.md too, they're just stripped down A LOT.
Human maintained knowledge base and documentation is good. Not even 100% human maintained, just make sure human eyes are reading it before you commit it! And if it's said in code it doesn't need to also be said in the docs, that's how they get stale and you get drift. This is causing LLMs problems.
The docs you don't need are big decision docs that keep getting appended to, multiple long lived markdown files, and verbose comments (I literally saw multiple MRs today with more lines of comments than lines of code, as bad as one file that had a single line changed but then 5 paragraphs of mostly unrelated details, overly specific details, and just plain claude slop). This stuff isn't helping.
And the one issue I'm mainly conflating this with are all the skills/plugins/etc that people are making thinking it's making their claude better. Like the super-powers repo that's been floating around for months, that thing is trash, you don't need it. Maybe you did before (I still don't really think so) but you especially don't need it now.
One guy is talking about running tests and parsing json with claude... but why would you even do that? Just make deterministic tools and then you run those for your testing.
Again, something else I'm seeing, someone wedging claude into ci/cd and then having claude use skills that run linters and trivy. Why is that not just in ci/CD???
The proliferation of prose is maddening
ChimpWithHat
4 days ago
I partially agree, way too many people are cargo culting overly complex AI workflows with little empirical data. My framing is a bit different though, I consider code the spec and actually keep a decent amount of docs for higher level concepts. So far this is working well for me across Claud and Codex.
kaydub
4 days ago
High level docs are fine and can be okay or I can see them as being helpful. I did say I'll keep an AGENTS.md/CLAUDE.md. There also are SOME docs and sometimes the LLM makes the docs anyways. I'm not an absolutist, but I'm REALLY fucking tired of seeing these giant .md files, "decision" documents, more comments than code, etc.
It's a LOT of cargo-culting overly complex AI workflows. It's devs/engineers making rube goldberg machines.
Everyone is doing their own little rain dance and when it rains they say "see, I told you it works"
chaostheory
4 days ago
Code can get much larger than the documentation that summarizes it. It also doesn't cover intent or rationale. Even if you inexplicable don't want documentation, at the very least use something like gitnexus to map out your code because relying on code alone isn't good enough
kaydub
4 days ago
That just sounds like a code-smell to me.
chaostheory
3 days ago
Vibe coding works only when your apps are small. Once they reach a certain threshold, whatever the AI was doing doesn’t scale since they only have so much context.
kaydub
3 days ago
Nah, they scale. How often are you working on a large project where you're working on the WHOLE project? The agents are good enough to get the context they need, doesn't matter the project size.
I like how so many of you get offended or go on the attack acting like I'm vibe coding some small apps. I'm working at a small/medium enterprise business with about 1B in ARR. We have 100s of repos, a couple of which are our original monolithic products. There's no issue working on any of these apps. The models have no problem using gitlab mcp to get pertinent information when they need it.
chaostheory
3 days ago
Forcing an agent to ingest more source code even if it’s just a portion of the project, needlessly eats up their context window and tokens vs giving them a decent summary of it. It results in poorer performing agents that are both slower and less capable with lower quality results at higher cost (as the context window fills up, performance declines).
Also you’re telling everyone that you don’t have proof that you actually understand what your AI built. It may not bite your company now, but it will bite them one day.
This is a prime example of bad laziness. Tbf most companies like reducing as much cost as fast as possible while ignoring the potential downsides of it, so I don’t blame anyone taking on as many projects as possible.
kaydub
3 days ago
Even WITH those docs the agent is going to review the codebase.
I've already DONE the large docs, the decision docs, the skills/plugins. I'm ripping those out now because I can see that they're garbage and they add little to zero value.
And I don't get where you think I'm telling everyone I don't have proof that I understand what my AI built. Doesn't even make sense. Because I don't have it generate tons of prose that's going to get stale? I'm okay not generating low value garbage and patting myself on the back for it.
chaostheory
2 days ago
> Even WITH those docs the agent is going to review the codebase.
But not as much of it since it has more context assuming the docs are good enough, which is my point.
> And I don't get where you think I'm telling everyone I don't have proof that I understand what my AI built.
Are you writing any of the code? Writing the tests? If not, then there isn’t any proof that you actually know the code base unless you have documentation. Even if you do know the code base one day, it’s not like you’re going to remember everything unless… you have documentation that quickly summarizes everything. Documentation isn’t just for AI. You’re likely also not going to be a company lifer so it’s also in your company’s interests for there to be documentation.
kaydub
2 days ago
> But not as much of it since it has more context assuming the docs are good enough, which is my point.
Prove it.
Somewhat related, I DID test some of this stuff today. Plain claude, telling it to use gitlab mcp to review our codebase vs a big claude skill (for mr reviews) with a lot of details encoded in the plugin, kinda like these overly verbose and other llm generated document rube goldberg machine.
Plain claude with a simple paragraph prompt + gitlab mcp outperformed the skill by quite a bit. Less false positives, one finding the skill didn't find, ran in 4min for less than $1 vs 45min and $20 for the skill (which is basically overly verbose documentation packaged up)
The llms don't need your prose or your documentation, they can figure things out more efficiently without it.
> Are you writing any of the code? Writing the tests? If not, then there isn’t any proof that you actually know the code base unless you have documentation
The code is the documentation. You can read the code. This makes zero sense.
Your two statements are contradictory here. Docs should be good enough so the llm doesn't need to read the codebase, but it should just summarize everything? You can't have both.
I even stated in my OG post that I still keep a claude/agents.md, they're just VERY SMALL. I'm railing against the decision docs, the explosion of markdown files, and the extremely verbose code comments. All the LLM generated docs are causing friction more than anything, the llms are just good enough to work through it.
chaostheory
2 days ago
> Prove it.
Anyone can make up a anecdote. In my case an agent was able to do a task in 10 min vs almost an hour without documentation. Citing an anecdote as “proof” isn’t good enough.
Here is my proof:
Hai, N. L., Nguyen, D. M., & Bui, N. D. Q. (2024). On the Impacts of Contexts on Repository-Level Code Generation. arXiv. https://doi.org/10.48550/arxiv.2406.11927 Cited by: 46
Jimenez, C. E., Yang, J., Wettig, A., et al. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv. https://doi.org/10.48550/arxiv.2310.06770 Cited by: 4998
Zhang, F., Chen, B., Zhang, Y., et al. (2023). RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. https://doi.org/10.18653/v1/2023.emnlp-main.151 Cited by: 698
> The code is the documentation.
No because it doesn’t capture the intent or the rationale why something was designed a certain way. Like why does it do some weird logic in this spot? What and why is being tested (a lot of tests are not straightforward)?
> Your two statements are contradictory here. Docs should be good enough so the llm doesn't need to read the codebase, but it should just summarize everything? You can't have both.
They are not contradictory. There’s this thing called a table of contents that summarizes large sections into a sentence or less, and links to said sections. Each section both large and small can also have a little paragraph that serves as a summary for the content to come. Sections can even have their own mini TOC. This enables an AI to be able to skip the parts it doesn’t need.
> All the LLM generated docs are causing friction
Ok, if you have a LLM generate your docs and you don’t vet it and make it your own, it doesn’t help you. Because again, how can you prove that you actually know the code base? Unless all you do is work on the same part of the same application day in and day out? Maybe the programs you’re working on is small or simple? How can you ever quickly ramp up when changing projects? This is assuming that you’re not just pairing vibe coding with manual monkey testing, or are you even validating anything when you’re not recording your goals in writing? Most of the cost in application development lies in maintenance and that is also where documentation benefits the most
kaydub
a day ago
I'm not asking for anecdotes. I'm asking for actual real proof. I'm not the one making extraordinary claims about how my rube goldberg machine of documents and hacks works. I'm just using the tool as designed.
You're really going to use papers from 2023 and 2024? It's 2026. The models are better these days... a lot better. Like come on man "The best-performing model, Claude 2" You're going to base your assessment on THAT?
Dude, maybe it's YOU guys working on small codebases that are worrying about "knowing the codebase" or maybe you guys just work on shit code, I don't know. This industry isn't new, there are known patterns and the models know them. It's really not hard to navigate a new codebase if you know what you're doing or if you know the framework or patterns of the language/platform. If you think you know your whole codebase though, I'd argue that it's YOU working on a small or simple app.
I'm not even understanding your line of questioning here. I don't know the codebase because I had an LLM write it and I didn't write documentation? You think writing documentation means you know the codebase? You think having your goals in writing is any different than putting it in the prompt?
chaostheory
14 hours ago
> I'm not asking for anecdotes. I'm asking for actual real proof.
I gave you real proof: real research in this field on the exact subject matter WITH data.
> You're really going to use papers from 2023 and 2024?
Vs what? Nothing? Ok, prove them wrong. anecdotes aren't proof
> I'm not even understanding your line of questioning here.
I've disproven your claim 3 times. You're not understanding anything because it goes against your narrative. If you didn’t write the code and you didn’t write any documentation about the code, then you have no proof that you know anything about said code base.
> Dude, maybe it's YOU guys working on small codebases
You're projecting. You can't vibe code large projects. It's not sustainable. if you are actually vibe coding actual large projects blind, it's only a matter of time before things go really wrong.
nomel
2 days ago
> The agents are good enough to get the context they need, doesn't matter the project size.
How big are your code bases that result in this opinion? I'm working on a few million lines of code, and the AI will very very often reinvent/reimplement whole libraries within the codebase.
soltanov
4 days ago
Versioned documentation is inspectable, reviewable, and easier to correct than opaque recalled snippets. Documentation should describe current truth; append-only events can preserve history.
ramesh31
4 days ago
Yup. Examples examples examples. All of the descriptive stuff is just nonsense that confuses the point. Makes perfect sense when you remember that these things are not intelligent, but truly just autocomplete on steroids.
enraged_camel
4 days ago
>> You don't need documentation or the 3rd party memory systems. The code IS the documentation.
We have heard this nonsense from the "we don't need to write comments, code should be self-documenting" types for decades. It was wrong in that context, and it is wrong in this one.
Code tells you how a system works. It does not tell you why it works that way. That is what memory is for. It exists so that your AI does not keep undoing past decisions when it writes or refactors code.
kaydub
4 days ago
I've had the LLMs undo past decisions MORE from the docs than from having no docs.
Old decisions always end up in the docs. If the LLM gets a whiff of an old decision, but doesn't get the update to that decision, well now you're doing things the old way again.
AppleBananaPie
4 days ago
I went through the same cycle as well.
I think it's going to be an incredibly common, maybe universal cycle people will go through working with AI until they realize it doesn't work long term.
jjfoooo4
4 days ago
What about big projects, where much of the code is not written yet?
kaydub
4 days ago
I did say a high-level CLAUDE/AGENTS.md is fine.
Make a high level .md, let it rip, iterate. You can give specifics in your original prompt.
I don't need to document the framework or the libraries etc. Pre-code, I tell it in the prompt one time. After it inits the project it's in the code.
mike-akdeniz
4 days ago
[flagged]
haukebri
4 days ago
[flagged]
shinokami
4 days ago
[flagged]