Anyone else develop a system for open horizon tasks with minimum supervision

1 pointsposted 4 hours ago
by K0balt

Item id: 49921230

3 Comments

K0balt

4 hours ago

In working my project I’ve kinda accidentally created a development “harness? It’s not even a program but rather just a set of structures and tools) that has become sufficient for open ended horizon software development, at least in my vertical.

It requires a minimal amount of steering and guidance and produces structurally sound, well documented, correct code, with less bugs and failures than we used to have with an all-human team.

The code has better test coverage, and is also easier to read and reason about in general (C++) than what we used to accomplish. The hardware testing is done automatically on the bench with minimal human intervention, and our bench test coverage is much, much higher than it ever was before, so that’s also good.

The failure mode seems to be just never finishing, with the local Overton window shifting to smaller and smaller issues until it’s writing bug reports and fixes about phrasing in the comments and formatting choices that are more aesthetic than functional, or fixing hypothetical bugs that would never be reachable without significant architectural changes.

Anyone else with similar experiences in your work?

favurdev

3 hours ago

You need to define an escape hatch for the LLM otherwise it will loop on increasingly small issues and even start making up issues that don't actually exist. A simple "fix all medium and high severity issues" guidance will prevent low-priority nits from consuming all your tokens.

K0balt

2 hours ago

yeah, im seeing that lol.

Im hoping i can still operate this way with Qwen3.8-Flash-Next, which we can run on-prem with existing hardware, but im pretty sure ill still have to keep some of the expert agents from frontier models, at least for now.