“This is mental”
To find out more about the reality of all the (generally overblown) promises about LLMs, I have been doing a very serious experiment on the poster child for Generative AI: coding. I wrote one critical post, and I am close to writing up my conclusions, but I am hitting some work that I have to do first before I can publish that (e.g. create a demo video which is a lot of work, and before that I need to finish some final functionalities and improvements.
This is Lab Note #3. Some basic explanations about my setup are in Lab Note #1.
Context: As an experiment to learn about the limits of LLM-coding, I am building a professional visual modelling application from the ground up. It can handle multiple grammars (definable in JSON). It has state of the art graphics. It is rich in what users can influence and how (e.g. schemes, style/colour/geometry properties). It is responsive. It is well documented. It has a solid software architecture. Every line of code (and many of the other texts) have been generated by the LLM (so far almost exclusively Claude Code). It runs on macOS and Windows (Linux will be added later). If I finish it and I am satisfied, it will be open sourced. This is a serious sized project: 140k lines C++, about 15k lines stable documentation and 30k lines of temporary documentation (analysis, design and planning for substantial changes over the last half year), so about one third of what is in play is not code but documentation.
Warning: That I am doing this should not be read as either pro or contra coding with LLM-agents. I am highly critical of the hype, but I am trying to establish the limits of LLM-coding based on as serious a use case as possible. I can say already that my experience so far has let me conclude that LLM-agents will in no way be able to replace actual human engineers.
A few days back, I had Claude Code (Opus 5.5, xhigh) create an analysis of work that Codex (GPT-6 Astra) had done:
> Our current branch holds autolayout work and a complete refactor of undo/redo based on bug findings in the original implementation. I want you to do a deep analysis, as deep as you can, on the work done. Focus on code quality, architecture quality, robustness, and performance. Write the result in docs/20260927-review.md
Claude Code dutifully fulfilled the task. and the result shocked me in parts (see previous Lab Note #2). The story, however, did not end there. I took that review back to Codex and let add its own analysis of the analysis.
> Now read docs/20260927-review.md. This is a review of the work you did on autolayout and undo/redo. I like you to add your view on the observations and conclusions
This refined the review, strengthened some observations from Claude, but weakened others. So, obviously I returned to Claude Code again and let it react to Codex’s assessment. I went back and forth a few time to see how close to agreement the analysis-generations would get. I am including the entire document (as PDF so you do not have to scroll past more than 1000 lines below)
You don’t have to understand the actual statements in this text about this ever more complicated software, but what you can get from skimming it is that a few of these rounds may actually withdraw some of what has been stated with high confidence in the first place.
So, another example of the dangers of trusting LLM-generated output blindly; it is something one has to do with care. I tend to have been the one doing this as ‘human in the loop’ (and I still was in this process of letting them react to each other)
Observations and Conclusions
The first observation of course is that this is helpful and that it improves the quality of the coding-with-LLMs process. You even don’t need multiple models for that, if you separate the sessions, you can have Claude Code critically assess (even destroy, I have seen it happen) the code Claude Code generated in the first place.
The second is of course that you are running a statistical process and thus results vary. Not ‘may vary’ but ‘will vary’, because that is the nature of the built-in randomness at the heart of LLMs. The variation may be constrained, but the core unpredictability remains. The review itself is also fundamentally unreliable because it is constrained random statistics. So it needs lots of extra statistics to converge. And even then, I have had situations where what seemed to be an obvious good choice, turned out to have been based on the non-understanding by an LLM of that code. It is one of the reason my harness now adds documentation that is about one third of the size of the code (not counting the self-documentation of the code I let generate). Without it the developments are all over the place.
The third is that this is going to be costly when used in larger code bases. In this case it takes me a lot of time, waiting for results reading and judging, etc. Also: especially Codex feels throttled and effort-limited — OpenAI makes a more ‘in dire straits’ impression than Anthropic, currently — so I have decided to use my current limits and go back to Claude only, also because of what came next when I went for planning and implementation, each model separately to see how they would differ in implementing the same (next Lab Note, probably, the result was very interesting).
Even when you would give the agents more autonomous leeway and you set up many agents with different tasks (code generation, review, etc.) and let them independently, automatically converge on a result, such convergences can take a lot of effort. In this case, I was the one decided that enough was enough. But when statistical token analysis will converge is anybody’s guess and if you add the convergence as a goal, the convergence becomes a ‘let me please the user’ aspect, i.e. it can converge but you don’t know why. The success of these systems (things like the Hugging Face incident included) depends on an enormous amount of visible and invisible trial and error and as Lab Note 1 noted, the models can get off course/topic, and there is no substitute for a human to prevent that from happening.
Mental observations
Running these architecture and code analyses via token generation is a costly mental detour. As an experienced software engineer knowledgeable of the details in the code base, much of this would happen efficiently in my head and turn into some efficient compact understandings that are easy to apply. It has value that it is written down (especially because humans tend to need constant engagement to keep up to snuff, go away from a project for 6 months and you tend to have to learn a lot all over again), but it costs energy to work through these texts and understanding and discussing them on a daily basis while change is ongoing, and the (personal) energy cost rises as the volume for me to digest increases. After all, you need to write this stuff closely because a wrong phrase can have very negative effects.
More about the human mental side (and the effects of ‘cognitive offloading’) in a later note.
First post: Anthropic/OpenAI may be spending more than $1000 for every $100 you pay them
Lab Note #1: Serious coding with LLMs. Lab notes 2026-09-25: Codex vs Claude Code, Usage Limit Resets (contains basic info on the project)
Lab Note #2: Serious coding with LLMs. Lab notes 2026-09-25: Codex vs Claude Code, Usage Limit Resets
Subscribe? I don’t know if I am going to publish the appearance of new notes every time everywhere (it is a drag), so best to keep up to date with my observations and analysis on the Information Revolution is to subscribe (top right). You’ll get an email whenever I publish. These Lab Note are on a pretty high frequency schedule so far, but normally I haven’t been doing more than 1-2 a month.
(I also like to be read. Share?)
[You do not have my permission to use any content on this site for training a Generative AI (or any comparable use), unless you can guarantee your system never misrepresents my content and provides a proper reference (URL) to the original in its output. If you want to use it in any other way, you need my explicit permission]