Notes from AI and Software Engineering in 2026

Introduction
It's hard to believe it was only a year ago that Anthropic released Claude Opus 4.5. So much has changed and it feels like we have always worked this way. I used Opus 4.5 over Christmas 2025 to catch up on adding a chat interface for one of my own apps and was blown away by the leap in capability. I knew right away we had to rush to change the platform at work to support LLMs.
I started at this current role around 2 years ago. We have a relatively small engineering org of 4 full-time engineers, so we used to be extremely tight on resources for any work that wasn't core user-facing features.
I had spent the first year applying general engineering practices to the existing platform - From adding typescript, to consolidating clouds, terraforming infrastructure, adding fully automated deploys, that kind of thing. Since Opus 4.5, AI has just dramatically accelerated our engineering transformation. From experience I know it would have taken years to get to where we are today without LLMs.
Since December 2025 our monthly commit volume increased by 4x. Over the same nine months, 500 errors across our application domains fell by around 60% as measured on Cloudflare.
Customer-created technical bug tickets fell by around 50%, and our total infrastructure costs fell by around 10%.
We still have four engineers, but measured by commits we now produce as much code as a 16-person "version" of the old team. Commits are NOT a great measure of engineering success, especially when LLMs write the code, but the increase is large enough that it demonstrates a real change in what the team can deliver due to embracing LLMs fully.

Here's my notes on some things we did that worked well and some things that are still really difficult in the age of AI engineering in 2026.
Parts of this will seem silly to those of you in large orgs where you get this stuff for "free", but other smaller teams will understand how awesome it is to be able to run and support all of this change. I'm really, really looking forward to revisiting all this in a year and see what has changed.
Note: our product has also added LLM usage to various features and has MCP etc. But this article is around internal engineering AI changes, rather than the user-facing product implementation.
Use the latest models and harnesses
We give everyone access to state of the art models. You need the best models for ambiguous tasks. You are leaving virtually free productivity on the table by nerfing your team and having them use sonnet instead of opus for example. Don't do it.
An AUD$200 plan looks expensive compared to a AUD$40 plan but it's nothing compared to an engineer losing two hours every day dealing with a crap model.
Do give everyone a budget, this is easy via one of the $100 or $200 plans from the main labs. Let each engineer figure out what works for them. Some tasks are fine with cheap models and we all make our plans last for a month with few issues.
Some people use fewer requests to expensive SOTA models whereas some devs use more frequent, faster requests to less expensive models. The process of choosing the right model is an art at the moment and it changes dynamically as the labs update model capability without changing the version(s).
Use a Monorepo
Problem: The platform had around 20 distinct app git repos and libraries. All the apps had evolved in slightly different ways over time - different node versions, different conflicting libraries, different CI and build systems, ESM vs commonJS etc etc.
I had seen that you will get better results from a lower-tier model with great context, than with crap context and a state of the art model so I figured giving the models access to all the code for each request would be best.
After the first 5-6 apps were migrated to one repo it was obvious the profound impact this had on LLMs ability to understand the whole platform and effect change across the entire stack in one prompt. We had apps that only one dev knew well, that only one dev really worked on but all those silos are gone now in the mono repo. Everyone just solves problems wherever they are in the stack.
LLM change accuracy went up, there were no more permission issues and starting up the whole platform for LLMs to verify changes became much easier. I've done changes like auth updates that required new infra via terraform, github repo changes, cloudflare changes, api and frontend changes across the entire platform and it can all happen in one PR.
If you can only do one thing with a smaller team I would recommend changing to a monorepo as soon as possible. It accelerates everything else.
The only major downside to a mono repo is managing CI time and cost. Suddenly a change to a package can fan out across 10 apps that you might not care about updating. But we made changes to fix that too which I'll cover later.
Comprehensive Observability, Logs and Telemetry
Problem: The original platform had comprehensive log streams but they were unstructured logs so they were difficult to search over. The logs went to an older platform that didn't have an api when I started planning, so there was no MCP either. There was no platform telemetry and no application tracing available on the platform.
So we moved all logs to structured Open Telemetry (otel) format and store them on a modern otel platform. All the servers and apps are now fully instrumented for telemetry. Important requests are traced through all infrastructure at various degrees of sampling.
There are some side cars to support collection on ECS and we run an otel collector inside the VPC to clean the data and for some polled stats and we self-host an observability platform. The cleaning and sanitization of auth tokens and PII is important here.
We ship all this data to S3 for cost effective storage. With significant sampling we ingest around 10TB of data in a quarter and this is compressed to roughly 200GB stored on s3 in parquet files, it's extremely cost effective compared to data dog or honeycomb.
I built a custom observability MCP server that proxies to the platform's existing API. This custom MCP server understands our environments, business and data structures to make LLM queries more efficient than using an out of the box MCP server.
This tooling has meant bugs are generally resolved within an hour or two of being reported. why did this post: <id> fail for <customerid> customer in the last 24h is enough of a prompt for any modern agent to resolve that issue now.
On top of that we have dramatically reduced bugs and increased performance across the entire platform just from having visibility. The team is adding around 2 new dashboards per week as we target new or existing features for monitoring.
Use Adversarial Code Review
LLMs still produce crap code in late 2026, even Sol and Astra. Anyone telling you otherwise is lying. So, to counter-act this we have copilot doing automatic reviews on every single PR. For trickier things claude can be triggered with a @claude review command.
My rubric is something like:
- Mechanical, reversible change: automated review may be enough.
- Product behaviour: dev verifies in local dev or a demo environment, automated review might be enough
- Infrastructure, authentication, billing or migrations: human engineer reviews and tests.
- Large architectural change: plans are shared for feedback before code is shipped
I personally also sometimes use opencode's /review with a medium-smart OpenAI model, even though I used codex to create the change. So this is more of a harness-adversarial review rather than model. It does work to find potential issues though.
So, you must still review all changes, sometimes another agent can do this, sometimes you need multiple agents and sometimes you still want another dev to look over the change.
This is one of our bottlenecks at the moment. We cannot fully automate mechanical changes like "Find and convert a styled component to tailwind" to production because 50-60% of them result in automated reviews finding issues, CI failing or they need a minor change.
Choose declarative solutions where possible
What I mean is choose terraform for infrastructure over using scripts and cli to configure your platform. LLMs absolutely love declarative coding. And it makes sense because the training data is so rich with examples.
Terraform has providers for all major cloud platforms and even some you might not expect like cloudflare or stripe. Use terraform and your AI tools will have a comprehensive understanding of your platform from infrastructure to frontend.
Using terraform will also allow you to have LLMs teach you complex infrastructure requirements and review everything to ensure you haven't made any mistakes.
Another less popular declarative coding method that LLMs absolutely love is state machines. State machines usually have a library on your coding platform (xstate for typescript/javascript).
State machines have a fixed vocabulary and a valid finite state machine ensure that all edge cases associated with implemented states and transitions must be considered or it wont compile. They are easily composed and easy to test.
Publishing child-parent data structures to third party api's can have surprising number of failure modes. State machines help us handle the worst ones. If your app has any complex flows you should ask AI to help you migrate to a state machine library.
Everyone can be a dev now
Knowledge of code syntax is not needed to make many types of changes to an existing app now. This means anyone in your org can be a developer if they want to be and if you give them the right tools. They might not make architectural decisions but UI changes and small bug fixes are well within reach.
You need to provide your agent LLM enough context so it can infer what a non-technical user might mean by their request. You also need to give the requestor the ability to test the change with say, configure a full local dev environment of a software platform.
Non-devs wont really read a PR diff, they must be able to use the application and see the change so they can verify it.
My goal was to enable anyone in the org to ask for a functional change and have an agent return an environment with that change on it. We have non-production data available for this kind of test environment so production data is not affected.
To facilitate this kind of workflow I created a slack bot that can create github issues and assign an agent to the task.
When the task is finished the same bot can be asked to create a temporary environment with that new change deployed on it. The user can check that the change is what they expect, and if it's ok they can ask a dev to triage and merge it.
These environments are protected through org SSO on top of the standard auth system we have. They are intentionally ephemeral and are killed off in just three hours. They don't run a full platform including all jobs etc, they are designed more around the kinds of things that non-devs can change and should be able to change - which is primarily user facing UI code and API contracts.
Today, technical users can ask LLMs the right questions to get good results but getting agents to create correct changes when asked more vague questions is still difficult. There are demo change environments created most days now in the org but I would like to see better agent quality for non-technical users' requests before allowing them to actually merge the changes.
Using typed languages
This might seem obvious in 2026 but some high profile devs have said that types are not needed with modern LLMs, I disagree.
Most of the code in this platform was javascript and there were hundreds of old undefined and other errors that could have been caught in static checks. Migrating almost all code to typescript gave LLMs instant feedback around correctness. This closes the loop quickly for LLMs and we haven't had any new type class of errors since.
The types and naming of types gives LLMs a huge amount of context about your business. We don't have to describe very much about how things work in READMEs because the types are so descriptive.
It is slower to compile for sure, but TypeScript 7 is quite fast and the benefits outweigh this disadvantage.
Engineering grunt work is not solved by AI
We have agents on GitHub and they have access to all of our MCPs. So for a while I tried automating some things - frontend changes, library updates, Javascript to typescript migrations, test cleanup.
I would wake up in the morning to 4-5 PRs but 50% of them needed manual changes so the promise of the software factory is over blown right no for us. It's not possible to trust AI code. This is using the latest models, with a decent harness, good MCPs.
PRs fail because they are:
- Visually broken, or not matching design system
- Technically correct but inconsistent with our linting
- Tests added and passed but fix was wrong
- Agent lacked enough product context
- Dependency update broke in runtime (ESM/CJS is common cause here)
So I turned off these automations again. I'm hoping this will become better in time with better models but ultimately I feel I will do a full sweep of the platform and upgrade whole features in one go myself, rather than try to have AI do small PRs itself each day.
Custom harnesses are extremely difficult to get right
Our slack bot scout can take questions and use the MCPs to investigate issues. It returns findings to the user.
It started with a completely custom for loop in a long running lambda. The results here were fast but not thorough.
- Lambda has execution limits
- context didnt survive past one run so the slack thread lost all previous thinking work for each new message
- no fielsystem really - github via api rather than local grep
- used cheap gemini for testing and results not great
It has since moved to CloudFlares harness with some custom instructions. I do limit the amount turns and we do use Gemini models there, the results are better but not as good as local state of the art harness and models, e.g. codex.
We don't have formal evals for the harness development but just using day to day queries I can see that Codex and Sol perform much better than cloudflare and Gemini at the moment.
We are working on the next version of this which will likely be Pi running on some kind of container with access to a full machine that can be easily torn down and replaced. We are speed running to the architecture that every other org seems to have converged on for long running agents.
Monitoring is easier
I used to manually monitor our customer inbound tickets when I got free time to look for patterns and fix priority issues before they grew into larger problems.
Now I have an agentic task that receives a webhook for each new ticket, it categorises it with a clustering algorithm and counts distinct tickets (tickets from multiple different customers). It will alert us on slack when one of these clusters is triggered so an engineer can investigate.
AI is verbose with comments and tests
Fable in particular was awful with comments. They are verbose and describe what the code does rather than why the code was written. Most comments are not needed. The code itself should be self documenting. This is an old engineering rule and it;s a bit frustrating that LLMs don't follow this old pattern.
All the LLMs seem to add useless tests. For example tests that check if a hardcoded tailwind CSS class is present are completely useless really. They make change hard because you have to update any new css in two places now. Don't allow these into your code because LLMs will copy the patterns it finds.
CI will slow you down
We run CI on GitHub Actions. Our CI is on fire. The cost has doubled every month for the last 6 months EXCEPT for the fact we keep adding optimisations to try to reduce the load.
I've added a custom turbo server instance. We cache as much as we can. We use carefully chosen instances for the task type. All the workflows have conditions around selection and running. All of the apps use vite 8 and vitest, rolldown written in rust. Everything is upgraded to TypeScript 7. We have deleted a chunk of old tests.
We have investigated using third party runners like BlackSmith but our total cost doesn't justify 10-20% potential savings just yet.
You will run out of "quick wins"
There is extreme productivity once you get the context set up. You can fix anything and everything and you do! These simple issues are easily diagnosed and fixed by LLMs. Noisy bugs in your logs that you have put off for years take just 10 minutes to PR. It's awesome.
But soon all these issues are gone and you are left with complex changes. Like your domain structures are out of date or your auth system needs to be refreshed. AI and LLMs still won't create good step by step architectural changes.
The changes I work on these days are still multi-week projects that need to be designed in diagrams and notes with requests for feedback before any code is produced.
Then I break down the job in to smaller tasks that LLMs can do (with many changes required to the first LLM output). So now we're kind of back to old school engineering again. It's just that now everyone is an "architect" and everyone on the team has an additional 6 virtual dev resources to implement their plans.
Not much has changed for these complex tasks really from a purely process perspective. I still need a good designer to help me. I still get a review from peers for tricky things. It's multiple PRs, feature flags, carefully staged releases to customers.
What's next
I think we need to have on-demand, long-lived agent environments with harnesses and models that are as good as clude and codex. We need anyone in the organisation to be able to spin these up to solve problems across code and logs and business context. Either each user gets an agent and their slack requests are sent to their agent, or each thread gets an agent instance.
Constantly monitoring costs. You want to increase your team throughput by 4x but you want to maintain costs at the old 1x. I'm not sure this is possible but any time anything spikes I go in and look for wasted compute and spend. Trying to get 4x throughput and 2x costs would be a good place to land.
Anyway, see you in 6-12 months! 🤖 🦾