Darragh ORiordan

  • About
  • Articles
  • Projects
  • Hire

Socials & Contact

  • Follow on Twitter
  • Follow on GitHub
  • Follow on LinkedIn
  • mailto:disco@darraghoriordan.com

Explore

  • About
  • Articles
  • Projects
  • Hire
  • Privacy Policy

© 2026 Darragh ORiordan. All rights reserved.

What worked and didn't work with AI-Software Engineering in 2026

  • #engineering
Photo by Martin Adams on Unsplash
Published September 27, 2026

Introduction

It's hard to believe it was a year ago that Anthropic released Claude Opus 4.5. So much has changed, it feels like we always worked this way. I used Opus 4.5 over Christmas 2025 to catch up on adding a chat interface for one of my own apps and was blown away by the leap in capability. I knew right away we had to rush to change the platform at work to support LLMs.

Holiday funtime

I started my current role around 2 years ago. We have a relatively small engineering org of 4 full-time engineers, so we used to be extremely tight on resources for anything that wasn't core user-facing features.

I had spent a lot of time applying general engineering practices to the existing platform - From adding typescript, to consolidating clouds, terraforming infrastructure, adding fully automated deploys, that kind of thing. It would have taken years to get to where we are today without LLMs. Since Opus 4.5 AI has just dramatically accelerated our engineering transformation.

Since December 2025 we increased engineering throughput by 4x, reduced errors by 50-60% and reduce infra costs by 30% in 9 months. We now operate as if we had the equivalent of 16 engineers so there's less resource bottleneck but a whole heap of new, interesting problems.

Org commits

Here's my notes on some things we did that worked well and some things that are still really difficult in the age of AI engineering in 2026. Some of this will seem silly to you in large orgs that have these things already, but other smaller teams will understand how awesome it is to be able to run and support all this change. I'm looking forward to revisiting all this in a year and see what has changed.

Use the latest models and harnesses

We give everyone access to state of the art models. You are leaving virtually free productivity on the table by nerfing your team and having them use sonnet instead of opus for example. Don't do it.

Do give everyone a budget, this is easy via one of the $100 or $200 plans from the main labs. Let them figure out what works for them. Some people use fewer requests to expensive SOTA models whereas some devs use more frequent, faster requests to less expensive models. Whatever works for them.

Use a Monorepo

Problem: The platform had around 20 distinct app git repos and libraries. All the apps had evolved in slightly different ways over time - different node versions, different conflicting libraries, different CI and build systems, ESM vs commonJS etc etc.

I had seen that you will get better results from a lower-tier model with great context, than with crap context and a state of the art model so I figured giving the models access to all the code for each request would be best.

After the first 5-6 apps were migrated to one repo it was obvious the profound impact this had on LLMs ability to understand the whole platform and effect change across the entire stack in one prompt. We had apps that only one dev knew well, that only one dev really worked on but all those silos are gone now in the mono repo. Everyone just solves problems wherever they are in the stack.

LLM change accuracy went up, there were no more permission issues and starting up the whole platform for LLMs to verify changes became much easier.

If you can only do one thing with a smaller team I would recommend changing to a monorepo as soon as possible. It accelerates everything else.

The only major downside to a mono repo is managing CI time and cost. Suddenly a change to a package can fan out across 10 apps that you might not care about updating. But we made changes to fix that too which I'll cover later.

Comprehensive Observability, Logs and Telemetry

Problem: The original platform had comprehensive log streams but they were unstructured logs so they were difficult to search over. The logs went to an older platform that didn't have an api when I started planning, so there was no MCP either. There was no platform telemetry and no application tracing available on the platform.

So we moved all logs and telemetry to open telemetry. All the servers and apps are fully instrumented for telemetry and structured logging. Important requests are traced at various degrees of sampling.

Observe

There are some side cars to support collection on ECS and we run an otel collector inside the VPC to clean the data and for some polled stats and we self-host an observability platform. The cleaning and sanitization of auth tokens and PII is important here.

We ship all this data to S3 for cost effective storage. With significant sampling we ingest around 10TB in a quarter and this is 200GB stored, it's extremely cost effective compared to data dog or honeycomb.

I built a custom observability MCP server that proxies to the platform's existing API. This custom MCP server understands our environments, business and data structures to make LLM queries more efficient than using an out of the box MCP server.

This tooling has meant bugs are generally resolved within an hour or two of being reported "why is x broken for y customer in the last 24h" is enough of a prompt for any modern agent to resolve most issues now.

On top of that we have dramatically reduced bugs and increased performance across the entire platform just from having visibility. The team is adding around 2 new dashboards per week as we target new or existing features for monitoring.

Use Adversarial Code Review

LLMs still produce crap code in late 2026, even Sol and Astra. Anyone telling you otherwise is lying. So to counter-act this we have copilot doing automatic reviews on every PR. For trickier things claude can be triggered with a @claude review command.

I personally also sometimes use opencode's /review with a medium-smart model, even though I used codex to create the change.

So, you must still review all changes, sometimes another agent can do this, sometimes you need multiple agents and sometimes you still want another dev to look over the change.

This is one of our bottlenecks at the moment. We cannot automate mechanical changes like "Find and convert a styled component to tailwind" because 50-60% of them need a minor change or human review.

Choose declarative solutions where possible

What I mean is choose terraform for infrastructure over using scripts and cli to configure your platform. LLMs absolutely love declarative coding. And it makes sense because the training data is so rich with examples.

Terraform has providers for all major cloud platforms and even some you might not expect like cloudflare or stripe. Use terraform and your AI tools will have a complete understanding of your platform from infrastructure to frontend.

Using terraform will allow you to have LLMs teach you complex infrastructure requirements and review everything to ensure you haven't made any mistakes.

Another less popular declarative coding method that LLMs absolutely love is state machines. State machines usually have a library on your coding platform (xstate for typescript/javascript).

State machines have a fixed vocabulary and a valid finite state machine ensure that all edge cases must be considered or it wont compile. They are easily composed and easy to test. If your app has any complex flows you should ask AI to help you migrate to a state machine library.

Everyone can be a dev now

Knowledge of code syntax is not needed to make many types of changes to an existing app now. This means anyone in your org can be a developer if they want to be and if you give them the right tools.

They need to be able to describe to an LLM what they're seeing and what they want to change. Then you need to provide the LLM enough context so it can infer what they mean, test the change and give it to them in a consumable format.

Non-devs wont read a PR diff, they must be able to use the application with the change to verify it.

To facilitate this kind of workflow I created a slack bot that can create github issues and assign an agent to the task.

When the task is finished the same bot can be asked to create a temporary environment with that new change deployed on it. The user can check that the change is what they expect, and if it's ok they can ask a dev to merge it.

These environments are protected through org SSO on top of the standard auth system we have. They are intentionally ephemeral and are killed off in just three hours. They don't run a full platform including all jobs etc, they are designed more around the kinds of things that non-devs can change and should be able to change - which is primarily user facing UI code and API contracts.

You will run out of "quick wins"

There is extreme productivity once you get the context set up. You can fix anything and everything and you do! These simple issues are easily diagnosed and fixed by LLMs. Noisy bugs in your logs that you have put off for years take just 10 minutes to PR. It's awesome.

But soon all these issues are gone and you are left with complex changes. Like your domain structures are out of date or your auth system needs to be refreshed. AI and LLMs still won't create good step by step architectural changes.

The changes I work on these days are still multi-week projects that need to be designed in diagrams and notes with requests for feedback before any code is produced.

Then I break down the job in to smaller tasks that LLMs can do (with many changes required to the first LLM output). So now we're kind of back to old school engineering again. It's just that now everyone is an "architect" and everyone on the team has an additional 6 virtual dev resources to implement their plans.

Not much has changed for these complex tasks really from a purely process perspective. I still need a good designer to help me. I still get a review from peers for tricky things. It's multiple PRs, feature flags, carefully staged releases to customers.

Using typed languages

This might seem obvious in 2026 but some high profile devs have said that types are not needed with modern LLMs, I disagree.

Most of the code in this platform was javascript and there were hundreds of old undefined and other errors that could have been caught in static checks. Migrating almost all code to typescript gave LLMs instant feedback around correctness. This closes the loop quickly for LLMs and we haven't had any new type class of errors since.

The types and naming of types gives LLMs a huge amount of context about your business. We don't have to describe very much about how things work in READMEs because the types are so descriptive.

It is slower to compile for sure, but TypeScript 7 is quite fast and the benefits outweigh this disadvantage.

Engineering grunt work is not solved by AI

We have agents on GitHub and they have access to all of our MCPs. So for a while I tried automating some things - frontend changes, library updates, Javascript to typescript migrations, test cleanup.

I would wake up in the morning to 4-5 PRs but 50% of them needed manual changes so the promise of the software factory is over blown right no for us. It's not possible to trust AI code. This is using the latest models, with a decent harness, good MCPs.

So I turned off these automations again. I'm hoping this will become better in time with better models but ultimately I feel I will do a full sweep of the platform and upgrade whole features in one go myself, rather than try to have AI do small PRs itself each day.

Custom harnesses are extremely difficult to get right

Our slack bot scout can take questions and use the MCPs to investigate issues. It returns findings to the user. It started with a completely custom for loop in a long running lambda and has since moved to CloudFlares harness with some custom instructions. I do limit the amount turns and we do use Gemini models there, the results are not great so far.

We are working on the next version of this which will likely be Pi running on some kind of container with access to a full machine that can be easily torn down and replaced. We are speed running what every other org seems to have converged on for long running agents.

Monitoring is easier

I used to monitor our customer inbound tickets when I got free time to look for patterns and fix priority issues before they grew into larger problems.

Now I have an agentic task that recieves a weebhook for each new ticket, it categorises it with a clustering algorithm and counts distinct tickets (tickets from multiple different customers). It will alert us on slack when one of these clusters is triggered so an engineer can investigate.

AI is bad at choosing architecture

LLMs will always find an architecture. And it might work. But it is often either too complex or too naive. For example I recently

AI is verbose with comments and tests

Fable in particular was awful with comments. They are verbose and describe what the code does rather than why the code was written. Most comments are not needed. The code itself should be self documenting. This is an old engineering rule and it;s a bit frustrating that LLMs don't follow this old pattern.

All the LLMs seem to add useless tests. For example tests that check if a hardcoded tailwind CSS class is present are completely useless really. They make change hard because you have to update any new css in two places now. Don't allow these into your code because LLMs will copy the patterns it finds.

CI will slow you down

We run CI on GitHub Actions. Our CI is on fire. The cost has doubled every month for the last 6 months EXCEPT for the fact we keep adding optimisations to try to reduce the load.

I've added a custom turbo server instance. We cache as much as we can. We use carefully chosen instances for the task type. All the workflows have conditions around selection and running. All of the apps use vite 8 and vitest, rolldown written in rust. Everything is upgraded to TypeScript 7. We have deleted a chunk of old tests.

We have investigated using third party runners like BlackSmith but our total cost doesn't justify 10-20% potential savings just yet.

Hey! Are you a developer?

🚀 Set Up Your Dev Environment in Minutes, Not Hours!

Tired of spending hours setting up a new development machine? I used to be, too—until I automated the entire process!

Now, I just run a single script, grab a coffee, and let my setup take care of itself.

Save 30+ hours configuring a new Mac or Windows (WSL) development environment.
Ensure consistency across all your machines.
Eliminate tedious setup and get coding faster!
Get Instant Access →