AI Transformation

An AI model can produce the code. We produce the confidence.

Generating software stopped being the hard part. Knowing it is correct, that somebody understands it, and that it can be changed safely next quarter, did not.

Triggers

The situations that lead here

Almost nobody arrives asking for an AI transformation. They arrive with one of these, and the cause turns out to be the same.

  • Output went up and confidence went down
  • Pull requests are too large and too fast for review to mean anything
  • The tests pass and you do not fully believe them
  • Nobody can explain a part of the system that shipped last month
  • Bugs are being fixed quickly and the same ones keep returning
  • A prototype somebody vibe-coded is now carrying real customers
  • Engineers disagree about how much to trust the tooling, and it is becoming cultural
  • You are asked whether the team is faster now and cannot honestly answer

Scope

What the work covers

The tooling is assumed. What gets built is the system that makes its output trustworthy.

  1. Guardrails

    The boundaries inside which anyone can move fast without asking permission: what can be generated freely, what needs a human decision first, and which parts of the system are simply off limits to a quick change.

  2. Review that catches things

    Review designed for code a person did not write line by line. Reviewing intent, contracts and edge cases rather than skimming a large diff that looks plausible because it was written to look plausible.

  3. Verification you can trust

    Tests that check behaviour against the requirement rather than mirroring the implementation. A model that writes both the code and its tests will make them agree, which proves nothing.

  4. The pipeline

    What has to be true before something reaches production, enforced automatically rather than remembered. As throughput rises, anything that depends on discipline alone stops holding.

  5. How the team actually works

    Where the tooling belongs in the day, where it does not, and how people keep building judgement rather than outsourcing it. This is the part that decides whether the change lasts.

  6. Architecture that survives the pace

    Clear boundaries and explicit contracts, so a fast change stays local. Without them, speed makes a codebase harder to reason about at exactly the rate it makes it bigger.

agent/desk: a completed multi-agent task run with its review summary
Guardrails, review and verification, so speed is something you can actually spend.

The argument

The constraint moved, and most delivery systems did not

For thirty years the expensive part of software was producing it. Nearly everything about how teams work was shaped by that: estimation, sprints, code review, the seniority ladder. All of it assumes writing code is slow and therefore the thing to economise on.

That assumption no longer holds, and the cost did not disappear. It moved. What is scarce now is confidence: knowing the thing is right, knowing someone understands it, knowing it can be changed next quarter without a week of archaeology first.

A team that adopts AI without changing anything else raises production and leaves verification where it was. The result is not faster delivery, it is a growing pile of work nobody has really checked, and the cost arrives later as incidents, rewrites and a system the team is afraid of.

So the work is not persuading people to use AI. It is rebuilding the parts of the delivery system that were quietly load-bearing, and are now carrying more than they were designed for.

Outcomes

What changes, and how you will know

Measured against the baseline written in the first month, not against impressions.

  1. Speed becomes something you can spend

    The gains stop being anecdotal. Work moves faster and reaches production without the quiet suspicion that it will come back.

  2. Review means something again

    Reviewers know what they are responsible for catching and have the time and the structure to catch it, whatever the size of the diff.

  3. The team understands what it ships

    Ownership is explicit. There is no part of the system where the honest answer is that a model wrote it and nobody has read it since.

  4. You can answer the question

    When the board asks whether AI has made the team faster, there is a real answer, with the baseline it is measured against.

What this is, and what it is not

What it is

  • Rebuilding review, testing and release for a higher rate of change
  • Guardrails written down, agreed, and enforced in the pipeline
  • Working practices that keep judgement inside the team
  • Honest measurement against a baseline
  • Led by Amir with senior engineers, alongside your team

What it is not

  • Tool selection or a licence procurement exercise
  • Training sessions on prompting
  • Building you an AI product or a model
  • A policy document nobody reads
  • Extra capacity to generate more code

If you are building an AI product, that is a different engagement and worth a separate conversation. This page is about a team that builds software with AI and needs the result to be trustworthy.

Questions executives ask

Is this about buying AI tools for the team?

No. Licences are the easy part and most teams already have them. The work is the system around the tooling: what gets reviewed, what gets tested, what an engineer is accountable for when a model wrote the first draft, and how you know the result is sound.

Are you going to tell us to stop using AI?

No. Teams that use it well move considerably faster. The failure mode is not using it, it is using it without changing anything else, so the volume of code goes up while the ability to verify it stays where it was.

Our engineers already use AI every day. What is left to do?

Usually quite a lot, and it is rarely visible from inside. Individual productivity is up, but review has become a formality, tests are generated alongside the code they are meant to check, and nobody can say which parts of the system anyone actually understands. That is a delivery risk, not a tooling gap.

How does this relate to vibe coding?

Vibe coding is a legitimate way to work, and it is very good at getting to something that runs. What it does not produce on its own is confidence that the thing is correct, that someone understands it, and that it can be changed safely in six months. That gap is the work.

Will you slow the team down?

Briefly, in the places where speed was being borrowed against later. The aim is throughput you can rely on, which is worth more than throughput you have to keep re-checking.

Results

See how this looks in practice

Real engagements, shared anonymously. The situation, the decisions, and what changed.

Explore the results

Find out where the confidence is leaking

Thirty minutes on how your team is working now, and a straight answer about where the confidence is actually leaking.