Skip to content
fghisi.com.br

· 6 min read

XP in 2026: the practices changed shape, the values didn't

What Extreme Programming still teaches in the age of LLMs: where AI sped up my team, where it stalled, and why fast feedback is still the game.

From 1999 to 2026

Extreme Programming was born in 1999 with a simple promise: embrace change instead of fighting it. Twenty-seven years later, the biggest change in how we build software did not come from a new framework, but from LLMs.

My thesis is that XP's practices changed shape, but its values became more important than ever. Communication, simplicity, feedback, courage and respect are exactly what separates a team that accelerates with AI from a team that just produces code faster.

I am not writing this from theory. I lead a technology team in payments, where a mistake costs real money, and we have seen AI get things right and wrong plenty of times.

What works: AI as support for coding

On my team, what has been working is using LLM-powered tools day to day, such as Cursor, Codex and Claude Code. They come in as support for the developer: generating a snippet, explaining legacy code, writing a test, suggesting a refactoring.

What they have in common is that the engineer stays in control of the cycle. They decide what to build, review what the tool proposes and integrate in small steps. AI speeds up each step, but it does not choose the path.

Through the lens of XP, this makes sense. The short feedback loop stays in the hands of those who know the domain, and AI only shortens the time between the idea and code that can be tested.

What didn't work: AI-DLC and the last 20%

We also tried the opposite path: letting AI drive the whole cycle. We used AI-DLC, an AWS methodology in which agents propose plans, code, tests and infrastructure, and humans validate each step across inception, construction and operation phases.

From zero to 80%, it was impressive. Requirements, design and a good part of the code came out fast. The final stretch did not: refinement, real tests and deploy stalled, and that is where the project lost its pace.

AI sped up the first 80%; the final stretch stalled AI sped up the first 80%; the final stretch stalled From 0 to 80% AI sped up: the problem was still generic Last 20% Stalled: our domain context RequirementsDesignCodeRefinementTestsDeploy
AI-DLC · where AI sped up and where it stalled

Looking back, it makes sense. The first 80% is where the problem is still generic and AI shines. The last 20% is where our system's context lives: partner integrations, payments business rules, edge cases and the production environment. That is exactly the knowledge XP has always argued should stay close to those who build.

XP calls this problem by its name. When feedback only arrives at the end, the surprise arrives with it. A process that generates a lot before integrating and testing trades short cycles for one big batch, and big batches fail in the final stretch.

Pair programming: the pair that never gets tired

An LLM is a pair that is always available, never gets tired and knows almost every library. In this arrangement, the human role is clear: they are the navigator, the one who sets the direction and questions the choices. AI becomes the driver.

But pair programming was never only about writing better code. It was also about spreading knowledge across the team and creating collective ownership of the code. A pair with AI does none of that. If everyone programs alone with their assistant, the team becomes a collection of islands.

Add AI to the pair, don’t replace the pair with AI Add AI to the pair, don’t replace the pair with AI One person + AI Knowledge stays with one person Two people + AI Knowledge flows across the team Personnavigates AIdrives Person APerson BAI
Pair programming · AI as the pair vs. AI added to the pair

That is why I would not replace the human pair with AI. I would add to it: two people and an assistant, or a mob with AI as one more participant. Knowledge keeps flowing, and speed comes along.

Test-first: the test becomes the specification

With LLMs, writing the test first gained a new role. The test is the most precise specification you can hand to the model: it says what the code must do, in a language the machine can verify on its own.

The risk is letting AI write the test and the code at the same time. Then it validates its own interpretation, and nobody is verifying anything. The test passes, but it may be testing the wrong thing.

The rule I propose is simple: the human defines the expected behavior in the test, and AI implements until it passes. In payments, where a wrong cent becomes an incident, this separation is not optional.

The person defines the behavior; AI implements until the test passes The person defines the behavior; AI implements until the test passes Person writes the testAI implementsTests runPassed?Person reviews and integrates no: AI adjusts the code yes
Test-first with AI · who does each step

This connects to our AI-DLC final stretch. The tests that mattered, with real data and real integrations, were left for the end. Had they come first, they would have guided the build instead of stalling it.

Small releases and simplicity

Generating code became cheap. Reviewing, integrating and shipping to production is still expensive. When AI produces a thousand lines in minutes, the bottleneck moves: away from writing and into review and deploy.

That is why small releases and continuous integration are back at the center. Small batches are reviewable, testable and reversible. Big batches, even AI-generated ones, carry the same risk as always.

Small batches bring feedback closer Small batches bring feedback closer Big batch Small batches Generate everythingIntegrate and testSurprises at the endGenerate, test, integrateGenerate, test, integrateGenerate, test, integrate Feedback only arrives once everything is done Feedback every cycle: problems show up early
Big batch vs. small batches · when feedback arrives

The same goes for simplicity. LLMs tend to generate more than needed: abstractions nobody asked for, handling for cases that don't exist, configuration for a hypothetical future. XP's YAGNI becomes a daily discipline: cut what is extra before integrating.

Documentation: from bureaucracy to context

XP has always been wary of heavy documentation. The idea was that clean code and well-written tests would explain the system better than any outdated document.

LLMs change that math. An agent is only as good as the context it receives. Instruction files, recorded architecture decisions and written business rules become fuel for the tool, not bureaucracy.

The difference is the audience. Before, documentation was for a human who might never read it. Now it is read on every interaction, by a tireless reader. This does not contradict XP: it is lean, living documentation with real use, exactly what it has always accepted.

Closing: fast feedback is still the game

The lesson I take from my team's experience is that AI works best when it speeds up short cycles, and worst when it tries to replace them. The support tools worked because they kept the engineer in the loop. AI-DLC stalled in the last 20% because real feedback came too late.

XP was never about the practices themselves. It was about getting fast feedback and having the courage to change course. LLMs don't make that obsolete. They make it cheaper for those who keep the discipline, and more expensive for those who abandon it.

Sources