Unit, Integration, and End-to-End Tests: Where a Test Belongs and Why the Pyramid Holds

By Sergey Nosov

3 October 2026

Ask two engineers whether a given test is a unit test or an integration test, and you can lose an afternoon. Ham Vocke opened The Practical Test Pyramid, published on Martin Fowler’s site in February 2018, with the same observation: “What I mean when I talk about unit tests can be slightly different from your understanding. With integration tests it's even worse.”

The argument is avoidable, because the layers are not defined by what a test asserts. They are defined by what a test is allowed to touch, how fast it runs, and where its failure points. This article gives a decision rule built on those three properties and explains why the test pyramid follows from them. It shows examples in C# and JavaScript that I ran before publishing, and it ends with what flaky tests and AI-generated tests do to the picture.

What a Unit Test Is

A unit test checks the smallest piece of behavior you can exercise on its own: a function, a method, or a small class. It runs inside the test process and touches nothing the code under test does not own. Because of that, it is fast, it is deterministic, and when it fails it names the line. Mike Cohn put the appeal in one sentence in his post on the test automation pyramid: “Automated unit tests are wonderful because they give specific data to a programmer—there is a bug and it's on line 47.”

Suppose a function computes the sales tax on a price. A unit test passes in a few prices and rates and checks that the arithmetic and the rounding come out right. It does not connect to anything. If the code is hard to test this way, the test is telling you something about the design: the function reaches for a database, a clock, or a global it should have been given.

The form of such a test is older than most of the code it protects. Kent Beck built a small framework to organize and run unit tests in Smalltalk. Then, as Martin Fowler tells it, “JUnit was born on a flight from Zurich to the 1997 OOPSLA in Atlanta.” Beck was flying with Erich Gamma, and the JUnit FAQ confirms that “JUnit was originally written by Erich Gamma and Kent Beck.” Ports followed; in Fowler’s words, “nearly every language has at least one JUnit port.” That heritage is why NUnit, xUnit.net, Jest, and PHPUnit all share the same shape: setup, assertions, and a runner.

One distinction inside the unit layer settles many arguments before they start. In his 2014 note on unit tests, Fowler borrows two terms from Jay Fields. A solitary unit test replaces the collaborators of the unit under test with test doubles. A sociable unit test lets the unit call its real collaborators, as long as everything stays in process and fast. Both are unit tests. A test that uses three real classes is not an integration test just because it uses three classes.

The vocabulary for the stand-ins comes from Gerard Meszaros, whose terms Fowler summarized in 2006: dummies, fakes, stubs, spies, and mocks. The two that get confused most are the last two. “Spies are stubs that also record some information based on how they were called.” “Mocks are pre-programmed with expectations which form a specification of the calls they are expected to receive.” Knowing which one you are reaching for keeps a test honest about what it proves.

What an Integration Test Is

An integration test crosses a boundary your code does not own: a database, the file system, a message broker, or the request pipeline of your web framework. Microsoft’s guidance for ASP.NET Core defines the layer this way: “Integration tests ensure that an app's components function correctly at a level that includes the app's supporting infrastructure, such as the database, file system, and network.”

What the test verifies is the contract between your code and the thing on the other side. The SQL really runs against the real engine. The route really binds and the JSON really round-trips. Suppose a repository class reads and writes orders. A unit test of it would prove only that it calls the database driver the way you expected; an integration test proves that an order you store is the order you get back.

Three kinds of tooling make this layer cheap enough to run on every build. The first is an in-process host for your application. In ASP.NET Core, the Microsoft.AspNetCore.Mvc.Testing package provides WebApplicationFactory, which boots the application on an in-memory TestServer so an HttpClient can call it without a port. The second is throwaway infrastructure. Testcontainers describes itself as “an open source library for providing throwaway, lightweight instances of databases, message brokers, web browsers, or just about anything that can run in a Docker container.” It has libraries for Java, Go, .NET, Node.js, Python, and more. The third is state reset. Jimmy Bogard’s Respawn “resets the database back to a clean, empty state by intelligently deleting data from tables,” so each test starts from a known point without rolling back a transaction.

Integration tests are slower than unit tests and their failures are less precise. A red test says the boundary is broken; it rarely says which side broke it.

What an End-to-End Test Is

An end-to-end test exercises the whole system the way a user or a client does: through the browser or the public API, against a deployed stack with its real configuration. It is the only layer that catches wiring mistakes, such as an environment variable that points at the wrong queue or a reverse proxy that strips a header. Microsoft’s Playwright is one current choice; its documentation says “Playwright Test is an end-to-end test framework for modern web apps. It bundles test runner, assertions, isolation, parallelization and rich tooling.” Cypress is another.

The price is steep. These tests are the slowest to run, the most expensive to maintain, and the hardest to debug. When one fails, the message is “checkout failed,” and the cause can sit anywhere in the stack.

A Decision Rule: Classify by What the Test Touches

Google solved the naming argument by refusing to have it. In December 2010, Simon Stewart described on the Google Testing Blog how Google labels tests by size rather than by kind, with rules a machine can enforce:

Stewart maps the sizes back to the familiar names: “A Small test equates neatly to a unit test, a Large test to an end-to-end or system test and a Medium test fits neatly in the middle.” The point is that the label follows from what the test does, so two engineers reading the same test reach the same answer.

My rule has three questions, in order.

  1. Does the test leave the process? If not, it is a unit test, however many real classes it uses.
  2. Does it reach something on this machine or in a container it started? A database, a file, a port on localhost, your framework’s in-memory host: that is an integration test.
  3. Does it go through the public surface of a deployed system, or call a system you do not run? That is an end-to-end test.

Two more properties confirm the answer. Where does a failure point? A unit test points at a line, an integration test at a boundary, an end-to-end test at a journey. And how long does it take? Google’s limits are ceilings, not targets. A “unit” test that takes a full second is almost always touching something it should not. Fowler’s 2014 note explains why speed matters so much at the base: “The common properties of unit tests—small scope, done by the programmer herself, and fast—mean that they can be run very frequently.” A suite you can run on every save changes how you work; a suite you run before lunch does not.

If your team still prefers other words, keep them, but write the definitions down and apply them consistently. Vocke gives the same advice. The entry on the test pyramid in my Software Development Principles series takes the same position: define the layers by speed, isolation, and determinism, not by vocabulary.

Why the Pyramid Follows

Fowler credits the test pyramid to Mike Cohn’s 2009 book Succeeding with Agile. Cohn drew three layers: unit tests at the base, service tests in the middle, and user-interface tests at the top. In his own words, “Service-level testing is about testing the services of an application separately from its user interface.” Of the top layer, he writes: “Automated user interface testing is placed at the top of the test automation pyramid because we want to do as little of it as possible.” The middle layer is the one he called forgotten: “Where many organizations have gone wrong in their test automation efforts over the years has been in ignoring this whole middle layer of service testing.”

Fowler’s 2012 summary gives the reasons the shape is not arbitrary. Of tests driven through the user interface, he writes, “Testing through the UI like this is slow, increasing build times.” Then: “Most importantly such tests are very brittle.” Once you have the decision rule above, the pyramid is just its consequence. Each step up the layers costs more in four ways at once:

Confidence runs the other way. Only the top layer proves the assembled system works as a user sees it. The pyramid is the trade: buy most of your coverage where it is cheap, and spend on the expensive layer only for the journeys that matter.

Two shapes show what happens when teams ignore the trade. Alister Scott named the inverted pyramid the software testing ice-cream cone in January 2012. Fabio Pereira of Thoughtworks described it in 2014: “This happens when there is not enough low-level testing (unit, integration and component), too many tests that run through the Graphical User Interface (GUI) and an even larger number of manual tests.” Pereira then added a second anti-pattern from his own projects, the cupcake, which is wide at every layer. It appears when “There are different teams who write different levels of tests,” and the teams do not talk: “This results in duplication—the same scenario ends up being automated at many different levels.”

The shape is negotiable; the principle is not. Kent C. Dodds argues for a testing trophy, with its widest band at integration tests, under a line he credits to a 2016 tweet by Guillermo Rauch: “Write tests. Not too many. Mostly integration.” Dodds’s case is that “Integration tests strike a great balance on the trade-offs between confidence and speed/expense.” His guiding rule is that “The more your tests resemble the way your software is used, the more confidence they can give you.” Where that is true for your system, widen the middle. What every shape agrees on is to keep the slow, brittle layer small, and to push each check down to the cheapest layer that can catch it. Vocke states the operating rule: “If a higher-level test spots an error and there's no lower-level test failing, you need to write a lower-level test.”

Examples in C# and JavaScript

I ran every sample below before publishing. The C# code was checked with xUnit.net v3 4.0.1 on .NET 10, created from the SDK’s xunit3 template. The JavaScript code was checked with Jest 30 on Node.js 22.

A Unit Test in C# with xUnit.net

Here is the sales-tax function from above, with the price validated and the result rounded to cents:

public static class SalesTax
{
    public static decimal Apply(decimal price, decimal rate)
    {
        if (price < 0)
            throw new ArgumentOutOfRangeException(nameof(price));
        return Math.Round(price * (1 + rate), 2,
            MidpointRounding.ToEven);
    }
}

And its tests. A [Fact] is a single case; a [Theory] runs once per [InlineData] row:

public class SalesTaxTests
{
    [Fact]
    public void Apply_AddsTax_ForPositivePrice()
    {
        var total = SalesTax.Apply(100m, 0.08m);

        Assert.Equal(108m, total);
    }

    [Theory]
    [InlineData(0, 0.08, 0)]
    [InlineData(19.99, 0.0825, 21.64)]
    public void Apply_RoundsToCents(
        decimal price, decimal rate, decimal want)
    {
        Assert.Equal(want, SalesTax.Apply(price, rate));
    }

    [Fact]
    public void Apply_Rejects_NegativePrice()
    {
        Assert.Throws<ArgumentOutOfRangeException>(
            () => SalesTax.Apply(-1m, 0.08m));
    }
}

Nothing here leaves the process, so by the rule above this is a unit test, and all four cases finish in well under a second. One trap for people who move between frameworks: [Test] and Assert.AreEqual are NUnit. xUnit.net uses [Fact], [Theory], and Assert.Equal, and a slide that mixes them will not compile. The names follow one convention, method, expected behavior, and condition, so that a failing test reads as a sentence in the runner’s output.

A Unit Test in JavaScript with Jest

const { sum } = require('./sum');

test('adds two numbers', () => {
  expect(sum(2, 3)).toBe(5);
});

test('treats a missing addend as NaN', () => {
  expect(sum(2)).toBeNaN();
});

Same principle, different syntax. Jest 30 shipped on 4 June 2025; Mocha with Chai remains a common alternative, and since Node.js 20 the runtime’s own node:test module has been stable, so a small library can test itself with no dependency at all.

An Integration Test in C# with WebApplicationFactory

Suppose a minimal API exposes GET /prices/{id}, returning a price for a known SKU and 404 otherwise. The test below boots the real application in memory and calls it over HTTP, so it proves routing, model binding, and serialization together:

using System.Net;
using Microsoft.AspNetCore.Mvc.Testing;

public class PricesTests
    : IClassFixture<WebApplicationFactory<Program>>
{
    private readonly HttpClient _client;

    public PricesTests(WebApplicationFactory<Program> factory)
    {
        _client = factory.CreateClient();
    }

    [Fact]
    public async Task Get_ReturnsPrice_ForKnownSku()
    {
        var response = await _client.GetAsync("/prices/sku-1");

        response.EnsureSuccessStatusCode();
        var body = await response.Content.ReadAsStringAsync();
        Assert.Contains("19.99", body);
    }

    [Fact]
    public async Task Get_Returns404_ForUnknownSku()
    {
        var response = await _client.GetAsync("/prices/nope");

        Assert.Equal(HttpStatusCode.NotFound, response.StatusCode);
    }
}

The test project references the API project and the Microsoft.AspNetCore.Mvc.Testing package, and the API’s Program.cs ends with public partial class Program { } so the factory can name it. By the rule, this is an integration test: it crosses the framework’s request pipeline, but it stays on this machine. Swap the in-memory host for a database in a container, and it is still an integration test, only a slower one.

An End-to-End Test with Playwright

This sample is illustrative. I did not run it, because it needs a deployed application to point at:

import { test, expect } from '@playwright/test';

test('checkout shows an order number', async ({ page }) => {
  await page.goto('/cart');
  await page.getByRole('button', { name: 'Check out' }).click();
  await expect(
    page.getByRole('heading', { name: /order/i })
  ).toBeVisible();
});

Note what it asserts. It finds elements by their accessibility role and name, and it checks that a heading about the order appears. It does not compare a block of page text against a fixed string, which is the fastest way to make an end-to-end test break on every copy change.

Who Writes Which Tests

Developers naturally live in the bottom two layers, because those tests sit next to the code and answer in seconds. Testers naturally live at the top, in end-to-end scenarios, exploratory testing, and the performance and security checks that no unit test can make. The failure mode is treating those as two separate programs. Pereira’s cupcake grows exactly there: separate teams, working in sequence, each automating the same scenario at its own level.

The fix is a shared test strategy, decided per feature. Which checks go in unit tests, which contracts need an integration test, and which one or two journeys earn an end-to-end test? The decision rule above gives that conversation a vocabulary.

Anti-Patterns

Flaky Tests and Test Data

A flaky test fails sometimes without any change to the code. Google published the scale of the problem in May 2016, in John Micco’s post Flaky Tests at Google and How We Mitigate Them. The definition first: “We define a ‘flaky’ test result as a test that exhibits both a passing and a failing result with the same code.” Then the numbers: “we see a continual rate of about 1.5% of all test runs reporting a ‘flaky’ result,” and “Almost 16% of our tests have some level of flakiness associated with them!” Worse, “about 84% of the transitions we observe from pass to fail involve a flaky test!” Micco names the root causes as “concurrency, relying on non-deterministic or undefined behaviors, flaky third party code, infrastructure problems, etc.”

The cost is not the reruns. It is trust. “It is human nature to ignore alarms when there is a history of false signals coming from a system.” A suite that cries wolf once a day trains the team to click “retry,” and the real failure hides among the false ones.

Google’s practices translate down to any team. Quarantine a flaky test as soon as it is identified: in Micco’s description, quarantining “removes the test from the critical path and files a bug for developers to reduce the flakiness.” The test keeps running, but it no longer blocks anyone, and it has an owner and a deadline. Then remove the sources of nondeterminism. Seed any randomness so a failure can be replayed. Reset shared state between tests. Never sleep to wait for something; poll for the condition or await the event. Google’s size rules forbid sleep statements in small and medium tests for exactly this reason.

Test data deserves the same care. Build it with factories or builders that produce one valid object with sensible defaults and let a test override only the field it is about. A test that constructs a forty-field object by hand hides its intent in the noise.

What AI-Generated Tests Change

Generating tests is now cheap. A coding assistant will scaffold a test class, propose edge cases, and write the mocks in seconds. That is a real gain, and it does not change what a test is for. A generated test can be syntactically correct and semantically weak: it may assert the behavior the code happens to have, including the bug, or it may miss the one boundary that matters to the business.

Meta’s experience shows what it takes to use generation responsibly. In February 2024, Nadia Alshahwan and colleagues described TestGen-LLM, a tool that uses language models to improve existing unit tests. Its design accepts a generated test only after it clears “a set of filters that assure measurable improvement over the original test suite, thereby eliminating problems due to LLM hallucination.” The yield is instructive: “75% of TestGen-LLM's test cases built correctly, 57% passed reliably, and 25% increased coverage.” Across Meta’s test-a-thons, “73% of its recommendations being accepted for production deployment by Meta software engineers” shows that human review stayed in the loop even after the filters.

Read those numbers as a process, not a verdict. The filters are a machine review: does it build, does it pass reliably, does it add coverage. The engineer’s review is the one that no filter replaces: is the assertion the right assertion? My book Code You Did Not Write treats this as a case of technical direction: review unfamiliar code for what it actually does, not for what it appears to do. A generated test is unfamiliar code. Treat the assistant as a fast junior who drafts; keep the judgment about what to assert.

Takeaways

Further Reading