Complicated vs. Complex: When to Analyze Up Front and When to Learn from Production

By Sergey Nosov

9 October 2026

You have probably read a postmortem that ends with this conclusion: “We should have thought of this.” It sounds like accountability. Often it is a category error. Suppose the outage came from an interaction that nobody could have seen in the parts alone. Then the conclusion asks for foresight about a system that shows its behavior only in hindsight.

The confusion starts with two words that software engineers use interchangeably: complicated and complex. In casual speech they are synonyms. In software they describe two different kinds of problems, and each kind rewards a different toolkit. My claim is that much of the frustration experienced engineers feel comes from mixing them up, not from a lack of skill. Think of the dashboards nobody reads, the specifications that do not survive contact with production, and the retrospective action items that feel hollow.

This article reduces the distinction to one question: is the answer knowable in advance, or only in retrospect? Ask it of each part of a problem, not of the whole system.

Complicated: Many Parts, Knowable Through Analysis

A complicated system has many parts, but the parts are knowable. Its behavior yields to analysis, specification, and careful reading. Given enough time and documentation, the right expert can work out its causes and effects on paper, before it ever runs.

A Swiss watch is complicated. So is a compiler, and a single, well-documented service usually is too. None of them is trivial, and understanding them can take real effort, but in principle all of them can be understood in advance.

Complex: The Behavior Lives in the Interactions

A complex system shows emergent behavior: the interactions between its parts produce outcomes that the parts alone cannot predict. No expert, however senior, can fully specify it in advance. Much of what it does shows up only under real load, with real users, over real time.

A fleet of services running in production is complex. So are a city and a financial market. In each of them, the behavior lives in the interactions, not in the components.

Where the Distinction Comes From

The distinction is not new. It is built into Cynefin (pronounced kuh-nev-in), a framework that Dave Snowden created in the late 1990s. In “A Leader’s Framework for Decision Making,” published in the Harvard Business Review in November 2007, Snowden and Mary E. Boone explained the difference with a car and a forest:

In a complicated context, at least one right answer exists. In a complex context, however, right answers can’t be ferreted out. It’s like the difference between, say, a Ferrari and the Brazilian rainforest. Ferraris are complicated machines, but an expert mechanic can take one apart and reassemble it without changing a thing. The car is static, and the whole is the sum of its parts. The rainforest, on the other hand, is in constant flux—a species becomes extinct, weather patterns change, an agricultural project reroutes a water source—and the whole is far more than the sum of its parts.

Each kind of context calls for a different response. In a complicated context, they wrote, leaders “must sense, analyze, and respond,” and the analysis often takes expertise. In a complex context, “we can understand why things happen only in retrospect.” Leaders there “need to probe first, then sense, and then respond,” and they learn from “experiments that are safe to fail.”

Cynefin has more domains than these two. The 2007 article names five: simple, complicated, complex, chaotic, and disorder, which “applies when it is unclear which of the other four contexts is predominant.” The framework’s wiki now calls the first one clear and the last one confused. For software teams, the boundary that matters most is the one between complicated and complex. My Software Development Principles series has an entry on Cynefin, with lists of signs that a team is treating one domain as another.

Same Code, New Questions

In software, the line between the two kinds runs through the same code. A single service you wrote is usually complicated. Put the same service behind a load balancer, next to a dozen other services, and let clients you have never met call it. Now it is part of something complex. You did not change the code. You changed which questions about it you can answer in advance.

Old code can also show new behavior when something around it changes. Amazon Web Services described such a case in its summary of the outage in its Northern Virginia region on 7 December 2021. It began when an automated activity to scale one AWS service “triggered an unexpected behavior from a large number of clients inside the internal network.” The surge of connections overwhelmed the devices between AWS’s internal and main networks, and the delays led to “even more connection attempts and retries.”

AWS wrote that the clients had “well tested request back-off behaviors,” but “a latent issue prevented these clients from adequately backing off during this event.” It added: “This code path has been in production for many years but the automated scaling activity triggered a previously unobserved behavior.” The code path was old. The behavior was new.

AWS’s account names several contributors: the scaling activity, the latent issue in the clients’ back-off, the devices between the two networks, and the retries that kept the congestion going. Once found, the latent issue was a complicated problem: a defect to fix and test. The outage was complex. It took the combination.

Richard I. Cook, whose paper I recommend below, describes latent defects as the normal state of complex systems: they “contain changing mixtures of failures latent within them.”

The Toolkit for Complicated Problems

Complicated problems reward the classical engineering toolkit. It works because the answer is knowable up front, so thinking up front pays off:

Everything on this list narrows the gap between what we intend and what the system will do. For a complicated problem, more of this work closes more of the gap. I have written separately about where each kind of test belongs and about the anatomy of a good code review.

Formal specification can answer questions about interactions, too. In a 2014 report, engineers at Amazon Web Services described model checking a specification of DynamoDB’s replication and fault-tolerance design. The checker found “a bug that could lead to losing data if a particular sequence of failures and recovery steps was interleaved with other processing.” The shortest trace that showed it took thirty-five high-level steps, and the bug “had passed unnoticed through extensive design reviews, code reviews, and testing.” Whether a protocol can lose data is a question you can answer in advance, with enough rigor, even inside a complex system. The answer holds only within the model: the properties specified, the failures modeled, and the configurations checked. Asked how they know “that the executable code correctly implements the verified design,” the authors write: “The answer is that we don’t.”

The Toolkit for Complex Problems

Complex problems reward a different toolkit, because you cannot specify in advance what you can only discover:

Notice what is missing from this list: more up-front prediction. Up-front work still matters, but its target changes: instead of predicting the behavior, you bound it, prepare to see it, and plan to undo it. Before a rollout, work out how far retries could multiply load, decide what normal looks like and which numbers would stop the rollout, and know how you would roll back.

Symptoms of the Wrong Toolkit

When a team treats a complex problem as a complicated one, the signs are familiar:

Each of these can have other causes. When several show up together, suspect the complicated toolkit applied to a complex problem. More effort inside that toolkit will not fix them.

The first symptom deserves a closer look, because hindsight distorts the postmortem itself. Cook explains why. “Knowledge of the outcome makes it seem that events leading to the outcome should have appeared more salient to practitioners at the time than was actually the case.” He also argues that “post-accident attribution to a ‘root cause’ is fundamentally wrong,” because an accident in a complex system needs several contributors at once. A postmortem that hunts for the one thing someone should have foreseen applies the complicated toolkit to the past.

Sometimes, though, the answer really was knowable: a missing input check, an unhandled error code, a timeout nobody set. So ask the same question of each finding, from where the team stood before the outage. If it was knowable then, strengthen the review, the test, or the checklist that catches that class of defect. If not, invest in seeing it sooner and recovering faster.

The Opposite Mistake

The mismatch also runs the other way. Treating a complicated problem as a complex one throws away the cheap answers that analysis would have given you. Suppose you ship a date-handling change behind a canary to learn whether it survives 29 February. The canary may never see that date. A unit test with a controlled clock covers leap days and century years in milliseconds. Skipping a design review because production will tell you anyway is the same mistake. When analysis is cheaper than an experiment, analyze first. Save experiments for the questions that analysis cannot settle, or cannot settle in time.

One Question Before You Choose Your Tools

Before you choose your tools, ask the question of each part of the problem in front of you: is it knowable in advance, or only in retrospect?

A question is usually complicated when its answer depends on things you can pin down in advance: the code, its inputs, and the documented behavior of what it calls. It is usually complex when the answer depends on what other people and systems will do: real users, real traffic, other services under load.

When you cannot tell, the problem is usually several questions in one. Snowden and Boone’s advice for that state, which they called disorder, works for software too: “break down the situation into constituent parts” and give each part its own answer. Keep the interactions between the parts as questions of their own. Give expensive analysis a time box. If it stalls even with the right expertise and evidence, treat the question as complex and probe. A few examples:

One exception: an incident that is still spreading leaves no time to analyze or to probe. Snowden and Boone’s advice for a chaotic context applies: “A leader must first act to establish order.” Roll back, shed load, or flip the kill switch, and learn afterward.

Moving the Line

The two toolkits are not rivals, and the line between the two kinds of problems is not fixed. Your answer to the question is a working judgment, not a permanent label. You can move the line in two ways.

Discoveries move it. What the complex toolkit discovers, the complicated toolkit can then pin down, as AWS set out to do with the back-off defect its outage revealed.

Constraints move it too. Timeouts, retry budgets, rate limits, and bulkheads shrink the space of possible interactions, so more of the system’s behavior becomes predictable. The Cynefin wiki defines order in these terms: “Order is constrained to the point where future outcomes are predictable as long as the constraints can be sustained.” AWS’s remediation added such constraints. It “immediately disabled the scaling activities that triggered this event.” It also deployed “additional network configuration that protects potentially impacted networking devices even in the face of a similar congestion event.”

Cook warns against “end-of-the-chain measures,” remedies that block the last step before an accident. “Instead of increasing safety, post-accident remedies usually increase the coupling and complexity of the system.” Constrain the mechanism an outage revealed, not the exact chain of events from last time, and keep each constraint simple enough to reason about.

The line also moves back. A system that changes can turn a settled question into an open one again, and one of Cook’s observations gives the reason: “Change introduces new forms of failure.”

Read Cook’s Paper

If this distinction is useful to you, read the paper that shaped my thinking about it: “How Complex Systems Fail.” Cook wrote it in 1998; the version in circulation dates from April 2000. Its eighteen numbered observations fill four pages and take about fifteen minutes to read. Cook wrote with medicine and other hazardous fields in mind, but the observations read like a description of a production fleet. One of them puts the complex half of this article in a sentence. “Safety is an emergent property of systems; it does not reside in a person, device or department of an organization or system.” The paper is likely to change how you read postmortems, and how you write them.

Checklist

Further Reading