Complicated vs. Complex: When to Analyze Up Front and When to Learn from Production
By Sergey Nosov
9 October 2026
You have probably read a postmortem that ends with this conclusion: “We should have thought of this.” It sounds like accountability. Often it is a category error. Suppose the outage came from an interaction that nobody could have seen in the parts alone. Then the conclusion asks for foresight about a system that shows its behavior only in hindsight.
The confusion starts with two words that software engineers use interchangeably: complicated and complex. In casual speech they are synonyms. In software they describe two different kinds of problems, and each kind rewards a different toolkit. My claim is that much of the frustration experienced engineers feel comes from mixing them up, not from a lack of skill. Think of the dashboards nobody reads, the specifications that do not survive contact with production, and the retrospective action items that feel hollow.
This article reduces the distinction to one question: is the answer knowable in advance, or only in retrospect? Ask it of each part of a problem, not of the whole system.
Complicated: Many Parts, Knowable Through Analysis
A complicated system has many parts, but the parts are knowable. Its behavior yields to analysis, specification, and careful reading. Given enough time and documentation, the right expert can work out its causes and effects on paper, before it ever runs.
A Swiss watch is complicated. So is a compiler, and a single, well-documented service usually is too. None of them is trivial, and understanding them can take real effort, but in principle all of them can be understood in advance.
Complex: The Behavior Lives in the Interactions
A complex system shows emergent behavior: the interactions between its parts produce outcomes that the parts alone cannot predict. No expert, however senior, can fully specify it in advance. Much of what it does shows up only under real load, with real users, over real time.
A fleet of services running in production is complex. So are a city and a financial market. In each of them, the behavior lives in the interactions, not in the components.
Where the Distinction Comes From
The distinction is not new. It is built into Cynefin (pronounced kuh-nev-in), a framework that Dave Snowden created in the late 1990s. In “A Leader’s Framework for Decision Making,” published in the Harvard Business Review in November 2007, Snowden and Mary E. Boone explained the difference with a car and a forest:
In a complicated context, at least one right answer exists. In a complex context, however, right answers can’t be ferreted out. It’s like the difference between, say, a Ferrari and the Brazilian rainforest. Ferraris are complicated machines, but an expert mechanic can take one apart and reassemble it without changing a thing. The car is static, and the whole is the sum of its parts. The rainforest, on the other hand, is in constant flux—a species becomes extinct, weather patterns change, an agricultural project reroutes a water source—and the whole is far more than the sum of its parts.
Each kind of context calls for a different response. In a complicated context, they wrote, leaders “must sense, analyze, and respond,” and the analysis often takes expertise. In a complex context, “we can understand why things happen only in retrospect.” Leaders there “need to probe first, then sense, and then respond,” and they learn from “experiments that are safe to fail.”
Cynefin has more domains than these two. The 2007 article names five: simple, complicated, complex, chaotic, and disorder, which “applies when it is unclear which of the other four contexts is predominant.” The framework’s wiki now calls the first one clear and the last one confused. For software teams, the boundary that matters most is the one between complicated and complex. My Software Development Principles series has an entry on Cynefin, with lists of signs that a team is treating one domain as another.
Same Code, New Questions
In software, the line between the two kinds runs through the same code. A single service you wrote is usually complicated. Put the same service behind a load balancer, next to a dozen other services, and let clients you have never met call it. Now it is part of something complex. You did not change the code. You changed which questions about it you can answer in advance.
Old code can also show new behavior when something around it changes. Amazon Web Services described such a case in its summary of the outage in its Northern Virginia region on 7 December 2021. It began when an automated activity to scale one AWS service “triggered an unexpected behavior from a large number of clients inside the internal network.” The surge of connections overwhelmed the devices between AWS’s internal and main networks, and the delays led to “even more connection attempts and retries.”
AWS wrote that the clients had “well tested request back-off behaviors,” but “a latent issue prevented these clients from adequately backing off during this event.” It added: “This code path has been in production for many years but the automated scaling activity triggered a previously unobserved behavior.” The code path was old. The behavior was new.
AWS’s account names several contributors: the scaling activity, the latent issue in the clients’ back-off, the devices between the two networks, and the retries that kept the congestion going. Once found, the latent issue was a complicated problem: a defect to fix and test. The outage was complex. It took the combination.
Richard I. Cook, whose paper I recommend below, describes latent defects as the normal state of complex systems: they “contain changing mixtures of failures latent within them.”
The Toolkit for Complicated Problems
Complicated problems reward the classical engineering toolkit. It works because the answer is knowable up front, so thinking up front pays off:
- documentation and design review;
- type systems and static analysis;
- thorough unit and integration tests;
- code review and pair programming;
- formal specification, where the stakes justify the cost.
Everything on this list narrows the gap between what we intend and what the system will do. For a complicated problem, more of this work closes more of the gap. I have written separately about where each kind of test belongs and about the anatomy of a good code review.
Formal specification can answer questions about interactions, too. In a 2014 report, engineers at Amazon Web Services described model checking a specification of DynamoDB’s replication and fault-tolerance design. The checker found “a bug that could lead to losing data if a particular sequence of failures and recovery steps was interleaved with other processing.” The shortest trace that showed it took thirty-five high-level steps, and the bug “had passed unnoticed through extensive design reviews, code reviews, and testing.” Whether a protocol can lose data is a question you can answer in advance, with enough rigor, even inside a complex system. The answer holds only within the model: the properties specified, the failures modeled, and the configurations checked. Asked how they know “that the executable code correctly implements the verified design,” the authors write: “The answer is that we don’t.”
The Toolkit for Complex Problems
Complex problems reward a different toolkit, because you cannot specify in advance what you can only discover:
- Observability: metrics, logs, and distributed traces, so you can see what the system is actually doing. Make sure it keeps working when the system it watches does not. In the AWS outage, the congestion impaired real-time monitoring, and operators relied on logs instead.
- Canary deployments and progressive rollout: a surprise that shows up at small scale reaches a slice of traffic, say 1 percent, before it reaches everyone. The slice limits direct exposure, not every effect. A canary that writes bad data to a shared database or floods a shared dependency can hurt everyone, so watch the shared resources as well as the canary.
- Feature flags and kill switches: they let you respond without redeploying, but, like monitoring, they help only if their control path does not share the failure. Flags carry risks of their own: my article on deleting dead code tells how a reused flag woke up code at Knight Capital that had been dormant for nine years.
- Experiments: A/B tests show how users respond, and chaos drills and game days test how the system responds to stresses you choose, so you learn before an incident teaches you. The Principles of Chaos Engineering list “non-failure events like a spike in traffic or a scaling event” among the events worth introducing. One of their principles is to “Minimize Blast Radius.”
- A culture that tolerates discovery: one that accepts learning how the system really behaves instead of insisting that it was all specified in advance. Snowden and Boone warn about the opposite pull, the temptation “to demand fail-safe business plans with defined outcomes.”
Notice what is missing from this list: more up-front prediction. Up-front work still matters, but its target changes: instead of predicting the behavior, you bound it, prepare to see it, and plan to undo it. Before a rollout, work out how far retries could multiply load, decide what normal looks like and which numbers would stop the rollout, and know how you would roll back.
Symptoms of the Wrong Toolkit
When a team treats a complex problem as a complicated one, the signs are familiar:
- Postmortems and action items that demand more foresight. “We should have thought of this” treats a discovery as a lapse and asks the team to predict what only production could show.
- Dashboards nobody reads. The instruments exist, but no decision depends on them.
- Specifications that do not survive contact with production. They described the parts, and production added the interactions.
- Monitoring bolted on after the first real outage. Nobody planned to watch the system until it failed.
Each of these can have other causes. When several show up together, suspect the complicated toolkit applied to a complex problem. More effort inside that toolkit will not fix them.
The first symptom deserves a closer look, because hindsight distorts the postmortem itself. Cook explains why. “Knowledge of the outcome makes it seem that events leading to the outcome should have appeared more salient to practitioners at the time than was actually the case.” He also argues that “post-accident attribution to a ‘root cause’ is fundamentally wrong,” because an accident in a complex system needs several contributors at once. A postmortem that hunts for the one thing someone should have foreseen applies the complicated toolkit to the past.
Sometimes, though, the answer really was knowable: a missing input check, an unhandled error code, a timeout nobody set. So ask the same question of each finding, from where the team stood before the outage. If it was knowable then, strengthen the review, the test, or the checklist that catches that class of defect. If not, invest in seeing it sooner and recovering faster.
The Opposite Mistake
The mismatch also runs the other way. Treating a complicated problem as a complex one throws away the cheap answers that analysis would have given you. Suppose you ship a date-handling change behind a canary to learn whether it survives 29 February. The canary may never see that date. A unit test with a controlled clock covers leap days and century years in milliseconds. Skipping a design review because production will tell you anyway is the same mistake. When analysis is cheaper than an experiment, analyze first. Save experiments for the questions that analysis cannot settle, or cannot settle in time.
One Question Before You Choose Your Tools
Before you choose your tools, ask the question of each part of the problem in front of you: is it knowable in advance, or only in retrospect?
- If it is knowable in advance, reach for the complicated toolkit. Specify it, test it, review it, and, where the stakes justify it, prove it. The up-front work will pay off.
- If it is knowable only in retrospect, reach for the complex toolkit. Instrument it, roll it out slowly, and build the feedback loop. Be honest with yourself that the system will teach you things no specification could have.
A question is usually complicated when its answer depends on things you can pin down in advance: the code, its inputs, and the documented behavior of what it calls. It is usually complex when the answer depends on what other people and systems will do: real users, real traffic, other services under load.
When you cannot tell, the problem is usually several questions in one. Snowden and Boone’s advice for that state, which they called disorder, works for software too: “break down the situation into constituent parts” and give each part its own answer. Keep the interactions between the parts as questions of their own. Give expensive analysis a time box. If it stalls even with the right expertise and evidence, treat the question as complex and probe. A few examples:
- Is the new search endpoint safe from SQL injection? Knowable in advance. Review it for queries built from strings, and test it with hostile input.
- How long should the checkout service wait for the fraud check? Partly knowable. That a timeout exists, what the customer sees when it fires, and a starting value from measured latency and the time budget can all be settled in advance. How retries and the fraud check’s other callers behave when it slows down at peak traffic is knowable only in retrospect. Rehearse that in a game day, and watch for it in production.
- Will the new cache hold up under real traffic? Partly knowable. Load-test what you can predict, then mirror production traffic to the cache or ramp it up in steps. Watch the hit rate, the tail latency, and the load on whatever sits behind the cache. A 1 percent slice will not show you the hit rate at full load, or what happens when a popular key expires.
- Will more customers finish the new checkout flow? Knowable only in retrospect. Probe early with usability sessions, then run an A/B test on the completion metric defined in advance.
One exception: an incident that is still spreading leaves no time to analyze or to probe. Snowden and Boone’s advice for a chaotic context applies: “A leader must first act to establish order.” Roll back, shed load, or flip the kill switch, and learn afterward.
Moving the Line
The two toolkits are not rivals, and the line between the two kinds of problems is not fixed. Your answer to the question is a working judgment, not a permanent label. You can move the line in two ways.
Discoveries move it. What the complex toolkit discovers, the complicated toolkit can then pin down, as AWS set out to do with the back-off defect its outage revealed.
Constraints move it too. Timeouts, retry budgets, rate limits, and bulkheads shrink the space of possible interactions, so more of the system’s behavior becomes predictable. The Cynefin wiki defines order in these terms: “Order is constrained to the point where future outcomes are predictable as long as the constraints can be sustained.” AWS’s remediation added such constraints. It “immediately disabled the scaling activities that triggered this event.” It also deployed “additional network configuration that protects potentially impacted networking devices even in the face of a similar congestion event.”
Cook warns against “end-of-the-chain measures,” remedies that block the last step before an accident. “Instead of increasing safety, post-accident remedies usually increase the coupling and complexity of the system.” Constrain the mechanism an outage revealed, not the exact chain of events from last time, and keep each constraint simple enough to reason about.
The line also moves back. A system that changes can turn a settled question into an open one again, and one of Cook’s observations gives the reason: “Change introduces new forms of failure.”
Read Cook’s Paper
If this distinction is useful to you, read the paper that shaped my thinking about it: “How Complex Systems Fail.” Cook wrote it in 1998; the version in circulation dates from April 2000. Its eighteen numbered observations fill four pages and take about fifteen minutes to read. Cook wrote with medicine and other hazardous fields in mind, but the observations read like a description of a production fleet. One of them puts the complex half of this article in a sentence. “Safety is an emergent property of systems; it does not reside in a person, device or department of an organization or system.” The paper is likely to change how you read postmortems, and how you write them.
Checklist
- Ask the question for each part of the problem, and ask it again once the code runs in production. When you cannot tell, split the problem and keep the interactions as questions of their own.
- For what is knowable in advance, specify it, test it, review it, and, where the stakes justify it, prove it. Do not reach for a canary where a unit test will do.
- For what is knowable only in retrospect, define normal before you ship, roll out gradually behind a kill switch, and experiment before an incident does it for you.
- Keep observability and kill switches working during failures, on paths that do not share the failure.
- When an incident is spreading, stabilize first. Analysis and experiments come afterward.
- Triage postmortem findings from where the team stood before the outage. Strengthen the defenses for what was knowable then, and improve detection and recovery for the rest.
- Move the line on purpose. Turn discoveries into fixes and tests, and add simple constraints, such as timeouts and retry budgets, that make behavior predictable.
Further Reading
- How Complex Systems Fail (Richard I. Cook, Cognitive Technologies Laboratory, University of Chicago, Revision D, 21 April 2000), also available as the original PDF
- A Leader’s Framework for Decision Making (David J. Snowden and Mary E. Boone, Harvard Business Review, November 2007)
- Cynefin Domains (Cynefin.io wiki, last modified 3 August 2022)
- Summary of the AWS Service Event in the Northern Virginia (US-EAST-1) Region (Amazon Web Services, 10 December 2021)
- Use of Formal Methods at Amazon Web Services (Chris Newcombe, Tim Rath, Fan Zhang, Bogdan Munteanu, Marc Brooker, and Michael Deardeuff, Amazon.com, 29 September 2014)
- Principles of Chaos Engineering (principlesofchaos.org, last updated March 2019)
- CanaryRelease (Danilo Sato, martinfowler.com, 25 June 2014)