Little’s Law: Why Starting Less Finishes Sooner
By Sergey Nosov
8 October 2026
In June 1878, the publisher Henry Holt asked William James whether he could deliver a psychology textbook in a year. James declined: “My other engagements and my health both forbid the attempt to execute the work rapidly.” He offered two years instead, “say the fall of 1880,” and signed the contract that month. On 1 January 1886, seven and a half years later, he wrote a letter to his friend, the philosopher Carl Stumpf:
I still try to write a little psychology, but it is exceedingly slow work. No sooner do I get interested than bang! goes my sleep, and I have to stop a week or ten days, during which my ideas get all cold again. Nothing so fatiguing as the eternal hanging on of an uncompleted task.
The book, The Principles of Psychology, came out in 1890 in two volumes, twelve years after James signed the contract.
Developers know the feeling on a smaller scale: the branch from last week, the review you meant to get to, the ticket that has said In Progress for a month. Each one costs a little attention every day, and none of them has delivered anything yet. James described what an unfinished task costs the person carrying it. Little’s Law, proved in 1961, describes what it costs the work: time. It fits on a napkin, and it explains why a team that starts less finishes sooner.
This article explains the law and the classroom dare behind its proof. It then reads the law backward and shows what a limit on work in progress can and cannot do. It adds a second result about busy reviewers, a simulation that checks the numbers, four habits, and a checklist.
The Law
Little’s Law ties three averages together:
L = λW
- L is the average number of items inside the system: the work in progress.
- λ (lambda) is the average rate at which items arrive. When the system is stable, with work not piling up inside, it is also the rate at which items leave: the throughput.
- W is the average time an item spends inside: the lead time.
A coffee shop that serves two customers a minute, where each customer stays five minutes, holds ten customers on average. The units have to agree: if λ counts items per day, W is in days.
The system can be anything with a boundary that items cross on the way in and on the way out. A coffee shop qualifies, and so do a hospital ward, a sprint board, and the open pull requests in a repository. In a 2008 chapter on the law, Little and Stephen Graves list what it never asks about. It does not ask how many servers there are, how long service takes, how arrivals are spaced, or in what order items are served.
Why does it hold? The proof fits in a paragraph and a picture. Take a shop that opens empty and closes empty, and draw the number of customers inside over the day. The line is a staircase that steps up at each arrival and down at each departure:
The area under the staircase is measured in customer-hours. Divide it by the hours the shop was open, and you get L. Divide the same area by the number of customers, and you get W. The number of customers divided by the hours is λ, so L = λW is plain arithmetic. In the picture, six customers spend twenty customer-hours in a shop open for thirteen hours. So L is 20/13, W is 20/6, and λ is 6/13. Multiply λ by W, and you get 20/13 again.
Little gave the physical reason for the double duty in his 2011 retrospective. In his words, “at the same time that a customer is standing in line and so can be counted, he or she is also accumulating minutes waiting.”
The same arithmetic sizes connection pools and thread pools, and the entry on Little’s Law in my Software Development Principles series covers that side. This article stays with the work people do.
A Proof That Should Not Have Been Too Hard
The formula is older than its proof. In his 1958 book Queues, Inventories and Maintenance, Philip Morse observed that it held across a wide variety of queues, and he left his readers a dare:
Those readers who would like to experience for themselves the slipperiness of fundamental concepts in this field and the intractability of really general theorems, might try their hand at showing under what circumstances this simple relationship between L and W does not hold.
John Little had been Morse’s doctoral student. From 1957 to 1962, he taught operations research at the Case Institute of Technology in Cleveland. In the 2011 retrospective, he recalls talking with students after class when one of them, Sid Hess, asked, “How hard would it be to prove it in general?” Little answered, “I guess it shouldn’t be too hard.” His own verdict on that answer: “Famous last words.” Hess replied, “Then you should do it!”
He did. “A Proof for the Queuing Formula: L = λW” appeared in Operations Research in 1961. Its key assumption was that the system is stationary: its statistical behavior does not drift over time. Later work removed even that. In 1974, Shaler Stidham published a proof that needs only the long-run averages to exist and be finite, under the title “A Last Word on L = λW.” Writing in 2002, Stidham added: “Of course it was not a last word.”
Little then “retired from queuing,” as he put it, and helped found the field of marketing science. He died on 27 September 2024, at age ninety-six.
Read It Backward
Turn the law around, and it says what lead time is made of:
W = L / λ
In words, lead time is work in progress divided by throughput. Suppose a team merges four pull requests a day and keeps twelve open. On average, a pull request takes three days from opening to merge. Keep six open at the same pace, and it takes a day and a half. Nobody typed faster, and nobody worked longer. The only thing that changed is how much was open at once.
That runs against instinct. When work feels slow, the urge is to start more, so that something is always moving. Little’s Law says the opposite: at the same throughput, every item you add lengthens the average lead time. To finish sooner, start less.
The practical tool is a limit: start a new item only when one finishes. Manufacturing calls this rule CONWIP, short for constant work in process, and Little and Graves call it “a very effective control policy.” The law also gives a starting point for the limit: throughput times the lead time you want. Take a team that merges four pull requests a day and wants an average of one day from opening to merge. It should aim for about four open and check that throughput holds.
Three Caveats
The waiting moves; it does not vanish. Work you have not started still waits, in the backlog instead of in progress. Draw the boundary around the backlog too, and the law covers the whole trip from request to merge: all unfinished work divided by throughput. If a limit only moves the same waiting upstream, whoever asked for the change waits just as long, on average.
The limit still pays. An item in the backlog holds no half-written code, needs no rebasing, and can be reordered or dropped cheaply. An item in progress ages, and someone has to keep it in their head. Starting less makes started work finish sooner. Requests finish sooner, on average, only when total unfinished work falls or throughput rises.
Too tight a limit costs throughput. The law promises only that the three numbers move together. While items wait, removing one shortens the average wait. Once nothing waits, the average lead time cannot fall below the time the work itself takes, and every item you remove comes out of throughput. Little and Graves state the coupling plainly: “any increase in output requires either an increase in work-in-process inventory or a reduction in cycle time or both.” Set the limit low enough that started work rarely waits and high enough that nobody runs out of useful work. When a limit leaves you without a next item, help finish something already started.
A queue costs more than the law counts. Little called the extra cost “queue-related overhead.” In his examples, someone has to keep track of every waiting item, and a computer switches more often between threads. On a software team, the overhead is rebasing stale branches, rereading code you have forgotten, and answering questions about work you started weeks ago. Cutting the queue cuts that overhead, so a limit can raise throughput as well as shorten lead time. The law cannot promise that gain, but it can measure it. James knew the overhead too: while his book waited, his ideas would “get all cold again.”
Busy Is Not Fast
Little’s Law relates the averages, but it does not predict them. How long an item waits before anyone starts on it depends on a second lever: how busy the people serving the queue are. The simplest model in queueing theory, known as M/M/1, has one server working first come, first served, and items that arrive at random, as a Poisson process. Their sizes are random too, drawn from an exponential distribution: most items are small, and a few are large. As long as the server is less than 100 percent busy, its average wait has a closed form:
wait = ρ / (1 − ρ) × S
Here ρ (rho) is the utilization, the fraction of time the server is busy, and S is the average amount of work per item. The wait is the time before work starts; the lead time W is the wait plus S. The first factor does the damage:
- At 50 percent busy, an item waits about as long as the work takes.
- At 80 percent, it waits four times as long.
- At 90 percent, nine times as long.
- At 95 percent, nineteen times as long.
Run a reviewer at 90 percent, and a one-hour review waits nine hours, on average, before anyone opens it. Little’s Law turns that wait into a line you can see. At nine pull requests every ten hours, each waiting nine hours, the line holds about eight on average: 0.9 × 9 = 8.1. Slack is not waste; it is what keeps the queue short. It means spare capacity, not idle people: reviews that fill half a reviewer’s day leave the other half for their own work.
The model is kinder than reality, though. It assumes that a review starts the moment the reviewer is free, and meetings and focused work only add to the wait.
Its exact multiples also depend on how random the work is. Real arrivals are not perfectly random, and real work sizes vary more or less than the model assumes. Kingman’s formula, named for J. F. C. Kingman and his 1961 paper on the single-server queue in heavy traffic, generalizes the result:
wait ≈ ρ / (1 − ρ)
× (ca² + cs²) / 2
× S
Here ca and cs are the coefficients of variation of the gaps between arrivals and of the work sizes: each one’s standard deviation divided by its mean. In the M/M/1 model, both equal one, and the middle factor drops out. The formula is an approximation for independent arrivals and work sizes, and it is most accurate when the server is nearly saturated, which is exactly the case that hurts. It names three levers:
- Utilization sets the shape. The wait grows slowly at first and then explodes as the server approaches 100 percent busy.
- Variability scales it. Random arrivals of uniform work halve the wait; work whose size varies wildly multiplies it.
- Size sets the unit. Every multiple is a multiple of S, so at the same utilization, a queue of small items moves in hours where a queue of big ones moves in days.
Check the Numbers Yourself
Formulas are easy to nod at and hard to believe, so here is a simulation of one reviewer and a million pull requests. It opens pull requests at random times, gives each a random review time averaging one hour, reviews them in order, and checks Little’s Law along the way. It was checked on Node.js 22 and runs in well under a second at the settings below.
// queue.js: one reviewer, pull requests reviewed first come, first served.
// Run: node queue.js 0.9 (the reviewer is busy 90 percent of the time)
// node queue.js 0.9 fixed (every review takes exactly one hour)
const busy = Number(process.argv[2] ?? 0.9);
const fixed = process.argv[3] === "fixed";
if (!(busy > 0 && busy < 1)) throw new Error("busy must be between 0 and 1");
const n = 1_000_000; // pull requests
const exponential = (mean) => -mean * Math.log(1 - Math.random());
const opened = new Float64Array(n);
const merged = new Float64Array(n);
let now = 0, free = 0, waited = 0, inside = 0;
for (let i = 0; i < n; i++) {
// The next pull request opens. Its review starts when the reviewer
// is free and takes one hour on average.
now += exponential(1 / busy);
const start = Math.max(now, free);
free = start + (fixed ? 1 : exponential(1));
opened[i] = now;
merged[i] = free;
waited += start - now;
inside += free - now;
}
// Count the open pull requests once an hour, the way a board shows them.
let samples = 0, counted = 0;
for (let t = 0, a = 0, d = 0; t < free; t++, samples++) {
while (a < n && opened[a] <= t) a++;
while (d < n && merged[d] <= t) d++;
counted += a - d;
}
// L: average number open. lambda: merged per hour.
// W: average hours from open to merge.
const L = counted / samples;
const lambda = n / free;
const W = inside / n;
console.log(`wait for review: ${(waited / n).toFixed(1)} hours`);
console.log(`L = ${L.toFixed(2)}, lambda * W = ${(lambda * W).toFixed(2)}`);
A run at 0.9 prints something like this:
wait for review: 9.1 hours
L = 9.04, lambda * W = 9.04
Run it at 0.5, 0.8, and 0.9, and the wait comes out close to one, four, and nine hours. At 0.95, it averages nineteen hours, but a single run can land an hour or more to either side: the busier the server, the noisier the queue. Add fixed to make every review take exactly one hour, and the wait at 0.9 drops to about four and a half hours. That is the factor of one-half in Kingman’s formula, which is exact when arrivals are random.
The second line compares two ways of measuring the same history: λ × W from the timestamps, and L from hourly counts, the way a board shows it. Both count every open pull request, including the one under review, so at 0.9 they land near nine: about eight in line and, most of the time, one in review. Over a run that starts and ends empty, they would agree exactly if someone watched continuously. Counting once an hour leaves a gap of about a hundredth. The wait changes with every setting; the law does not.
The model leaves out rework, second reviewers, and priorities, which change the numbers and the exact curve. The warning survives: when work varies, waits climb steeply as the server nears 100 percent busy.
Measure Your Own Queue
The same arithmetic works on a real board, as long as you count consistently. Draw one boundary where your question is: from a pull request’s opening to its merge, from In Progress to Done, or from request to production. Then measure all three numbers against it over the same window:
- L: count the items inside the boundary at the same time every day, and average the counts.
- λ: count the items that left during the window, and divide by its length.
- W: average the time inside for the items that left.
The shop in the proof starts and ends empty. Over such a window, Little’s 2011 paper shows, the law is exact, with no assumption of steady state. A team’s board is rarely empty. For that case, Little and Graves name two conditions that make the law a good approximation over a long enough window. The work in progress must be “roughly the same at the beginning and end of the time interval,” and its average age must be “neither growing nor declining.” Underneath both sits conservation: everything that enters eventually leaves.
The second condition is the one boards break. Little and Graves illustrate it with a doll store that keeps one hundred dolls in stock and sells about two a week. Dolls with mauve hats almost never sell, and each sold doll has a one-in-ten chance of being replaced by one. Over the years, the mauve-hat dolls pile up and age, while the rest turn over faster and faster. The naive estimate, one hundred dolls divided by two a week, or fifty weeks, overstates how long the sold dolls actually stayed. Every stale ticket and abandoned pull request on a board is a mauve-hat doll. Finish it or close it, count a closure as an exit, and the averages start describing the work again.
Two more uses follow from the same arithmetic. Little’s summary covers the first: “If you know two of {L, λ, W}, you can quickly calculate the third.” A count of open pull requests and a merge rate give the average time to merge without a single timestamp. And if a dashboard shows all three and they disagree badly, suspect the measurement: mismatched boundaries, uncounted exits, or a window that breaks the two conditions above.
Finally, the law looks backward, and it speaks only of averages. Little puts it bluntly: “we are in the measurement business, not the forecasting business.” Last month’s averages describe last month. They make a reasonable baseline for planning, not a promise about any single item. An average of three days can also hide a few pull requests that took three weeks, so watch the slowest items separately.
Finish Before You Start
On a team, the arithmetic turns into four habits:
- Review before you start. An open review is someone else’s work in progress. Finishing it shortens their lead time; starting something new only adds to yours. Google’s code review guide, which calls a pull request a CL (short for changelist), recommends reviewing between focused tasks. It also names the cost of waiting. In its words, “new features and bug fixes for the rest of the team are delayed by days, weeks, or months as each CL waits for review and re-review.” My article on the anatomy of a good code review covers the rest of the reviewer’s side.
- Set yourself a limit, such as one item in progress and one in review. At the limit, finish or unblock something before you pick up the next one.
- Cut smaller slices. A small pull request flows through review; a big one is several small ones waiting together. Google’s guide to small CLs gives the reviewer’s side. “It’s easier for a reviewer to find five minutes several times to review small CLs than to set aside a 30 minute block to review one large CL.” Kingman’s formula gives the queue’s side: smaller, more uniform work waits less at the same utilization. When you compare before and after, remember that splitting a change multiplies the pull request count without delivering more.
- Keep the board honest. If the In Progress column shows five items with your name on them, you have five. If that is your usual load, then at the same pace your items take five times as long, on average, as they would one at a time.
The Kanban community has a slogan for all four: “Stop starting, start finishing.”
Checklist
- Know the law both ways: L = λW, and lead time equals work in progress divided by throughput.
- Start less to finish sooner. Begin with a limit of throughput times the lead time you want, start new work only when something finishes, and check that throughput holds.
- Remember where the waiting goes. A limit moves it to the backlog, where it is cheaper; only less total unfinished work or higher throughput speeds up the whole trip.
- Do not cut so far that people run out of useful work. When you have no next item, help finish someone else’s.
- Keep review load well below review capacity. In the simplest model, a reviewer at 90 percent makes a one-hour review wait nine hours on average.
- Shrink the work and even out its size. Both shorten the queue at the same utilization.
- Measure L, λ, and W against one boundary and one window. Check that they agree, watch the slowest items separately, and read the results as measurements, not forecasts.
- Finish or close stale items, so that the average age of your work in progress stops drifting.
- Start today: count your work in progress now and again at the end of the week, and see which way it moved.
Further Reading
- Little’s Law as Viewed on Its 50th Anniversary (John D. C. Little, Operations Research, May–June 2011)
- Little’s Law (John D. C. Little and Stephen C. Graves, in Building Intuition, Springer, 2008)
- A Proof for the Queuing Formula: L = λW (John D. C. Little, Operations Research, 1961)
- The Single Server Queue in Heavy Traffic (J. F. C. Kingman, Mathematical Proceedings of the Cambridge Philosophical Society, October 1961)
- Speed of Code Reviews and Small CLs (Google Engineering Practices Documentation)
- The Letters of William James, Volume 1 (edited by Henry James, Atlantic Monthly Press, 1920; Project Gutenberg)