Idempotency in APIs: Making Retries Safe with HTTP Semantics and Idempotency Keys

By Sergey Nosov

6 October 2026

A customer presses Pay. The request reaches the server, the card is charged, and the response never makes it back: the connection drops, the client times out, or a proxy gives up first. The client now holds the worst kind of error, one that does not say whether the work happened. If it retries, it may charge the card twice. If it gives up, it may lose a paid order.

Idempotency is what lets the client retry without guessing. This article covers what the word means in HTTP, why retries make it mandatory, and how to design for it. Most of it is about idempotency keys. The usual sketch of one is a response cache, which hides six decisions and gets most of them wrong. It ends with an ASP.NET Core middleware and a JavaScript client that I tested against each other on .NET 10 and Node.js 22.

What Idempotency Means

The word comes from algebra. Benjamin Peirce introduced it in 1870, in Linear Associative Algebra. An expression, he wrote, “may be called idempotent” when, “raised to a square or higher power, it gives itself as the result.” Programmers write the same property as f(f(x)) = f(x): applying the operation a second time changes nothing. Taking an absolute value is idempotent; adding to a balance is not.

I like to think of an elevator call button. Pressing it ten times does not bring the elevator sooner or summon ten elevators. The first press registers the call, and the system absorbs the rest.

For APIs, the definition that matters is HTTP’s. RFC 9110 is the HTTP semantics standard, published in June 2022. It defines a method as idempotent “if the intended effect on the server of multiple identical requests with that method is the same as the effect for a single such request.” Two details in that sentence settle many arguments.

The effect matters, not the response. Delete an order twice, and the second response may be 404 Not Found instead of 204 No Content. The order is gone either way, so the method is still idempotent. RFC 9110 says as much when it explains retries: the client “knows that repeating the request will have the same intended effect, even if the original request succeeded, though the response might differ.”

The intended effect matters, not every side effect. The same section adds that “a server is free to log each request separately, retain a revision control history, or implement other non-idempotent side effects for each idempotent request.” A log line per request does not break idempotency. A second charge does.

What HTTP Already Promises

RFC 9110 sorts its own methods into three groups, and RFC 5789 adds a fourth case for PATCH:

Whether a PATCH is idempotent depends on the patch format. A JSON Merge Patch adds or replaces the members it names and removes the ones it sets to null, so applying the same document twice gives the same result. A JSON Patch add operation whose path ends in /- appends to an array, so applying it twice appends twice.

Nothing enforces these promises. A PUT handler that adds the submitted amount to a balance instead of storing it breaks one, and every client and library that trusts the method may repeat the request.

Idempotency is not concurrency control, either. A delayed PUT can overwrite a change that another client made in the meantime, and a delayed DELETE can remove a resource that someone has since re-created. Conditional requests close that gap. RFC 9110 offers If-Match “to prevent accidental overwrites when multiple user agents might be acting in parallel on the same resource.” For a PUT that should only create, If-None-Match: * keeps it from overwriting a resource that the client believed did not exist.

Why Retries Make Idempotency Mandatory

In a 2017 post on Stripe’s engineering blog, Brandur Leach lists three ways a call between two machines can fail. The connection can fail; the call can fail midway, “leaving the work in limbo”; or the call can succeed and the connection break before the client hears about it. The first failure is usually clear, because nothing reached the server. The other two look the same to the client, and the only way to learn the outcome is to ask again. Asking again is a retry.

Retries happen whether you design for them or not:

RFC 9110 draws the line for automatic retries. “A client SHOULD NOT automatically retry a request with a non-idempotent method unless it has some means to know that the request semantics are actually idempotent, regardless of the method, or some means to detect that the original request was never applied.” There are two ways to give clients that knowledge. Make the operation idempotent by design, or accept an idempotency key.

Idempotent by Design

The cheapest idempotency is the kind you do not have to build. Before I reach for keys, I look for a way to make the operation itself repeatable:

My Software Development Principles series treats idempotency as a principle of its own, with these strategies, the indicators of proper application, and the common violations.

Some operations resist all of these: the server assigns the payment ID, or the work calls a system that cannot recognize a repeat, such as an email provider. Those need an idempotency key. The key protects your API from its own clients’ retries, though; it cannot make a call to a system without deduplication safe to repeat. When such a call’s outcome is unknown, Leach’s advice is to “take the conservative route and fail the operation.”

How Idempotency Keys Work

An idempotency key is a unique value that the client generates for one logical operation and sends with every attempt at it, usually in an Idempotency-Key request header. The server records the key, a fingerprint of the request, and the outcome. When an attempt arrives with a key the server has already seen, it returns the recorded outcome instead of doing the work again.

Stripe’s API reference describes its implementation precisely. “Stripe’s idempotency works by saving the resulting status code and body of the first request made for any given idempotency key, regardless of whether it succeeds or fails.” Stripe accepts keys of up to 255 characters, suggests version 4 UUIDs, and marks a replayed response with an Idempotent-Replayed: true header.

The header has not become a standard. A working-group draft at the IETF, The Idempotency-Key HTTP Header Field, set out to define it, and it credits the API documentation of PayPal and Stripe as its inspiration. Its abstract says the header “can be used to make non-idempotent HTTP methods such as POST or PATCH fault-tolerant.” The latest revision, dated 15 October 2025, expired on 18 April 2026, and as of October 2026 the IETF Datatracker lists the draft as expired rather than as an RFC. It is still the best single description of the pattern, and its list of implementations shows how much the names vary. Stripe, Adyen, and Dwolla use Idempotency-Key. PayPal uses PayPal-Request-Id, and Square puts an idempotency_key property in the request body. Amazon EC2 calls the same thing a client token, “a unique, case-sensitive string of up to 64 ASCII characters.”

What a Response Cache Gets Wrong

The sketch fits in one sentence: look up the key, return the stored response if there is one, and otherwise run the request and store its response. Each step of that sentence hides a decision.

Claim the Key before Doing the Work

Looking up and then running is a check-then-act race. Two attempts with the same key that arrive together, such as a retry that overtakes a slow original, both find nothing and both run. Claim the key atomically before doing any work instead. In Redis, SET with the NX option will “only set the key if it does not already exist.” In SQL, an insert into a table with a unique index on the key does the same. The request that wins the claim does the work. The draft says what the loser should hear: “If the request is retried, while the original request is still being processed, the resource SHOULD reply with an HTTP 409 status code.” A 409 Conflict tells the client to wait and ask again.

Give the claim a lease. If the server crashes in the middle of a request, a claim without one blocks every retry until the record expires. Leach’s Postgres implementation of Stripe-style keys keeps a locked_at column and takes over a lock only when it “has expired because the original request was long enough ago.” A lease is not a lock, though. A slow request can outlive its lease while a retry takes over. Give each claim an owner token, make completing or releasing it conditional on that token, and keep the downstream keys described below stable.

Fingerprint the Request

A key promises that the request is the same, and clients break that promise by accident. A bug reuses a key, or a user edits the form and the client retries with the old key. A cache that looks only at the key returns the old response to the new request, and the client believes the new one succeeded. Store a fingerprint of each request and refuse a mismatch. It should cover the method, the URL, the content type, a hash of the body, and any other header that changes the meaning, such as an API version. The draft asks for 422 Unprocessable Content. Stripe’s idempotency layer “compares incoming parameters to those of the original request and errors if they’re not the same,” and EC2 fails such a retry with an IdempotentParameterMismatch error. In the Amazon Builders’ Library, Malcolm Featonby gives the reasoning: “We find that it is safest to assume that the customer intended a different outcome, and that this might not be the same request.”

A hash of the raw body treats reformatted JSON as a different request. I accept that: it errs toward refusing, and refusing is the safe direction.

Scope and Validate the Key

If the server looks keys up globally, two clients that choose the same key collide, which happens easily when keys come from order numbers rather than random UUIDs. Worse, a client that learns another client’s key can read that client’s stored response. The draft’s security considerations warn that with low-entropy keys, “attackers MAY determine other keys and use them to fetch existing idempotent cache entries, belonging to other clients.” Its remedy is a composite lookup key that combines the client’s key “with other client specific attributes known only to the resource.” Use an identifier that is unique and stable, such as the authenticated user’s ID, not a display name. Leach’s schema makes the key unique per user. Stripe’s v2 API treats a request as a replay only if it happens “in the scope of the same account or sandbox.”

Validate the key before it reaches any lookup. Publish a format, such as a UUID, enforce a maximum length, and reject anything else with 400 Bad Request. The draft puts it plainly: “Always validate the key as per its published specification before processing any request.” Stripe adds a privacy rule worth copying: “Avoid using sensitive data (for example, email addresses or personal identifiers) as idempotency keys.”

Decide What a Failure Leaves Behind

This is the hardest decision. When the first attempt fails, should a retry with the same key replay the failure or run the operation again? Stripe has shipped both answers. Its v1 API “always returns the previously-saved response of the first API request, even if it was an error.” Its advice for a 500 follows from that: “The client can retry the request with a new idempotency key, but we advise against it because the original key may have produced side effects. You should treat the result of a 500 request as indeterminate.” Its v2 API instead “attempts to retry any failed requests without producing side effects” and returns an updated response.

Running a failed request again is safe only if the failure left nothing behind. For work inside your own database, that means one transaction that records the key and makes the change together. Featonby states the requirement: “the process that combines recording the idempotent token and all mutating operations related to servicing the request must meet the properties for an atomic, consistent, isolated, and durable (ACID) operation.” Without that, a crash can record the key without the change, or make the change without recording the key.

Calls to other systems cannot join your transaction. Leach puts it bluntly: “once we make our first foreign state mutation, we’re committed one way or another.” His design commits local work in what he calls atomic phases between external calls and resumes from the last completed phase on a retry. When his example charges a card through Stripe, it sends a key derived from its own idempotency record. That key stays the same on every retry and, unlike the client’s value, is unique across all of its accounts.

A failure before the work starts is different. Stripe does not save a result when “incoming parameters fail validation, or the request conflicts with another request that’s executing concurrently,” and adds, “You can retry these requests.” My rule: store every outcome by default, failures included, and release the key only when the handler knows that a retry cannot repeat anything. Before the work starts, that is easy to know. After a failure, it takes two things. The attempt’s own writes must have rolled back, and every call to another service must have carried a key derived from this request’s key that the service still remembers.

Decide What a Duplicate Gets Back

Replaying the stored status code and body is the simplest choice, and it is what Stripe’s v1 API does. It is not the only one. Featonby describes delivering “a semantically equivalent response in every case for the same unique request identifier for some interval.” A retried EC2 launch can report the instance as running rather than pending. The EC2 documentation warns that “the result might contain updated information, such as the current creation status.” Either way, the retry’s answer has to work as the answer to the original request. Featonby’s counterexample is a retry that answers ResourceAlreadyExists: nothing happened twice, but the client cannot tell whether its own request created the resource.

Decide How Long a Key Lives

A key that expires too soon turns a late retry into a new operation. Stripe’s v1 API treats requests as replays only within 24 hours of each other, and its v2 API within 30 days. For EC2 instances, Featonby describes a retention period of “the lifetime of the resource, plus an interval after which it is reasonable to assume that any late arriving requests would either have arrived or would no longer be valid.” There is no universal number. Choose a window longer than the longest path a retry can take, including mobile clients that come back online and jobs that wait in a queue, and publish it. The draft says the same: “The resource SHOULD define such expiration policy and publish it in the documentation.”

A Server-Side Implementation in ASP.NET Core

The middleware below puts these decisions into code. Together, the three blocks in this section form one .NET 10 file-based app: save them in order as payments.cs and run dotnet run payments.cs. It is a teaching sample, and the list after the code says what production needs on top of it. The first block is the application. Its payment endpoint waits 2 seconds to stand in for a slow payment provider:

#:sdk Microsoft.NET.Sdk.Web
#:property PublishAot=false

using System.Collections.Concurrent;
using System.Security.Claims;
using System.Security.Cryptography;

var builder = WebApplication.CreateBuilder(args);
builder.Services.AddSingleton<IIdempotencyStore, InMemoryStore>();
var app = builder.Build();

app.UseMiddleware<IdempotencyMiddleware>();

var payments = new ConcurrentDictionary<string, Payment>();

app.MapPost("/payments", async (PaymentRequest request) =>
{
    await Task.Delay(2000); // stands in for a slow payment provider
    var payment = new Payment(Guid.NewGuid().ToString(), request.Amount);
    payments[payment.Id] = payment;
    return Results.Created($"/payments/{payment.Id}", payment);
});

app.MapGet("/payments", () => payments.Values);

app.Run();

record PaymentRequest(decimal Amount, string Currency);
record Payment(string Id, decimal Amount);

The second line turns off native AOT, which file-based apps enable by default. Under native AOT, minimal APIs need source-generated JSON serializers, and this sample keeps to reflection for brevity.

The store is the part to replace in production. This one keeps records in memory for a single process. Its claim is ConcurrentDictionary.GetOrAdd, which returns the existing record when another request got there first. A shared store needs the same atomic claim, plus an expiry for completed records and a lease for claims:

record StoredResponse(
    int Status, string? ContentType, string? Location, byte[] Body);

// Response is null while the first request is still running.
record IdempotencyRecord(string Fingerprint, StoredResponse? Response);

interface IIdempotencyStore
{
    // Atomically claims the key. Returns null when this request won the
    // claim, or the record that another request stored first.
    Task<IdempotencyRecord?> TryClaimAsync(string key, string fingerprint);
    Task CompleteAsync(string key, IdempotencyRecord record);
    Task ReleaseAsync(string key);
}

// For one process only. A shared store needs the same atomic claim
// (Redis SET NX, or an INSERT against a unique key), an expiry for
// completed records, and a lease with an owner token for each claim.
class InMemoryStore : IIdempotencyStore
{
    private readonly ConcurrentDictionary<string, IdempotencyRecord> records =
        new();

    public Task<IdempotencyRecord?> TryClaimAsync(
        string key, string fingerprint)
    {
        var claim = new IdempotencyRecord(fingerprint, null);
        var current = records.GetOrAdd(key, claim);
        return Task.FromResult(
            ReferenceEquals(current, claim) ? null : current);
    }

    public Task CompleteAsync(string key, IdempotencyRecord record)
    {
        records[key] = record;
        return Task.CompletedTask;
    }

    public Task ReleaseAsync(string key)
    {
        records.TryRemove(key, out _);
        return Task.CompletedTask;
    }
}

The middleware does the rest:

class IdempotencyMiddleware(RequestDelegate next, IIdempotencyStore store)
{
    public async Task InvokeAsync(HttpContext context)
    {
        // PUT and DELETE are idempotent already; POST and PATCH are not.
        var request = context.Request;
        var needsKey = HttpMethods.IsPost(request.Method) ||
            HttpMethods.IsPatch(request.Method);
        if (!needsKey ||
            !request.Headers.TryGetValue("Idempotency-Key", out var header))
        {
            await next(context);
            return;
        }

        // One published format: a hyphenated UUID, sent bare (as Stripe
        // does) or in one pair of quotes (as the IETF draft specifies).
        var value = header.ToString();
        if (value.Length == 38 && value[0] == '"' && value[^1] == '"')
            value = value[1..^1];
        if (!Guid.TryParseExact(value, "D", out var id))
        {
            await Problem(context, 400, "Idempotency-Key must be a UUID.");
            return;
        }

        // Scope the key to the caller's user ID, so that no client can
        // replay another client's response. Authenticate before this runs.
        var caller = context.User.FindFirst(ClaimTypes.NameIdentifier)?.Value
            ?? "anonymous";
        var key = $"{caller}:{id}";
        var fingerprint = await FingerprintAsync(request);

        var existing = await store.TryClaimAsync(key, fingerprint);
        if (existing is not null)
        {
            if (existing.Fingerprint != fingerprint)
                await Problem(context, 422,
                    "Idempotency-Key was used for a different request.");
            else if (existing.Response is null)
                await Problem(context, 409,
                    "A request with this Idempotency-Key is in progress.");
            else
                await ReplayAsync(context, existing.Response);
            return;
        }

        // A caller that hangs up will retry and expect this result, so
        // finish the work and record the response even if it is gone.
        context.RequestAborted = CancellationToken.None;

        var response = context.Response;
        var originalBody = response.Body;
        using var buffer = new MemoryStream();
        response.Body = buffer;
        try
        {
            await next(context);
        }
        catch
        {
            // The outcome is unknown, so every retry gets this 500 rather
            // than a second run of work that may already have happened.
            await store.CompleteAsync(key, new IdempotencyRecord(
                fingerprint, new StoredResponse(500, null, null, [])));
            throw;
        }
        finally
        {
            response.Body = originalBody;
        }

        // A handler answers 429 or 503 only when it did nothing, so a
        // retry may run. Every other answer, errors included, is final.
        if (response.StatusCode is 429 or 503)
        {
            await store.ReleaseAsync(key);
        }
        else
        {
            var stored = new StoredResponse(response.StatusCode,
                response.ContentType, response.Headers.Location,
                buffer.ToArray());
            await store.CompleteAsync(key,
                new IdempotencyRecord(fingerprint, stored));
        }

        buffer.Position = 0;
        await buffer.CopyToAsync(originalBody);
    }

    // Method, URL, content type, and a hash of the body: the same key
    // must always come with the same request.
    static async Task<string> FingerprintAsync(HttpRequest request)
    {
        request.EnableBuffering();
        var hash = await SHA256.HashDataAsync(request.Body);
        request.Body.Position = 0;
        var url = $"{request.PathBase}{request.Path}{request.QueryString}";
        return $"{request.Method} {url} {request.ContentType} " +
            Convert.ToHexString(hash);
    }

    static async Task ReplayAsync(HttpContext context, StoredResponse stored)
    {
        var response = context.Response;
        response.StatusCode = stored.Status;
        response.ContentType = stored.ContentType;
        if (stored.Location is not null)
            response.Headers.Location = stored.Location;
        response.Headers["Idempotent-Replayed"] = "true";
        await response.Body.WriteAsync(stored.Body);
    }

    static Task Problem(HttpContext context, int status, string title) =>
        Results.Problem(statusCode: status, title: title)
            .ExecuteAsync(context);
}

The sample leaves out production plumbing:

Two more cautions apply. The middleware must run inside response compression, or it stores compressed bytes without their Content-Encoding header. And the handler’s own calls still need timeouts, because a hang-up no longer cancels them. Count replays, 409s, and 422s. A client that keeps receiving 422 has a bug, and a stream of 409s means that retries are outrunning the work.

A Client That Retries Safely

The client half is shorter, and its most important rule is about the key. Create one key per operation and reuse it for every attempt; a new key on each retry would make every retry a new payment. The function takes the key as a parameter rather than creating it, because the operation outlives the call. A caller that holds the key can ask about the payment again after the function gives up. Stripe suggests deriving a key “from a user-attached object, like the ID of a shopping cart,” which “provides a relatively straightforward way to protect against double submissions.” With a server like the one above, which accepts only UUIDs, create a UUID along with the cart and send that. Leach suggests a hidden form field, filled in when the page is rendered, so that a second click on Pay sends the same key.

const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));

// 409: the first attempt is still running. 429 and 503: the server did
// nothing yet. Other 5xx: worth another try, since a replay is final.
const retryable = (status) =>
  status === 409 || status === 429 || status >= 500;

// The caller creates the key once per operation and keeps it until the
// outcome is known, so that a later call can ask about the same payment.
export async function createPayment(baseUrl, payment, key, {
  attempts = 5,
  timeoutMs = 10_000,
} = {}) {
  const body = JSON.stringify(payment); // the same bytes on every attempt
  for (let attempt = 1; ; attempt++) {
    let response;
    try {
      response = await fetch(`${baseUrl}/payments`, {
        method: "POST",
        headers: {
          "Content-Type": "application/json",
          "Idempotency-Key": key,
        },
        body,
        signal: AbortSignal.timeout(timeoutMs),
      });
      if (response.ok) return await response.json();
    } catch (error) {
      // A timeout or a network failure, even while reading the body,
      // leaves the outcome unknown: try again with the same key.
      if (error instanceof SyntaxError || attempt >= attempts) throw error;
      response = undefined;
    }

    // A replayed response is the final answer for this key.
    const final = response &&
      (response.headers.has("Idempotent-Replayed") ||
        !retryable(response.status));
    await response?.body?.cancel();
    if (final) throw new Error(`Payment ended with HTTP ${response.status}`);
    if (attempt >= attempts) {
      throw new Error(`No result after ${attempts} attempts`);
    }
    // Exponential backoff with full jitter: 0-1 s, 0-2 s, 0-4 s...
    await sleep(Math.random() * 1000 * 2 ** (attempt - 1));
  }
}

The rest of the function follows from the earlier sections:

The crypto.randomUUID() function returns a version 4 UUID made by “a cryptographically secure random number generator.” RFC 9562 gives a version 4 UUID 122 random bits, so a collision is possible but negligible. My article on combinatorics in software engineering works through the birthday bound behind that claim. In browsers, the function is available only in secure contexts, such as pages served over HTTPS.

I ran the two halves against each other with a 1-second client timeout and the 2-second handler, on .NET 10. Every run went the same way. The first attempt timed out while the server kept working. Retries that arrived while it was still running got 409. The first retry after it finished received the stored 201, marked as replayed, with the same payment ID. Each run created exactly one payment. In the other tests, a concurrent duplicate got 409, and a malformed key got 400. A changed body, query string, or content type under the same key got 422. Two users who sent the same key got separate payments, and a failed request was replayed rather than run again.

Checklist

Further Reading