Make your Service Bus message handlers idempotent

ยท 5 min read

Azure Service Bus delivers every message at least once, and duplicates can arrive at the same moment. Here's how to make handling the same message twice harmless, even when two copies race.

Azure Service Bus guarantees that every message is delivered at least once. That wording is deliberate: sometimes you get the same message twice.

It happens more often than you'd think. Your handler processes an order, writes to the database, and then crashes, times out or loses its lock before it completes the message. Service Bus sees a message that was never completed and delivers it again. If the handler charges a card or sends an email, the customer gets it twice.

The fix isn't to prevent redelivery. It's to make processing the same message a second time harmless. That property is called idempotency.

How duplicates happen

With the default peek-lock mode, receiving a message locks it rather than removing it. Your handler has until the lock expires to complete it. If the handler throws, the process dies or the lock runs out first, the message becomes available again and its delivery count goes up. After MaxDeliveryCount attempts (10 by default) it moves to the dead-letter queue.

Duplicates can also start at the sender. If a send times out after the broker has already stored the message, the sender retries and the queue now holds two copies.

Neither kind of duplicate waits politely for the first copy to finish:

  • With several instances of your service, or a processor running with MaxConcurrentCalls above 1, two copies of a sender-side duplicate can be picked up within milliseconds of each other.
  • If a handler runs past its lock, Service Bus hands the message to another receiver while the first is still working on it. The processor renews locks for you, up to MaxAutoLockRenewalDuration (5 minutes by default), but a slow dependency or a long pause can still outlast it.

So your handler has to be safe not only when the same message comes back later, but when two copies are being processed at the same moment.

Duplicate detection helps, but only on the way in

Service Bus has a duplicate detection feature. When it's on, the broker drops any message whose MessageId it has already seen within a time window (10 minutes by default, up to 7 days). That stops the sender-side duplicates above, as long as the retry reuses the same MessageId.

It does nothing for redelivery to your receiver. It also has to be turned on when the queue or topic is created. You can't enable it later.

Why "check, then do the work" isn't enough

The obvious approach is to keep a table of processed message IDs and check it first:

if (await db.ProcessedMessages.AnyAsync(m => m.Id == messageId, ct))
    return; // already handled

// ... do the work, then record the message ID ...

This handles a message that comes back a minute later. It fails when two copies arrive together. Both handlers run the check before either has saved anything, both see "not processed", and both do the work. Any check followed by a separate write has this gap, however fast the code is.

The check and the claim have to be one operation that the database decides, so only one copy can win.

Claim the message first, and let the database pick the winner

Insert the processed-message record before doing the work, inside the same transaction, and make the message ID its primary key:

processor.ProcessMessageAsync += async args =>
{
    await using var scope = services.CreateAsyncScope();
    var db = scope.ServiceProvider.GetRequiredService<ShippingDb>();
    var order = args.Message.Body.ToObjectFromJson<OrderPlaced>();
    var ct = args.CancellationToken;

    await using var tx = await db.Database.BeginTransactionAsync(ct);
    try
    {
        // 1. Claim the message. The primary key on Id lets only one transaction hold it.
        db.ProcessedMessages.Add(new ProcessedMessage(args.Message.MessageId, DateTimeOffset.UtcNow));
        await db.SaveChangesAsync(ct);

        // 2. Do the work in the same transaction.
        db.Shipments.Add(Shipment.For(order));
        await db.SaveChangesAsync(ct);

        await tx.CommitAsync(ct);
    }
    catch (DbUpdateException ex) when (IsDuplicateKey(ex))
    {
        // Another copy already claimed this message and did the work.
        await tx.RollbackAsync(ct);
    }

    // 3. Complete only after the transaction has finished.
    await args.CompleteMessageAsync(args.Message, ct);
};

static bool IsDuplicateKey(DbUpdateException ex) =>
    ex.InnerException is SqlException { Number: 2627 or 2601 }; // SQL Server

Here's what happens when two copies arrive at the same instant:

  1. Both handlers try to insert the same message ID.
  2. SQL Server lets the first insert through and makes the second wait on that key until the first transaction finishes.
  3. When the first commits, the second insert fails with a duplicate key error. The handler rolls back, having done no work, and completes its copy.
  4. If the first handler crashes and rolls back instead, the second insert goes through and that copy does the work. Either way the work happens exactly once.

The database's unique key is the only thing both handlers can see at the same time, which is why it has to make the decision.

A few details matter:

  • Use a fresh DbContext for each message, as the scope above does. After a failed SaveChangesAsync, the context still holds the rejected entities, and handlers running in parallel must not share one.
  • Create the processor with AutoCompleteMessages = false, so you decide when to complete. Completing before the commit, then crashing, loses the message.
  • Senders must set MessageId to something stable, such as the order ID plus the event type. A random ID makes a retried send look like a brand-new message, and none of this can catch it.
  • Keep the transaction short. The claim holds a lock until you commit, so the losing copy waits for as long as the work takes.

On MongoDB: use the message ID as the _id of a processed-messages document. A second insert with the same _id fails with duplicate key error 11000. To keep the claim and the work together, run both in a multi-document transaction. That needs a replica set; every MongoDB Atlas cluster is one.

Add a business-level constraint as a backstop

The message ID protects you from the same message twice. It can't help when the sender publishes the same event twice with different IDs. A unique index on the business key, such as one shipment per order, catches that case as well:

modelBuilder.Entity<Shipment>()
    .HasIndex(s => s.OrderId)
    .IsUnique();

Handle a violation of that index the same way: roll back and complete the message.

Or process one order at a time with sessions

If messages about the same entity must never be processed in parallel, sessions let Service Bus do the serializing. Set SessionId to the order ID when sending and receive with a ServiceBusSessionProcessor. Only one receiver holds a session at a time, so messages for the same order are handled one after another, in order, while different orders still run in parallel.

Sessions have to be enabled when the queue is created, and they cap throughput per session. They don't replace the claim either: a redelivered message still comes back through the same session. What they remove is the race.

When the work isn't in your database

Some work can't join your transaction, like calling a payment API. Claiming the message first still stops a second copy from starting, but if your handler crashes after the payment and before the commit, the redelivered message will try again. For those calls, pass an idempotency key to the other system (many payment and email APIs accept one) and build it from the same message ID. Retrying with the same key then returns the original result instead of repeating the action.

Takeaway

Assume every message can arrive more than once, and sometimes at the same moment. A check followed by a write can't stop two copies racing each other. Claim the message ID with an insert that a unique key protects, do the work in the same transaction, and complete the message only after the commit. Add a unique business key as a backstop, and use sessions when order matters.