system reliability field note

At-least-once delivery: why a job can run twice

A broker can preserve a job across worker crashes, but it cannot know whether the job's external side effect already happened. The application must make repetition safe.

By Torollo · Published September 8, 2026 · 13 min read

The acknowledgement gap

  1. 01

    Worker receives a job

    The broker keeps it unacknowledged while processing runs.

  2. 02

    The side effect succeeds

    A card is charged or another external system changes state.

  3. 03

    The worker dies

    No acknowledgement reaches the broker.

  4. 04

    The broker redelivers

    Another worker cannot tell that the earlier side effect completed.

What does at-least-once delivery mean?

At-least-once delivery means the system prefers repeating a message over silently losing it. A consumer acknowledges a delivery only after it has completed the work for which it accepts responsibility. If the connection closes first, the broker can make the unacknowledged message available again.

That recovery behavior creates a clear contract: the same business job may reach a consumer more than once.

RabbitMQ states this directly in its reliability guide: acknowledgements provide at-least-once delivery. Its consumer acknowledgement guide also explains that unacknowledged deliveries are requeued when their channel or connection closes.

At-least-once does not mean every message always runs twice. It means the application must remain correct if one does.

How a duplicate charge happens

Consider a payment job with manual acknowledgement:

  1. A worker receives order A1042.
  2. It sends the charge to the payment provider.
  3. The provider accepts the charge.
  4. The worker crashes before it records completion and acknowledges the message.
  5. RabbitMQ sees the closed connection and requeues the delivery.
  6. Another worker receives order A1042 and sends the charge again.

The broker cannot inspect the payment provider and infer whether step three happened. Redelivery is the safe choice for the message, even though repeating the side effect is unsafe for the customer.

This interval between the side effect and the acknowledgement is the acknowledgement gap. Moving the acknowledgement earlier closes one side of the gap by creating another failure: if the worker acknowledges before charging and then crashes, the payment job is lost.

Duplicates can enter before redelivery

Broker recovery is only one source of repetition. The same business operation can appear more than once because:

  • a browser or API client retries after a timeout;
  • a publisher loses the broker’s response and publishes again;
  • a producer and its database commit are not coordinated;
  • the broker redelivers an unacknowledged message;
  • two workers race on duplicate messages already in the queue;
  • an operator replays a dead-lettered job.

The redelivered flag is useful evidence, but it is not a complete deduplication key. Two separately published messages for the same order can both be first deliveries.

What is idempotency?

An operation is idempotent when repeating the same logical request has the same externally visible effect as performing it once.

For payments, this usually requires a stable idempotency key chosen from the business operation, such as the payment attempt ID. A delivery tag, worker ID or retry number identifies one execution attempt, not the charge the customer intended.

The idempotency record commonly stores four things:

  • the stable operation key;
  • a state such as processing, succeeded or failed;
  • the result or provider reference that later retries should return;
  • enough request identity to reject accidental reuse with different parameters.

Stripe’s idempotent request documentation gives a concrete external API contract: a client sends an idempotency key, and later requests with the same key receive the saved result. Stripe also compares parameters to prevent a key from being reused for a different request.

Why check-then-act is not enough

A first implementation often reads like this in prose: check whether the order was charged, charge it if not, then mark it charged.

Two workers can both pass the check before either writes the result:

Time Worker A Worker B
1 Reads “not charged”
2 Reads “not charged”
3 Sends charge Sends charge
4 Writes result Writes result

The key exists, but the decision is not atomic. Idempotency needs a single winner at the boundary where ownership is claimed. Depending on the store, that can be a unique constraint, a conditional insert, a compare-and-set operation or a transaction that makes duplicate claims conflict.

The losing worker should not guess. It needs a defined response for work that is already complete, still processing, retryable or permanently failed.

The result must be durable before acknowledgement

Acknowledgement order carries the reliability guarantee:

  1. claim the stable operation key atomically;
  2. perform or resume the side effect using the same key downstream;
  3. persist the outcome needed by a retry;
  4. acknowledge the message.

If the result exists durably before the acknowledgement, a redelivery can find it and return or skip the completed work. If the acknowledgement happens first, a crash can erase the only remaining path to completion.

Publisher safety is a separate boundary. Durable queues and persistent messages do not tell a publisher that a broker accepted responsibility for a message. RabbitMQ publisher confirms address that direction of transfer; consumer acknowledgements address the other.

What a lease solves

A lease is a lock with an expiry. It prevents a crashed worker from holding ownership forever. Another worker can take over after the lease expires.

That property helps liveness, but expiry creates a new race. A worker can pause longer than its lease, perhaps because of a long runtime pause, a network delay or a slow dependency. A second worker acquires the expired lease and starts the same operation. The first worker then resumes with an old belief that it still owns the job.

Redis describes this timing boundary in its distributed lock documentation: the lock validity time is also the time available to finish before another client may acquire the resource.

Lease renewal can reduce the chance of expiry during healthy work. It cannot make an old owner harmless after ownership has already moved.

What a fencing token adds

A fencing token is a monotonically increasing number issued when ownership is acquired. Every protected write carries the token. The resource remembers the newest token it has accepted and rejects an older one.

If worker A holds token 41, pauses, and worker B later acquires token 42, the storage layer accepts 42 and rejects A’s late write with 41. The resource must enforce this rule atomically. Comparing the token only inside the worker leaves the same timing gap.

Martin Kleppmann’s article How to do distributed locking illustrates the expired-lease problem and the role of fencing at the storage boundary. Redis also recommends fencing tokens when consistency matters.

A fence does not undo an external charge

A fencing token can keep a stale worker from corrupting a ledger record. It cannot travel backward in time and cancel a card charge the stale worker already sent.

The external payment call needs its own idempotency contract, using the same stable business key on every retry. The local atomic claim protects local ownership. The provider’s idempotency key protects the external side effect. A production design often needs both.

This distinction matters when testing. A dashboard that shows one accepted ledger write does not prove the payment provider saw only one charge.

Exactly-once is an end-to-end property

A queue can provide durable storage, acknowledgements and redelivery. A database can provide unique constraints and transactions. A payment API can provide idempotency keys. None of those components alone knows the whole business operation.

The useful target is an exactly-once effect at the boundary the customer cares about. Achieving it requires the producer, broker, worker, state store and external service to agree on identity and recovery.

AWS makes the same client-intent distinction in Making retries safe with idempotent APIs: the service needs an explicit identifier for the caller’s request rather than trying to infer whether two similar requests mean the same thing.

What to observe in production

Record enough context to reconstruct one operation across several attempts:

  • business operation ID and idempotency key;
  • message ID and broker redelivery flag;
  • delivery attempt and worker identity;
  • lease owner, expiry and fencing token;
  • provider request ID and provider idempotency key;
  • local state transition and acknowledgement time;
  • duplicate, in-progress and stale-owner decisions.

Count redeliveries separately from duplicate business operations. A redelivered message may be handled safely. Two first-delivery messages can still represent one accidental duplicate payment.

A practical design review

Before placing a customer-facing side effect behind a queue, answer these questions:

  1. Which identifier names the customer’s intended operation?
  2. Can two publishers create different messages for that operation?
  3. Where does one worker claim ownership atomically?
  4. What result does a retry receive after success?
  5. What happens while the first attempt is still running?
  6. Can a worker continue after its lease expires?
  7. Which resource rejects a stale fencing token?
  8. Does the external service support the same idempotency key?
  9. Is the outcome durable before the consumer acknowledges?
  10. Can logs connect every delivery attempt to one business operation?

If one answer depends on “the retry probably arrives later,” the failure remains timing-dependent.

Sources and further reading

put the mechanism under load

free foundation

Workers & the Redis job queue

Build the producer, queue and worker baseline before adding payment semantics.

Open the free roadmap

production incident

Charged twice

Crash a payment worker, expose the races and validate the repaired system.

Inspect the incident brief