External services are not always fast or reliable. Webhooks, queues and background jobs help applications handle integrations without blocking users or losing important events.

An integration must assume the other system will sometimes be slow, unavailable or unexpectedly repetitive. Webhooks, queues and background jobs are tools for handling that uncertainty, not decorations added after the first timeout.

Receive events defensively

A webhook endpoint should authenticate the sender, validate the payload, record an event identifier and return promptly. Work that takes time belongs in a Job. Never trust an unverified callback merely because it came to a hard-to-guess URL. Keep a record of events so an operator can see what arrived and what was processed.

The same event may arrive more than once and events may arrive out of order. Identify them with provider event IDs or a stable business key, then make handlers idempotent. When a “paid” event arrives before an earlier “pending” event, the state machine must not move backwards.

Use a Queue to separate response time from work time

A queued job can synchronise data, send a message or generate a file without making the user wait. Set appropriate timeouts and retry policies. Retry temporary failures, but stop and surface permanent validation errors. Monitor failed jobs and give the team a controlled way to replay them after the underlying issue is fixed.

  • Persist an inbound event before acknowledging it when loss would matter.
  • Keep job payloads small and avoid storing secrets in them.
  • Use unique keys for side effects that must happen once.
  • Log correlation IDs to follow an event across services.

Treat external state as a contract

Document which system owns each field, how conflicts are resolved and what happens during an outage. A nightly reconciliation may be more valuable than adding another immediate retry because it catches events that never arrived. Use test environments to simulate delayed and duplicate deliveries.

Reliable integration is an operational design: clear ownership, safe retries, observability and a recovery path. The Queue is only one piece of it.

Trace one event from arrival to outcome

Suppose a shipping provider reports that a parcel was delivered. The webhook handler verifies the signature and stores the provider event ID. A job loads the order, checks whether delivery is a valid next state, updates it once and records an audit entry. A notification follows only after that update succeeds. If the job retries, the event ID prevents a second notification.

This example exposes the questions a simple “receive and update” handler misses: what if the order is cancelled, the payload is incomplete, or an earlier transit event arrives later? The answers belong in the integration contract and tests.

Choose retry policies by failure type

A temporary network error may deserve exponential backoff. A malformed payload or invalid signature should be rejected immediately. A provider rate limit needs a different delay from a database outage. A single generic retry count treats these cases as equivalent and can make an incident worse.

For outbound calls, use timeouts and record an idempotency key where the provider supports it. For inbound events, retain enough raw information to replay safely but protect personal data and set a retention period. Dead-letter or failed-job views should identify the owner of the next action.

Monitor queue age as well as failure count. A queue that never technically fails but runs hours behind still breaks the business promise. Reconciliation compares local and external state to catch the events that no retry can recover because they were never received.

See how I approach reliable integrations and workflow automation.