An AI agent task fails, goes back to the queue, and fails again. The worker may then repeat the same call, discard the message, or leave an operator to discover the problem later. The bigger risk is not knowing whether an earlier attempt already changed an external system.
After a task reaches its retry limit, move it to a dead-letter queue (DLQ) and preserve the evidence needed for a decision. A DLQ should not be a forgotten pile of messages. It is a quarantine state: a person or an explicit policy decides whether the task should be repaired, resumed, compensated, or rejected.
This complements durable execution for AI agents with queues and workflows. Durable execution stores progress and checkpoints. A DLQ organizes what happens when that progress cannot safely continue.

Short answer
- Set retry limits at both the task and worker level.
- Separate transient failure, permanent failure, and uncertain external effects.
- When the budget is exhausted, record the task, cause, attempts, checkpoint, and possible effects in a DLQ.
- Redrive only after checking the cause and the operation's idempotency.
- If repeating is not provably safe, route the task to repair, compensation, or human review.
What does a dead-letter queue solve for an AI agent?
A DLQ gives a task a destination when it should leave the normal path. It keeps a poison message from consuming retries forever and keeps an operator from reconstructing the incident from unrelated logs.
It does not interpret the failure by itself. A broker knows that delivery was not acknowledged. It does not know whether a tool created an order before the connection dropped, whether the model produced invalid arguments, or whether business state changed partially. That meaning belongs to the agent contract and the system that ran the tool.
Agent libraries already provide loop limits for an individual run. The OpenAI Agents JS documentation describes maxTurns and currently documents a default of 10; the Vercel AI SDK documents a default of 20 steps for ToolLoopAgent. Those limits end an execution. They do not retain or investigate the task that ended. The guide to preventing an AI agent from looping forever covers the limit inside the loop; this article focuses on what happens after it. (OpenAI Agents JS, Vercel AI SDK, checked October 8, 2026.)
Which failures belong in the DLQ?
Do not send every exception through the same path. The policy should separate the chance that another attempt will work from the risk that repeating an effect will cause harm.
| Situation | Example | Next state | Automatic redrive? |
|---|---|---|---|
| Transient with no confirmed effect | timeout before a response arrives | delayed retry | maybe, within budget |
| Permanent with no effect | invalid schema or removed tool | DLQ for repair | no |
| Uncertain external effect | timeout after a payment call | DLQ for lookup or compensation | no |
| Policy limit | too many steps or repeated calls | DLQ or cancellation | no |
| Recoverable state | failure after a valid checkpoint | resume workflow | only with a resume contract |
A transient failure is not defined by an HTTP code alone. The same timeout can happen before or after an external service confirms a mutation. Keep the task position, idempotency key, and effect receipt when available. Without them, the worker turns uncertainty into a duplicate.
Temporal's retry-policy documentation distinguishes permanent from transient errors and supports non-retryable errors and attempt limits. A framework decision is input to your contract, not proof that a particular tool is safe to repeat. (Temporal, retry policies, checked October 8, 2026.)
What should the task record contain?
The DLQ record should let an operator make a decision without running the agent blindly. It usually needs:
- Identity:
taskId,runId, task type, tenant, and idempotency key. - Controlled input: a reference to the required payload or snapshot, without copying secrets into the queue.
- History: attempt count, timestamps, agent and model versions, relevant prompt or policy version, and tool names.
- Classified cause: timeout, limit, validation, permission, unavailable dependency, or uncertain effect.
- Checkpoint and artifacts: last confirmed progress, structured output, tool receipt, and related trace.
- Next action:
retry,resume,repair,compensate,reject, orhuman_review, with owner and reason.
Do not retain a full transcript by default. Keep a pointer to a protected artifact and apply the product's retention policy. The guide to what belongs in an AI agent audit trail helps separate operational evidence from a conversation that does not prove the effect. Detailed logs also need access control because an agent failure may contain customer data or tool arguments.
Redrive, resume, repair, or compensation?
These terms are not interchangeable. Redrive creates another attempt from the queue. Resume continues from a checkpoint. Repair changes the input or configuration before trying again. Compensation reverses or neutralizes an effect that already happened.
Use this sequence before moving a task back to the normal queue:
- Find the failure boundary. Did it happen before the call, during the call, or after a partial confirmation?
- Check the external system. A timeout confirms neither success nor failure. Query by idempotency key, receipt, or a safe lookup.
- Classify the cause. Fix schema, permission, or invalid data before redrive; do not raise the limit to hide the diagnosis.
- Choose the starting point. Resume from a checkpoint when the workflow guarantees it. Redrive from the beginning only when every effect is idempotent or the agent can prove it has not happened.
- Record the decision. The next operator needs to know why the task returned, was compensated, or was rejected.
The tool-call verification contract before retrying covers the closest case: an uncertain tool result. The DLQ extends that check to the moment when the retry policy is already exhausted.
Where do queue providers help, and where do they stop?
Provider DLQs reduce operational work, but they do not create recovery semantics for an agent.
- Amazon SQS sends messages to a DLQ after the configured
maxReceiveCountis reached. AWS recommends analyzing DLQ contents and configuring alarms; retention and ordering depend on queue configuration. (AWS, dead-letter queues, checked October 8, 2026.) - Google Cloud Pub/Sub forwards undeliverable messages to a dead-letter topic and tracks delivery attempts approximately. The application still has to interpret the task and define redrive behavior. (Google Cloud, dead-letter topics, checked October 8, 2026.)
In both cases, “it went to the DLQ” is an infrastructure event. The agent contract still needs taskId, external-effect state, and the next allowed action. Do not treat broker delivery count as the full count of turns, tool calls, or side effects.
A small decision contract
The example is illustrative. It does not call a queue or decide whether a payment is idempotent. It makes the information required by the policy explicit before choosing the next state.
type NextAction = "retry" | "resume" | "repair" | "compensate" | "reject" | "human_review";
type Failure = {
attempts: number;
maxAttempts: number;
effect: "none_confirmed" | "confirmed" | "unknown";
checkpoint: boolean;
retryable: boolean;
inputFixed: boolean;
};
export function decideNext(failure: Failure): NextAction {
if (failure.effect === "unknown") return "human_review";
if (failure.effect === "confirmed") return "compensate";
if (failure.checkpoint) return "resume";
if (failure.retryable && failure.attempts < failure.maxAttempts) return "retry";
if (failure.inputFixed) return "repair";
return "reject";
}
The contract is incomplete if effect is always set to none_confirmed because the worker found no receipt. “No confirmation found” is different from “proved that no effect happened.” When that difference matters, use unknown and stop the task. If the run also went through compaction, the context-compaction checklist helps preserve the boundary between confirmed fact and hypothesis.
How should you verify a DLQ before releasing redrive?
Use controlled fixtures, including cases where the tool may have run. A minimum matrix should prove that:
- a timeout before the call can retry within the budget;
- a timeout after a mutation goes to
human_reviewor an idempotency lookup; - a schema error does not consume retries forever;
- a valid checkpoint resumes at the correct step without repeating earlier effects;
- repeated redrive keeps the same idempotency key;
- an operator can see cause, version, artifacts, and owner without the full transcript.
You do not need a live model to prove deterministic state transitions. You do need a controlled integration test when the property depends on the broker, external system, or tool. This article includes no benchmark or production test: the contract is a testable proposal, not a measured result.
Common failures a DLQ does not fix
Treating the DLQ as a trash can
Without an alert owner, the queue only hides the failure. Define retention, alerting, triage owner, response time, and an exit for tasks that cannot be redriven.
Redriving in bulk
A batch can mix authentication errors, invalid data, and uncertain external effects. Redrive by cause class and preserve the decision history. One repaired task does not prove that every other task is safe.
Deleting the payload to meet retention
Short retention can remove the evidence needed to explain the incident. Long retention can increase data risk. Separate the operational envelope from sensitive artifacts and set the rule with security and product owners.
Confusing a turn limit with recovery
maxTurns or a step limit prevents a local loop. It does not explain what happened to a call that lost its response. That question requires idempotency, a receipt, a lookup, or review.
Frequently asked questions
Does every agent failure need a DLQ?
No. A demonstrably transient failure can retry within a budget. The DLQ is for work that exhausted the normal path or needs a different decision before continuing.
Does a DLQ replace checkpoints?
No. A checkpoint says where a workflow can continue. A DLQ says why the task left the normal path and what decision is pending. They complement each other.
Can I automatically redrive after fixing the cause?
Only when the contract proves that the new run is safe, the cause is fixed, and the task has no uncertain external effect. Otherwise, keep explicit review and record the authorization.
How long should a task stay in the DLQ?
That depends on risk, operational agreements, and data retention. Set the period with the system owner and retain the artifacts required to investigate a decision. There is no honest universal number.
Conclusion
A repeatedly failing task needs a better state than “try once more.” A dead-letter queue creates that state, but it is useful only when it preserves execution history, possible external effects, and the next authorized action.
Start by classifying failures, limiting attempts, recording checkpoints, and treating an unknown effect as unknown. Then test redrive, resume, repair, and compensation with fixtures that exercise the dangerous transitions. The system does not need to eliminate every error. It needs to stop a hard failure from becoming an infinite loop, a lost message, or a silent duplicate.
How this article was produced
Samuel Fajreldines is accountable for this article. The research compared current OpenAI Agents JS, Vercel AI SDK, AWS, Google Cloud, and Temporal documentation with public results and recent practitioner discussions. The TypeScript is an illustrative decision model, not a production test. AI assistance supported research organization, drafting, image generation, and localization; it did not run an agent, measure a queue, or supply operational experience the author has not demonstrated. The author also uses RemoteCode as a work tool; this is not a performance claim about the article.