Temporal is useful for an AI agent when the execution must survive worker crashes, service restarts, long waits, and retryable failures without restarting the whole job. It does not make an agent's reasoning correct. It makes the surrounding workflow recoverable and inspectable.
The important mechanism is not "saving Python memory after every line." Temporal records workflow events and Activity results. When a Worker must recover a Workflow Execution, it replays that history so the deterministic Workflow code reconstructs its state and continues from the latest recorded event. Temporal's Workflow Execution documentation describes this replay model and the deterministic constraint behind it.
That distinction matters for AI systems. LLM calls, database writes, HTTP requests, and tool invocations belong in Activities, where their failure and retry behavior can be controlled. Workflow code coordinates those effects; it should not perform them directly.
When Temporal earns the complexity
Use Temporal when several of these conditions are true:
- a run lasts longer than a normal request/response window;
- losing completed steps would waste meaningful model cost or operator time;
- multiple effects need separate timeout and retry policies;
- a workflow must wait for human approval or an external event;
- the process needs a durable audit trail and explicit recovery state;
- deployment or worker failure must not erase in-flight work.
Do not add Temporal for a single bounded model call with one retry. A queue, job runner, or application-level retry may be enough. The companion guide on when an AI pipeline needs Temporal compares that boundary with cron, task queues, and Airflow.
The execution model
Four parts carry different responsibilities:
- Workflow: deterministic coordination code. It decides which Activity to run, waits on Temporal-provided awaitables, and holds logical workflow state.
- Activity: a side-effecting operation such as an LLM call, database write, file fetch, or external API request.
- Worker: an application process that polls a Task Queue and runs Workflow or Activity code.
- Temporal Service: stores execution history, schedules Tasks, and tracks workflow progress. It can be self-hosted or supplied by Temporal Cloud.
The Workflow coordinates durable state; Activities contain external effects.
If a Worker disappears, another Worker can replay the history and reconstruct the Workflow state. Completed Activity results are read from history rather than executed again during Workflow replay. An Activity itself can still run more than once when a Worker completes the effect but fails before Temporal records the completion. That is why Activity idempotency is a production requirement, not an optimization.
Put LLM and tool effects in Activities
The Python SDK separates Workflow and Activity definitions. This sketch shows the boundary; llm_gateway represents an application-owned provider client.
from dataclasses import dataclass
from temporalio import activity
@dataclassclass AnalysisRequest: document_id: str text: str
@activity.defnasync def analyze_document(request: AnalysisRequest) -> str: info = activity.info() idempotency_key = f"{info.workflow_run_id}-{info.activity_id}"
return await llm_gateway.summarize( text=request.text, idempotency_key=idempotency_key, )An idempotency key is most useful when the downstream service honors it. If the provider does not, store the operation key and accepted result in an application database, or design the Activity so a retry cannot create a second irreversible effect. Temporal documents Activities as at-least-once execution and recommends idempotent Activity design.
Do not retry every error. Authentication failures, invalid inputs, policy rejections, and unsupported model parameters are usually permanent until something changes. Rate limits and transient network errors may justify backoff. The retry classification should match the provider and the cost of a duplicate attempt.
Coordinate Activities from a deterministic Workflow
from datetime import timedelta
from temporalio import workflowfrom temporalio.common import RetryPolicy
with workflow.unsafe.imports_passed_through(): from activities import AnalysisRequest, analyze_document
@workflow.defnclass DocumentAnalysisWorkflow: @workflow.run async def run(self, documents: list[AnalysisRequest]) -> list[str]: results: list[str] = []
for document in documents: result = await workflow.execute_activity( analyze_document, document, start_to_close_timeout=timedelta(minutes=3), retry_policy=RetryPolicy( initial_interval=timedelta(seconds=2), backoff_coefficient=2.0, maximum_interval=timedelta(seconds=30), maximum_attempts=4, ), ) results.append(result)
return resultsThe retry numbers are examples, not universal defaults. Measure provider latency, rate-limit behavior, per-attempt cost, and the business deadline. A model call that costs dollars per attempt needs a different retry budget from a cheap metadata lookup.
Workflow code must remain deterministic because replay executes it again and checks the generated Commands against recorded history. Direct network calls, uncontrolled randomness, wall-clock access, and other nondeterministic effects belong behind Temporal SDK APIs or Activities. The current Python SDK guide links the Workflow, Activity, Worker, testing, sandbox, and deployment contracts.
Recovery is not exactly-once side effects
Temporal can give a Workflow the effect of durable, resumable execution. It does not turn every external system into an exactly-once database.
For every Activity, decide:
- What happens if the effect succeeds but completion is not recorded?
- Can the effect be safely repeated with the same operation key?
- Which failures are retryable, non-retryable, or ambiguous?
- Which timeout represents execution time, queue delay, or the whole Activity lifecycle?
- What evidence tells an operator whether the external effect occurred?
This is especially important for agents with write-capable tools. "Send the email," "approve the refund," and "update the customer record" require a deduplication contract outside the model. A prompt instruction is not that contract.
Heartbeats are for progress, not decoration
Long-running Activities can heartbeat so Temporal can detect progress, expose heartbeat details, and deliver cancellation. A heartbeat at the start of a short LLM request adds little. For an Activity that processes a large file or polls an external batch job, heartbeat after durable progress and store a safe resume marker.
Heartbeat timeouts, Start-to-Close timeouts, and retry policy solve different failure questions. Configure them deliberately rather than applying one copied policy to every tool call. The Temporal Python error-handling guide distinguishes transient, intermittent, and permanent failures.
Keep execution history bounded
A naive loop over 10,000 documents can create a large history even when every step succeeds. Do not assume a single Workflow Run should accumulate forever.
Use one or more of these patterns:
- split independent work into bounded child Workflows;
- process stable batches with explicit aggregation state;
- use Continue-As-New to start a fresh Run in the same Workflow Execution chain when history grows;
- store large payloads in application storage and pass references, not document bodies, through history;
- bound agent iterations, tool calls, elapsed time, and model cost.
Temporal's Continue-As-New Python guide describes the fresh-history boundary. Continue-As-New is not a substitute for an execution budget: an agent that is making no semantic progress should stop, escalate, or wait for new input rather than renew itself indefinitely.
Production checklist
- Keep Workflow code deterministic and external effects in Activities.
- Give write-capable Activities an idempotency or deduplication contract.
- Classify retryable, permanent, and ambiguous failures.
- Set attempt, time, and cost budgets for LLM calls.
- Heartbeat only when an Activity has meaningful long-running progress.
- Bound history with batches, child Workflows, or Continue-As-New.
- Test replay compatibility before deploying changed Workflow code.
- Version Worker deployments so in-flight executions remain compatible.
- Instrument business progress, model/tool outcomes, retry cost, and stuck states—not only Workflow status.
The deeper operational topics have separate guides: Activity retry patterns for LLM APIs, Workflow versioning for AI pipelines, and Temporal observability beyond workflow status.
The architectural decision
Temporal solves a specific production problem: recoverable coordination across time and failure. It does not choose the right tools, validate model output, or make external side effects exactly once. A durable agent still needs bounded authority, evaluation, idempotent writes, cost controls, and a recovery policy.
If a live system already mixes orchestration, retry, and write-path failures, a Production AI Audit should isolate those boundaries before another framework layer is added.