Failure Recovery & Duplicate Prevention for AI Workflows by Hardi ReiljanFailure Recovery & Duplicate Prevention for AI Workflows by Hardi Reiljan

Failure Recovery & Duplicate Prevention for AI Workflows

Hardi Reiljan

Hardi Reiljan

Role: Python Workflow Reliability Engineer
Hardened a persistent job runtime against failure and replay problems that appear after the happy path: lease expiry, worker restart, bounded retries, partial completion, and duplicate-sensitive external actions.
The runtime uses explicit state, attempt ceilings, idempotency keys, lease recovery, structured failure behavior, and regression tests around replay semantics. One self-audit exposed a path where a single-attempt external job could become executable again after lease expiry; the fix now enforces the configured attempt ceiling across recovery.
Like this project

Posted Aug 30, 2026

Hardened persistent workflows against lease expiry, worker restarts, bounded retries, partial completion, and duplicate-sensitive external actions.