Reducing the time between incident detection, investigation, and remediation is a critical priority for organizations running production workloads on AWS. When an issue arises, on-call engineers often need to quickly diagnose the problem across application components, identify the root cause, and apply the fix, often in the middle of the night.
AWS DevOps Agent, an AI powered agent that autonomously triages incidents all day based on correlated metrics, logs, and application topology, addresses the first part of this priority by providing root cause analysis (RCA) and recommended actions for resolution. However, to retain control and help prevent unintended changes, organizations typically keep their observability agents, including AWS DevOps Agent, in an observe-and-report mode, where the agent diagnoses issues but doesn’t modify production resources directly. In this post, we demonstrate how to use AWS Lambda Durable Functions, a capability of AWS Lambda, Amazon…


