Skip to content
Work

02 · Workday

Eight minutes to two

Agentic incident comms on Gemini, PagerDuty and Slack — plus a Claude runbook agent.

Gemini, PagerDuty and Slack in a loop — the agentic comms path from eight minutes to two.

8 → 2Comms, minutes

How I ran it

The sequence on the night.

  1. 01

    PagerDuty fires the P1. The room exists; the broadcast does not.

  2. 02

    Gemini drafts who is hit, what is down, who is on the bridge, what happens next.

  3. 03

    Slack ships it. Eight minutes becomes two — executives act while engineers still pull logs.

  4. 04

    A human still sends it. The agent does not talk to customers alone.

  5. 05

    Claude retrieves the runbook for that service, then parses MTTD / MTTI / MTTR so the PIR is not a memory test.

At Workday I own global P1/P2 major-incident and problem governance for a SaaS platform used by more than half of the Fortune 500. Restoration still depends on engineers and vendors. What used to waste the first eight minutes was the broadcast: who is impacted, what is down, who is on the bridge, what happens next.

I built an agentic workflow on Google Gemini, PagerDuty and Slack that writes and ships those notifications. Delivery time dropped from eight minutes to two. That is not a demo. That is the difference between executives guessing and executives acting while the technical room is still pulling logs.

Alongside it I trained a Claude agent on our operational runbooks, so the team can get step-by-step recovery guidance for the specific service or infrastructure that is down — not a generic chatbot, the book we already wrote, retrieved under pressure. A second Claude model parses major-incident data for MTTD, MTTI, MTTR, owners, gaps and problem trends, so the PIR is not a memory test.

With the automation team I also put proactive ticket alarming into PagerDuty and Jira. Mean time to identify fell 14% over three quarters. The point is not that we use AI. The point is the clock moved.