Global IT infrastructure status

GITHUB

🟢 Operational

Incident with Actions

Resolved

Started
2026-08-26 15:12 UTC
26 days ago
Last update
2026-08-27 02:24 UTC
25 days ago
Duration
2h 49m

Timeline

  1. 27 Aug · 02:24 UTC 🟡 Resolved atom

    Aug 26, 18:01 UTC
    Resolved - On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system.

    At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC.

    3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency for larger-runner jobs.

    Customers using concurrency groups saw longer impact due to a separate issue where runners assigned to a subset of jobs disconnected before the force-revoke mitigation was deployed, which prevented runner acquisition from progressing and left jobs in a waiting-for-runner state. This was resolved at 01:00 UTC on August 27.

    Some runs triggered during the 15:02-15:45 UTC incident window encountered a bug that left them showing as queued eve

  2. 26 Aug · 18:01 UTC 🟡 Resolved atom

    Aug 26, 18:01 UTC
    Resolved - This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

    Aug 26, 18:00 UTC
    Update - All inbound queues have recovered and Actions is operating as expected. 3.7% of jobs assigned to larger runners during the early stage of this incident are stuck waiting for runner assignment. Those will be canceled within the hour. Other runners are successfully processing all new jobs.

    Aug 26, 17:54 UTC
    Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

    Aug 26, 17:32 UTC
    Update - We are continuing to observe recovery and expect actions inbound queues to be back to normal in

  3. 26 Aug · 17:54 UTC 🟡 Monitoring atom

    Aug 26, 17:54 UTC
    Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

    Aug 26, 17:32 UTC
    Update - We are continuing to observe recovery and expect actions inbound queues to be back to normal in

  4. 26 Aug · 17:32 UTC 🟡 Identified atom

    Aug 26, 17:32 UTC
    Update - We are continuing to observe recovery and expect actions inbound queues to be back to normal in

  5. 26 Aug · 16:50 UTC 🟡 Identified atom

    Aug 26, 16:50 UTC
    Update - We are continuing to observe recovery and delayed queues are burning down. Some customers will continue to see increased delays until all throttled work has been completed - we expect this within the next hour.

    Aug 26, 16:49 UTC
    Update - Pages is operating normally.

    Aug 26, 16:14 UTC
    Update - We believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.

    Aug 26, 15:48 UTC
    Update - primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues

    Aug 26, 15:23 UTC
    Update - We've identified an issue with a database primary and are failing over to a replica immediately

    Aug 26, 15:12 UTC
    Update - Pages is experiencing degraded performance. We are continuing to investigate.

    Aug 26, 15:11 UTC
    Investigating - We are investigating reports of degraded availability for Actions

  6. 26 Aug · 16:49 UTC 🟡 Identified atom

    Aug 26, 16:49 UTC
    Update - Pages is operating normally.

    Aug 26, 16:14 UTC
    Update - We believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.

    Aug 26, 15:48 UTC
    Update - primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues

    Aug 26, 15:23 UTC
    Update - We've identified an issue with a database primary and are failing over to a replica immediately

    Aug 26, 15:12 UTC
    Update - Pages is experiencing degraded performance. We are continuing to investigate.

    Aug 26, 15:11 UTC
    Investigating - We are investigating reports of degraded availability for Actions

  7. 26 Aug · 16:14 UTC 🟡 Identified atom

    Aug 26, 16:14 UTC
    Update - We believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.

    Aug 26, 15:48 UTC
    Update - primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues

    Aug 26, 15:23 UTC
    Update - We've identified an issue with a database primary and are failing over to a replica immediately

    Aug 26, 15:12 UTC
    Update - Pages is experiencing degraded performance. We are continuing to investigate.

    Aug 26, 15:11 UTC
    Investigating - We are investigating reports of degraded availability for Actions

  8. 26 Aug · 15:48 UTC 🟡 Identified atom

    Aug 26, 15:48 UTC
    Update - primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues

    Aug 26, 15:23 UTC
    Update - We've identified an issue with a database primary and are failing over to a replica immediately

    Aug 26, 15:12 UTC
    Update - Pages is experiencing degraded performance. We are continuing to investigate.

    Aug 26, 15:11 UTC
    Investigating - We are investigating reports of degraded availability for Actions

  9. 26 Aug · 15:23 UTC 🟡 Identified atom

    Aug 26, 15:23 UTC
    Update - We've identified an issue with a database primary and are failing over to a replica immediately

    Aug 26, 15:12 UTC
    Update - Pages is experiencing degraded performance. We are continuing to investigate.

    Aug 26, 15:11 UTC
    Investigating - We are investigating reports of degraded availability for Actions

  10. 26 Aug · 15:12 UTC 🟡 Identified atom

    Aug 26, 15:12 UTC
    Update - Pages is experiencing degraded performance. We are continuing to investigate.

    Aug 26, 15:11 UTC
    Investigating - We are investigating reports of degraded availability for Actions

Source

Official GitHub status — view original incident ↗

IT Status normalizes vendor wording into a common vocabulary and rewrites machine-shaped titles for readability. The vendor's page remains authoritative.