CircleCI outage: transaction log filled its disk and stalled pipelines for 5 hours
On October 6, CircleCI suffered a five-hour outage: the transaction log of its orchestration database filled a dedicated volume, blocking all writes. Pipeline starts were delayed from 15:30 to 20:30 UTC and stopped entirely from 15:48 to 16:38. Moving the logs to the main volume took the database offline for 50 minutes; full throughput returned by 20:30 UTC.
- Outage lasted about 5 hours: delays from 13:30 UTC, no starts at all from 15:48 to 16:38
- Cause: the database's transaction log filled its dedicated volume faster than the archiver could drain it
- Moving the logs to the main volume required 50 minutes of database downtime
- After recovery throughput was about half normal until a setting change at 18:19 UTC
Read next
Software