Post-mortem
Impact window: ~10:08 AM – 12:30 PM CT, May 31, 2026
Summary: Starting around 10:08 AM CT, some customers experienced delays and failures when sending messages and receiving inbound message events, along with elevated timeouts or 5xx API responses. Message sending and inbound delivery were restored by ~11:25 AM CT; we declared full resolution at ~12:30 PM CT after confirming stability.
Cause: A component in our message-processing pipeline became overloaded during a peak in traffic and stopped processing correctly, interrupting message delivery for a subset of traffic.
Resolution: We rebalanced and restored the affected component, then verified end-to-end recovery. The platform is fully operational.