Metrics in the UI are now restored. Users may notice metrics data continue to be unavailable during the time of the incident, from the timeframe 7/16/26 15:45 UTC until approximately 7/16/26 21:07 UTC
Throughput and latency have recovered and the system should be fully operational again.
The core issue is related to the part of the system that persists run state. Some operations experienced and increase in errors and retries, causing a backlog in some parts of the system and some failed operations like checkpoints or signals. We are continuing a thorough post-mortem investigation to prevent reoccurrance.