Throughput and latency have recovered and the system should be fully operational again.
The core issue is related to the part of the system that persists run state. Some operations experienced and increase in errors and retries, causing a backlog in some parts of the system and some failed operations like checkpoints or signals. We are continuing a thorough post-mortem investigation to prevent reoccurrance.
Resolved
Throughput and latency have recovered and the system should be fully operational again.
The core issue is related to the part of the system that persists run state. Some operations experienced and increase in errors and retries, causing a backlog in some parts of the system and some failed operations like checkpoints or signals. We are continuing a thorough post-mortem investigation to prevent reoccurrance.
Monitoring
We identified issues in the system and have deployed changes to mitigate the issues. Throughput has increased across the system. We continue to monitor latency as well as we expect it to reduce with the throughput increases. We are continuing to investigate for any additional issues that may have occurred.
Investigating
We are actively investigating an issue with latency and execution errors. We will provider further updates as we identify the cause and resolve the issue.