The incident is now resolved and the system is full operational.
Delays began around 14:03 UTC caused by an issue publishing to a new Kafka topic. Recovery began around 14:40 UTC as the team was able to isolate this topic. The consuming services which schedules new function runs was scaled out and began to consume the backlog. No events were dropped during this incident, all functions were scheduled, but with delays during this time window. No action is needed for manual recovery.
Resolved
The incident is now resolved and the system is full operational.
Delays began around 14:03 UTC caused by an issue publishing to a new Kafka topic. Recovery began around 14:40 UTC as the team was able to isolate this topic. The consuming services which schedules new function runs was scaled out and began to consume the backlog. No events were dropped during this incident, all functions were scheduled, but with delays during this time window. No action is needed for manual recovery.
Monitoring
New function execution has resumed and nearly caught up from the incurred backlog. Wait for event and cancel on handlers are still in a backlog and are catching up with processing.
No events were dropped during this issue, all functions will be executed, but with delays.
The root cause has been confirmd and changes are underway to prevent reoccurrence.
Identified
We have identified the cause of the issue - the team has deployed a fix and is scaling to catch up on function delays.
Investigating
We are actively investigating an issue with function execution that is leading to delays starting function runs after events are received. We will provider further updates as we identify the cause and resolve the issue.