The issue is now fixed, the postmortem of the incident is available below.
--
Summary
On the morning of October 8, Siit stopped processing most of our background tasks between about 07:45 and 09:40 CEST. During that time, notifications, emails, Slack and Teams messages, workflows and integration syncs were delayed. Some interactive actions, such as clicking a button or talking to the AI Agent in a Slack or Teams message, couldn't be completed and may have needed a retry.
While we caught up on the delayed work that had piled up, the app was slower than usual until about 11:00 CEST.
All delayed work was queued safely and processed once service was restored. No data was lost.
What happened
A routine configuration cleanup after a recent release left our background processing unable to restart during scaling phases.
The effect built up gradually overnight and became visible in the morning when activity picked up.
How we resolved it
We found the cause a few minutes after the investigation started and restored processing, handling the most time-sensitive tasks first. Catching up on the backlog put heavy load on our database, so we added database capacity and spread out the remaining work. Everything was back to normal by 11:00 CEST.
Where we failed
We had no alert for this specific failure, and the alerts we did have relied on the affected system itself
As a result, the problem went unnoticed until it impacted you.
What we're changing
We are revamping our Alerting, so such cases are caught right away.
Stricter safeguards and automated checks on production configuration changes.
We’ve upgraded our database capacity
Improving out-of-hours alerting and escalation
We're sorry for the disruption to your teams. If an action from that morning still looks incomplete, retrying it should work. If not, our support team is happy to help.
Resolved
The issue is now fixed, the postmortem of the incident is available below.
--
Summary
On the morning of October 8, Siit stopped processing most of our background tasks between about 07:45 and 09:40 CEST. During that time, notifications, emails, Slack and Teams messages, workflows and integration syncs were delayed. Some interactive actions, such as clicking a button or talking to the AI Agent in a Slack or Teams message, couldn't be completed and may have needed a retry.
While we caught up on the delayed work that had piled up, the app was slower than usual until about 11:00 CEST.
All delayed work was queued safely and processed once service was restored. No data was lost.
What happened
A routine configuration cleanup after a recent release left our background processing unable to restart during scaling phases.
The effect built up gradually overnight and became visible in the morning when activity picked up.
How we resolved it
We found the cause a few minutes after the investigation started and restored processing, handling the most time-sensitive tasks first. Catching up on the backlog put heavy load on our database, so we added database capacity and spread out the remaining work. Everything was back to normal by 11:00 CEST.
Where we failed
We had no alert for this specific failure, and the alerts we did have relied on the affected system itself
As a result, the problem went unnoticed until it impacted you.
What we're changing
We are revamping our Alerting, so such cases are caught right away.
Stricter safeguards and automated checks on production configuration changes.
We’ve upgraded our database capacity
Improving out-of-hours alerting and escalation
We're sorry for the disruption to your teams. If an action from that morning still looks incomplete, retrying it should work. If not, our support team is happy to help.
Monitoring
All our monitoring data now appears back to normal, we will continue monitoring and provide a post mortem shortly.
Monitoring
We are still having impactful latency until we reach full recovery
Monitoring
Our fix appears to be working - our process are still catching their queue so you could still expect some latency in the delivery of pilled up notifications.
We’ll come up with a detailed reports once we confirm that things settled down for good
Investigating
We have pushed a fix and expect recovery to happen during the next 10-40 minutes
Investigating
We are receiving multiple reports and alerts of issues affecting our app