Resolved
The issue didn't reproduce - we implemented a few counter mesure for future occurences.
Monitoring
Infrastructure is back to "normal" state - we are still working on improving the delay to resorb such spikes even more.
Timeline (CEST):
13:51 - issues affecting our read cluster started pilling up
13.56 - Autoscalers started kicking in
14:01 - Autoscalers fully resorbed the issue
Investigating
We are investigating an increase in latency