Get the latest tech news
A Server Lost Power at 00:32. We Found Out at 08:18
One machine stopped, and S3, the container registry and every deployment stopped with it for eight hours. Our monitoring caught it in three minutes and told nobody. Here is the full postmortem: why single-machine redundancy was not what we thought it was, the near-miss we nearly caused ourselves, and the pager and hardware watchdog we shipped the same day.
None
Or read this on Hacker News
