INCIDENT: Service goes down in production; investigation concludes the code was right and the clock was wrong
Four hours of investigation, two dashboards open, and an emergency call ended with the discovery that the server clock had been three minutes fast since the last reboot.
- incident
- ntp
- infrastructure
- clock
notice: This is satire. Names, companies and telemetry are invented. Any resemblance to your environment is your problem.
The incident began at 2:14 a.m. and ended at 6:31 a.m. According to the post-incident report, the code responsible for the closing routine was correct, the automated tests covered the case, and the database behaved "exemplarily."
The problem was the clock.
The server in question had been three minutes and forty-two seconds fast since its last reboot, eleven days earlier. That margin was enough for the processing window to overlap another window, processing two batches in reverse order — and the order, in this case, mattered a great deal.
"We spent four hours looking in the wrong place, which is the place we always look," said the person in charge of the investigation. "The code was the honest part of the system. The clock was the creative part."
The server's time synchronization was restored and clock-drift monitoring was added, with an alert starting at two seconds. The team reported that the same drift had been seen on other occasions, but never with a visible symptom, "which is the worst kind of problem."
The final report closed with a one-line recommendation: "check the clock before checking the code."
Satire. Incidents invented from a real and embarrassing cause: a clock out of sync.