Live data from Hacker News

NTP at NIST Boulder Has Lost Power

lists.nanog.org

71–80 of 218 posts

Re: NTP at NIST Boulder Has Lost Power

#71
post #47

This makes me wonder, if you take the average time of all wristwatches on the planet, accounting for timezones and throwing out outliers, how close would you get to NTP time? And how many randomly chosen wristwatches would you need to get anything reasonable?

I have a hunch my casio wrist watch is designed to be running a bit too quick to make resetting the seconds easier. Your averaging assumes manufacturers try to make their watches as accurate as possible for average conditions

I think it runs quick to be on the safe side, so you never miss appointments, trains, etc. because of your watch.

But yes, good point.

Re: NTP at NIST Boulder Has Lost Power

#72
post #10

Earlier quoted context omitted.

Time travel is extremely dangerous right now. I highly recommend deferring time travel plans except for extreme temporal emergencies.

Would traveling to the past in order to put in place a preemptive fix for this outage be wise or dangerous? Asking for a friend.

Safety not guaranteed.

Re: NTP at NIST Boulder Has Lost Power

#73
post #13
post #5

Can anybody expand on the implications of this? Being unfamiliar with it, it's hard to tell if this is a minor blip that happens all the time, or if it's potentially a major issue that could cause cascading errors equal to the hype of Y2K.

Google has their own fleet of atomic clocks and time servers. So does AWS. So does Microsoft. So does Ubuntu. They're not going to drift enough for months to cause trouble. So the Internet can ride through this, mostly. The main problem will be services that assume at least one of the NIST time servers is up. Somewhere, there's going to be something that won't work right when all the NIST NTP servers are down. But wh…

Can't they point these dns records to working servers meanwhile to avoid degradation?

Re: NTP at NIST Boulder Has Lost Power

#74
post #10

Earlier quoted context omitted.

Time travel is extremely dangerous right now. I highly recommend deferring time travel plans except for extreme temporal emergencies.

Same for database transaction roll back and roll forward actions. And most enterprises, including banks, use databases. So by bad luck, you may get a couple of transactions reversed in order of time, such as a $20 debit incorrectly happening before a $10 credit, when your bank balance was only $10 prior to both those transactions. So your balance temporarily goes negative. Now imagine if all those amounts were ten th…

To clear, most banks "sort" transactions from to high to low to create more NSF fees on purpose.

That purpose equates to over $12 billion in fees for 2024

https://finhealthnetwork.org/research/overdraft-nsf-fees-big...

Re: NTP at NIST Boulder Has Lost Power

#75
post #73
post #13

Earlier quoted context omitted.

Google has their own fleet of atomic clocks and time servers. So does AWS. So does Microsoft. So does Ubuntu. They're not going to drift enough for months to cause trouble. So the Internet can ride through this, mostly. The main problem will be services that assume at least one of the NIST time servers is up. Somewhere, there's going to be something that won't work right when all the NIST NTP servers are down. But wh…

Can't they point these dns records to working servers meanwhile to avoid degradation?

My understanding is that people who connect specifically to the NIST ensemble in Boulder (often via a direct fiber hookup rather than using the internet) are doing so because they are running a scientific experiment that relies on that specific clock. When your use case is sensitive enough, it's not directly interchangable with other clocks.

Everyone else is already connecting to load balanced services that rotate through many servers, or have set up their own load balancing / fallbacks. The mistakenly hardcoded configurations should probably be shaken loose anyways.

Re: NTP at NIST Boulder Has Lost Power

#76
post #17
post #13

Earlier quoted context omitted.

Google has their own fleet of atomic clocks and time servers. So does AWS. So does Microsoft. So does Ubuntu. They're not going to drift enough for months to cause trouble. So the Internet can ride through this, mostly. The main problem will be services that assume at least one of the NIST time servers is up. Somewhere, there's going to be something that won't work right when all the NIST NTP servers are down. But wh…

Atomic clock non-expert here, what does having a fleet of atomic clocks entail and why would the hyperscalers bother?

There's a lot of focus in this thread on the atomic clocks but in most datacenters, they're not actually that important and I'm dubious that the hyperscalers actually maintain a "fleet" of them, in the sense that there are hundreds or thousands of these clocks in their datacenters.

The ultimate goal is usually to have a bunch of computers all around the world run synchronised to one clock, within some very small error bound. This enables fancy things like [0].

Usually, this is achieved by having some master clock(s) for each datacenter, which distribute time to other servers using something like NTP or PTP. These clocks, like any other clock, need two things to be useful: an oscillator, to provide ticks, and something by which to set the clock.

In standard off-the-shelf hardware, like the Intel E810 network card, you'll have an OXCO, like [1], with a GPS module. The OXCO provides the ticks, the GPS module provides a timestamp to set the clock with and a pulse for when to set it.

As long as you have GPS reception, even this hardware is extremely accurate. The GPS module provides a new timestamp, potentially accurate to within single-digit nanoseconds ([2] datasheet), every second. These timestamps can be used to adjust the oscillator and/or how its ticks are interpreted, such that you maintain accuracy between the timestamps from GPS.

The problem comes when you lose GPS. Once this happens, you become dependent on the accuracy of the oscillator. An OXCO like [1] can hold to within 1µs accuracy over 4 hours without any corrections but if you need better than that (either more time below 1µs, or more accurate than 1µs over the same time), you need a better oscillator.

The best oscillators are atomic oscillators. [2] for example can maintain better than 200ns accuracy over 24h.

So for a datacenter application, I think the main reason for an atomic clock is simply for retaining extreme accuracy in the event of an outage. For quite reasonable accuracy, a more affordable OXCO works perfectly well.

[0]: https://docs.cloud.google.com/spanner/docs/true-time-externa...

[1]: https://www.microchip.com/en-us/product/OX-221

[2]: https://www.u-blox.com/en/product/zed-f9t-module

[3]: https://www.microchip.com/en-us/products/clock-and-timing/co...

Re: NTP at NIST Boulder Has Lost Power

#77
post #53
post #48

Earlier quoted context omitted.

What happens in the event all the sites for time.nist.gov go down? is it included in the spec? Also thank you for that link, this is exactly the kind of esoteric knowledge that I enjoy learning about

Most high-availability networks use pool.ntp.org or vendor-specific pools (e.g., time.cloudflare.com, time.google.com, time.windows.com). These systems would automatically switch to a surviving peer in the pool. Many data centers and telecom hubs use local GPS/GNSS-disciplined oscillators or atomic clocks and wouldn’t be affected. Most laptops, smartphones, tablets, etc. would be accurate enough for days before drift…

Great list! Just double-checked the CAT timekeeping requirements [1] and the requirement is NIST sync. So a subset of all UTC.

You don’t need to actually sync to NIST. I think most people PTP/PPS to a GPS-connected Grandmaster with high quality crystals.

But one must report deviations from NIST time, so CAT Reporters must track it.

I think you are right — if there is no NIST time signal then there is no properly auditable trading and thus no trading. MFID has similar stuff but I am unfamiliar.

One of my favorite nerd possessions is my hand-signed letter from Judah Levine with my NIST Authenticated NTP key.

[1] https://www.finra.org/rules-guidance/rulebooks/finra-rules/6...

Re: NTP at NIST Boulder Has Lost Power

#78
post #14

Wind gusts were reaching 125 MPH in Boulder county, if anyone’s curious. A lot of power was shut off preemptively to prevent downed power lines from starting wildfires. Energy providers gave warning to locals in advance. Shame that NIST’s backup generator failed, though.

Notably, we had the marshal fire here 4 years ago and recently Xcel settled for $680M for their role in the fire. So they're probably pretty keen not to be on the hook again

Re: NTP at NIST Boulder Has Lost Power

#79
post #55

Earlier quoted context omitted.

If your computer was using it as your time server and you didn't have alternatives configured your clock my have drifted a few seconds.

I never checked it, but how much a typical's pc/server's clock does actually drift over a week or a month? I always thought it's well under a second.

Clocks do drift. Seconds a week is definitely possible. I think there are varying quality of internal clocks in electronic devices, and the cheaper the quality the more drift there is. I think small cheap microcontrollers can drift seconds per day.

Re: NTP at NIST Boulder Has Lost Power

#80
post #10
post #5

Can anybody expand on the implications of this? Being unfamiliar with it, it's hard to tell if this is a minor blip that happens all the time, or if it's potentially a major issue that could cause cascading errors equal to the hype of Y2K.

Time travel is extremely dangerous right now. I highly recommend deferring time travel plans except for extreme temporal emergencies.

Uhh, here's the problem, I'm sort of stuck travelling into the future at a more or less constant rate. I don't know how to stop doing that...
Post reply on HN