Live data from Hacker News

How Precision Time Protocol is being deployed at Meta

engineering.fb.com

61–70 of 104 posts

Re: How Precision Time Protocol is being deployed at Meta

#61
post #43

Earlier quoted context omitted.

At Meta you choose what team you want to work on during a 3 month “bootcamp” period, very different from Google, Amazon, etc. There’s enough teams solving hard distributed system problems that you’re guaranteed to never need to think about button copy

You can also be directly hired onto a team, skipping bootcamp (some roles).

you skip team selection, not bootcamp. every engineer and EM go through bootcamp

Re: How Precision Time Protocol is being deployed at Meta

#62
I have to say I am not a fan of Facebook in general (their primary business that is) but their engineering is always impressive. And equally impressive is that they do everything in house.

An interesting movie might be a dystopian future where Facebook itself falls out of use but some type of rogue group takes over the Facebook datacenter infrastructure and uses it for their own purposes.

Re: How Precision Time Protocol is being deployed at Meta

#63
Given Meta's resources I'm surprised they didn't go whole-hog with the high accuracy profile (derived from CERN White Rabbit) Why not set up SyncE while you're at it? Not like the hardware cost is any issue and Meta can hire hardware engineers.

Might as well get ps level precision with sub-ns accuracy

Re: How Precision Time Protocol is being deployed at Meta

#64

Companies have been using PTP to synchronize (frame sync) networked cameras for many years. It is much better than wiring up a separate ttl or differential signals. It is amazing the accuracy you can get with this protocol.

I've also seen seismologists use it to sync geophones and other instruments. Synchronization is very important for doing time sensitive measurements for things like localization of earthquakes.

Re: How Precision Time Protocol is being deployed at Meta

#65
post #41
post #34

Earlier quoted context omitted.

It's an effective marketing campaign for "you should work here" - VERY effective. And lots of this stuff is NOT secret sauce, it's basic business building-blocks that they need. It's not the advertising formulas. I'm sure FAANG is very VERY happy that they can just run Linux everywhere and don't have to pay Sun or Microsoft a massive per-CPU fee for everything they do.

I think you'd normally be right, but Meta just laid off 10k+ people, and is currently in a company wide hiring freeze until at least Q1 of 2023. Much of the rest of FAANG is either doing one or both as well. In Meta's case, they fired a lot of "boot campers" as well, some only a few days into their job and before they had a team. Some returning interns even had their offers rescinded. Not to sound cynical, but this a…

It's likely that this article was in the works well before the layoffs. Depending on field it can take a ridiculously long time to get a tech blog post approved at Meta.

Re: How Precision Time Protocol is being deployed at Meta

#66

So they state: > One could argue that we don’t really need PTP for that. NTP will do just fine. Well, we thought that too. But experiments we ran comparing our state-of-the-art NTP implementation and an early version of PTP showed a roughly 100x performance difference: While I'm not necessarily against more accuracy/precision, what problems specifically are experiencing? They do mention some use cases of course: > Th…

First, 300mm is not the real measure in practice for the common use case. PTP is often used to distribute GPS time for things that need it but don't have direct satellite access, and also so you don't have to have direct satellite access everywhere.

For that use case, 1ns of inaccuracy is about 10ft all told (IE accounting for all the inaccuracy it generates).

It can be less these days, especially if not just literally using GPS (IE a phone with other forms of reckoning, using more than just GPS satellites, etc). You can get closer to the 1ns = 1ft type inaccuracy.

But if you are a cell tower trying to beamform or something, you really want to be within a few ns, and without PTP that requires direct satellite access or somet other sync mechanism.

Second, I'm not sure what you mean by special. Expense is dictated mostly by holdover and not protocol. It is true some folks gouge heavily on PTP add-ons (orolia, i'm looking at you), but you can ignore them if you want. Linux can do fine PTP over most commodity 10G cards because they have HW support for it. 1G cards are more hit or miss.

For dedicated devices: Here's a reasonable grandmaster that will keep time to GPS(/etc) with a disciplined OCXO, and easily gets within 40ns of GPS and a much higher end reference clock i have. https://timemachinescorp.com/product/gps-ntpptp-network-time...

It's usually within 10ns. 40ns is just the max error ever in the past 3 years.

Doing PTP, the machines stay within a few NS of this master.

If you need better, yes it can get a bit expensive, but honestly, there are really good OCXO out there now with very low phase noise that can more accurately stay disciplined against GPS.

Now, if you need real holdover for PTP (IE , yes you will probably have to go with rubidium, but even that is not as expensive as it was.

Also, higher end DOCXO have nearly the same performance these days, and are better in the presence of any temperature variation.

As for me, i was playing with synchronizing real-time motion of fast moving machines that are a few hundred feet apart for various reasons. For this sort of application, 100us is a lot of lag.

I would agree this is a pretty uncommon use case, and I could have achieved it through other means, this was more playing around.

AFAIK, the main use of accurate time at this level is cell towers/etc, which have good reasons to want it.

I believe there are also some synchronization applications that have need of severe accuracy (synchronous sound wave generation/etc) but no direct access to satellite signal (IE underwater arrays).

Re: How Precision Time Protocol is being deployed at Meta

#67
post #2

This is incredible. If I knew I'd be working on things like this (as opposed to optimizing ad revenue or "user engagement" by A/B testing button copy), I'd possibly consider working at Meta.

At Meta you choose what team you want to work on during a 3 month “bootcamp” period, very different from Google, Amazon, etc. There’s enough teams solving hard distributed system problems that you’re guaranteed to never need to think about button copy

In good times, yes, there are more roles than bootcampers. In bad times (now) there are more bootcampers than roles...

Re: How Precision Time Protocol is being deployed at Meta

#68

So they state: > One could argue that we don’t really need PTP for that. NTP will do just fine. Well, we thought that too. But experiments we ran comparing our state-of-the-art NTP implementation and an early version of PTP showed a roughly 100x performance difference: While I'm not necessarily against more accuracy/precision, what problems specifically are experiencing? They do mention some use cases of course: > Th…

Probably high frequency hedge funds because being first matters a lot. That 300mm is the difference between winning and losing, it’s pretty binary up there.

Re: How Precision Time Protocol is being deployed at Meta

#69

So they state: > One could argue that we don’t really need PTP for that. NTP will do just fine. Well, we thought that too. But experiments we ran comparing our state-of-the-art NTP implementation and an early version of PTP showed a roughly 100x performance difference: While I'm not necessarily against more accuracy/precision, what problems specifically are experiencing? They do mention some use cases of course: > Th…

Most professional network cards have PTP support and a grandmaster is cheap enough or your colo provider will provide PTP as a service.

Re: How Precision Time Protocol is being deployed at Meta

#70
post #59

I’m wondering whether the extremely careful GNSS part is really needed. A microsecond of offset between two servers in the same datacenter could easily matter, but I suspect that, if an entire datacenter were off by a microsecond, everything would be fine — communicating from that datacenter to anywhere else will take well over a microsecond, so an offset of this type would be a bit like the datacenter wiggling aroun…

Having rooftop mounted GNSS receive antennas for GPS+GLONASS is extremely common in telecom and ISP infrastructure applications. It's sort of a belt and suspenders approach to obtaining time from low stratum NTP sources and also having a local GNSS timing source to reference from. Or for use in a case where the network has a total absence of connectivity to any internet-based NTP sources (maybe because your managemen…

Meta seems to have gone with the expensive option. I think it’s this:

https://www.mouser.com/ProductDetail/HUBER%2bSUHNER/Direct-G...

Admittedly, at their scale, this is peanuts. But I wouldn’t buy one of these for a scrappy startup :). SparkFun will sell a perfectly serviceable kit for a few hundred dollars.

(If you are worried about lightning, then GPSoF looks like cheap insurance.)

Post reply on HN