So from what I am seeing in this with a brief look over it, the only cases in which data loss seemed to occur were when two clients were editing the same file temporally close to each other? I.e. you end up creating something similar to a git merge conflict, which cannot be solved automatically well, and thus can generate loss of data.
Mysteries of Dropbox: Testing of a Distributed Sync Service (2016) [pdf]
11–20 of 28 posts
Re: Mysteries of Dropbox: Testing of a Distributed Sync Service (2016) [pdf]
#12I was lead on Syncplicity's desktop client. File synchronization has a myriad of corner cases that are difficult and non-intuitive to think through; and non-programmers often thoroughly underestimate just how difficult these are to anticipate and mitigate. The fact that they found bugs that rely on sensitive timing doesn't surprise me.
Can you share which difficult and non-intuitive corner cases there are? I guess debouncing, etc.
- Receiver tried to create a file before receiving attributes of the directory containing the file. Receiver author assumed it would always receive directory attributes first and create the directory, so it crashed.
- Receiver created a file before receiving attributes of the directory containing the file. Parent directory was created automatically, but with default attributes so the file was too accessible on the receiver when it should not have been.
- Bidirectional sync peers got into a non-terminating protocol loop (livelock) when trying to agree if a directory deep in a tree should be empty or removed (garbage collected) after synchronising removal of contents. It always worked if one side changed and sync settled before the next change, but could fail if both sides had concurrent changes.
- Mesh sync among multiple peers, with some of them acting as publish-subscribe proxies forwarding changes to others as quickly as possible merged with their own changes, got into a more complicated non-terminating protocol loop when trying to broadcast and reconcile overlapping changes observed on three or more nodes concurrently. The solution was similar to distributed garbage collecting and spanning tree protocols used in Ethernet switch networks.
- Transmission of commands halted due to head of line blocking (deadlock) on a multiplexed sync stream because a data channel was going to a receiver process whose buffer filled while waiting for a command on the command channel, which the transmitter process had issued but couldn't transmit. The fault was separate, modular tasks assuming data for each flowed independently. The solution was to multiplex correctly with per-channel credits like HTTP/2 and QUIC, instead of incorrectly assuming you can just mix formatted messages over TCP.
- Rendered pages built from mesh data-synchronised components, similar to Dropbox-style sync'd files but with a mesh of 1000s of peers, showing flashes of inconsistent data, e.g. tables whose columns should always add to 100% showing a different total (e.g. "110% (11050 of 10000) devices online"), displayed addresses showing the wrong country, numbers of devices exceeeding the total number shipped, devices showing error flags yet also "green - all good" indication, number of comments not matching the shown commments, number of rows not matching rows in a table, etc. Usually for only a few seconds, sometimes staying on screen for a long time if the 3G network went down, or if rendered to a PDF report. Such glitches made the underlying systems look like they had a lot of bugs when they really didn't, especially when captured in a PDF report. It completely undermined trust in the presented data being something you could rely on. All for want of more careful synchronisation protocol.
Re: Mysteries of Dropbox: Testing of a Distributed Sync Service (2016) [pdf]
#13Anything written by John Hughes is worth a read. He also also wrote quickcheck.
Re: Mysteries of Dropbox: Testing of a Distributed Sync Service (2016) [pdf]
#14Earlier quoted context omitted.
Can you share which difficult and non-intuitive corner cases there are? I guess debouncing, etc.
The way I used to explain it: Imagine that you are on a plane, (and don't have an internet connection). You edit a file. At the same time, I edit that file. What should we do? We can't possibly know every file format out there, and implement operational transform for all of them. Now, imagine that we both edit the same file, at the same instant. One of us is going to submit the change first, and the other will submit…
Re: Mysteries of Dropbox: Testing of a Distributed Sync Service (2016) [pdf]
#15Re: Mysteries of Dropbox: Testing of a Distributed Sync Service (2016) [pdf]
#16Earlier quoted context omitted.
The way I used to explain it: Imagine that you are on a plane, (and don't have an internet connection). You edit a file. At the same time, I edit that file. What should we do? We can't possibly know every file format out there, and implement operational transform for all of them. Now, imagine that we both edit the same file, at the same instant. One of us is going to submit the change first, and the other will submit…
oh and the parent folder is on a shared NAS with some caching.
The root cause of the problem is that in .net, there is a bug with File.Exists. If there is a filesystem / network error, instead of getting an exception, the error is swallowed and the call just returns false. I'm not sure if newer versions of .net fix it or not; I only learned about this when we were implementing a driver / filesystem.
Re: Mysteries of Dropbox: Testing of a Distributed Sync Service (2016) [pdf]
#17Earlier quoted context omitted.
oh and the parent folder is on a shared NAS with some caching.
We had to add logic to block network and USB drives. (They were an ever-present source of customer issues.) The root cause of the problem is that in .net, there is a bug with File.Exists. If there is a filesystem / network error, instead of getting an exception, the error is swallowed and the call just returns false. I'm not sure if newer versions of .net fix it or not; I only learned about this when we were implemen…
Re: Mysteries of Dropbox: Testing of a Distributed Sync Service (2016) [pdf]
#18One of the authors, John Hughes did a talk on property-based testing at Clojure West some number of years back. Worth a watch if you're interested: https://www.youtube.com/watch?v=zi0rHwfiX1Q
Re: Mysteries of Dropbox: Testing of a Distributed Sync Service (2016) [pdf]
#19There was a discussion of a self-built dropbox on the frontpage ( https://news.ycombinator.com/item?id=47673394 ). This is just to show that dropbox is thoroughly tested for all kinds of wierd interactions and behaviours across OS using a very formal testing framework.
Not everything is a CRUD app website.
I was running my own hacky sync thing to the cloud a decade ago. I would never in my boldest dreams compared it to dropbox.
Even if you know the use cases, the edge cases could be 99% of the work. POCs are 100x easier than working production multi-user applications. Don’t confuse getting to a POC in 2 hours with getting a final product in 4 hours.
Re: Mysteries of Dropbox: Testing of a Distributed Sync Service (2016) [pdf]
#20I was lead on Syncplicity's desktop client. File synchronization has a myriad of corner cases that are difficult and non-intuitive to think through; and non-programmers often thoroughly underestimate just how difficult these are to anticipate and mitigate. The fact that they found bugs that rely on sensitive timing doesn't surprise me.
Can you share which difficult and non-intuitive corner cases there are? I guess debouncing, etc.
You could have it mirror an entire subdirectory, including external drives.
If you booted up long enough and that external drive was not mounted, the service registered that as a subdirectory delete (bad). When you then mounted it again, the sync agent saw it as out of sync with the newer server-side delete and proceeded to clear the local external drives.
They also implemented versioning so poorly that a deleted directory was not versioned, only the files within it. So you could recover raw files without the directory structure back in a giant bundle of 1000s of files. Horrible.
See: https://dynamicsgpland.blogspot.com/2011/11/one-significant-...