Ch.19: System Design: File Sync Without Lost Edits
Outline
- 0:00 Two edits, one missing
- 0:29 The sync path
- 0:58 What the system must do
- 1:34 No silent loss
- 2:06 File identity beyond paths
- 2:43 Metadata and immutable content
- 3:52 Upload before commit
- 4:26 Conditional revision writes
- 4:58 Preserve both edits
- 5:47 Catch up from a cursor
- 6:21 When notifications fail
- 6:54 Apply remote changes safely
- 7:35 Client crash windows
- 8:17 Deletes and moves
- 8:51 Permissions beyond content hashes
- 9:25 History versus storage cost
- 9:50 Test failure states
- 10:29 Mira and Leo, replayed
- 11:01 Return to the architecture
Transcript
0:00 Mira edits a shared design file on a plane. Leo edits the same file offline. They started from the same revision. Mira reconnects, then Leo's save silently overwrites hers. Two successful saves, one missing edit. That is the file-sync failure we need to prevent. And her screen may still show a green checkmark. That's the cruel part: we'd trust the icon while her paragraph is gone. How do we keep both versions recoverable and show the collision? Before fixing that overwrite, here is the map. Each device keeps a local journal; a metadata service owns revisions; an immutable block store keeps the bytes; and a change feed tells other devices what happened.
0:41 We will zoom in, then return to this whole path. The public system write-ups behind this exercise are linked below. On this map, the revision service has to decide whether Leo's stale save can replace Mira's visible version. Let's pin down the promise before we choose how. Before picking components, the functional promise is concrete: Mira can upload and download; Leo can edit offline; both devices receive missed changes; and when edits collide, both versions remain recoverable. If Mira renames the folder while Leo is offline, he should still recognize the same file.
1:15 Deletion and sharing have to follow that identity too; a path change must not duplicate the document by accident. So the visible result matters: one current file when changes agree, two recoverable versions when they do not. Those are the functional requirements; next, what must remain safe through a failure? Now the non-functional requirements begin with durability: never silently lose work we promised to preserve. Before upload, Leo's laptop holds the only copy of his edit. That's a different risk from the server overwriting Mira's committed revision.
1:49 Right, the remote side can't rescue bytes it hasn't seen. We also need verified content, resumable transfer and permission checks. If both people changed the file, I'd show two intact versions rather than pick a winner. Speed can wait. Which identifier survives Mira's rename? That rename needs a file ID that survives path changes. If I use `team/design.docx` as the identity, renaming the `team` folder would make one file look like two. Say Mira moves that folder on her laptop. If the server keys by path, it may interpret the old address as a deletion and the new one as a different file.
2:27 Keep that ID; version its parent and name separately. Yeah. A rename should update where that ID lives, not create an unrelated file. Dropbox's published sync-engine work uses stable identity across moves; the source link is in the description. Once a move has a stable identity, separate its mutable record from its content. Mira's rename changes the folder entry in metadata; her committed revision points to content pieces that never change in place. My first thought was to update the file in place.
2:59 But wait, a half-finished transfer could change what readers see. Dropbox documents metadata and immutable content storage separately; our arrangement is an exercise. The new bytes can wait until a revision is ready to publish. With immutable blocks, we can check integrity. A digest lets the client verify downloaded bytes and may reveal pieces it already holds. For example, if Mira changed one paragraph, unchanged pieces may not need to cross the network again, depending how the file was split. A common mistake is to treat matching hashes as proof that a file is current.
3:34 Leo can have perfectly intact bytes from a stale revision, or know a digest after access is revoked. Oh, intact bytes could still be the wrong revision. The hash says they arrived intact; it doesn't say Leo has the current version or permission to publish. Let's follow his edit onto the network. Now follow Leo's pending journal to the network. He uploads missing pieces in resumable chunks; the server verifies them. Only after the bytes are accepted does he ask to commit a metadata revision pointing at them. Why can't the small metadata update go first?
4:08 Airport Wi-Fi may die halfway through a transfer. A published revision must not point to bytes the server lacks. Leo's journal keeps the pending work; Microsoft documents upload sessions that can resume. But having the bytes is only preparation, not permission to replace Mira's revision. With both uploads staged, suppose Mira and Leo started from revision 7. Mira commits first, making revision 8. Leo now requests a replacement based on 7. What prevents the overwrite from our opening? A conditional metadata write: accept Leo's new revision only if the current one is still 7.
4:45 It is 8, so the condition fails. Microsoft Graph documents a stale `If-Match` tag as a precondition failure. Our service makes that compare-and-swap decision as one operation, so revision 8 stays visible. That failed revision check leaves Leo's uploaded blocks uncommitted. How does his client make sure his edit survives? If we clear his pending edit and show an error, we have merely delayed the loss. A product manager would ask, "Can we quietly merge the files and spare people a second filename?"
5:16 I understand the ask. A generic binary file gives us no safe merge we can promise. Keep Leo's edit in the local journal while his client reads Mira's version. Then commit a separate conflict copy with its own file ID, using a retry-safe operation key. Clear the journal entry only after the server acknowledges that copy. So the extra filename looks like sync failed, but it is the receipt that Leo's edit reached durable storage. Now both people can inspect their versions and decide what to combine.
5:46 Once Leo's edit is safe, his laptop still needs Mira's revision. Scanning every remote file after a day offline would be wasteful. What can it ask for instead? Leo's laptop saved a cursor, a marker for the last change it applied before the flight. On reconnect it asks the ordered workspace log for everything after that marker, including Mira's commit. A missed event cannot hide behind an empty notification inbox. Dropbox's public API exposes cursor-based change enumeration. We can use that pattern, but we still have to decide how long a cursor remains usable. Once the cursor handles reconnect, a push could wake an online device sooner.
6:25 But what if that push disappears while the device stays online? A periodic check or foreground refresh still reads the authoritative log, so a duplicate push only triggers one extra query. The notification is just a hint, never the durable event. With a valid cursor, that log check catches the missed change. If the cursor expired, the device takes a current snapshot and reconciles it without discarding pending local edits. The push only affects how soon we ask. And that snapshot cannot erase what Leo still has pending locally.
6:58 Even when Mira's revision reaches his laptop, can we replace the visible file as soon as its blocks arrive? Not with Leo's edit pending. For example, if his application is still writing while sync swaps the file underneath it, the server may be correct while the laptop destroys unsent work. Stage and verify remote bytes first; inspect local state before any visible replacement. Right. Apply an uncontested update atomically. If local and remote versions diverged, keep separate recoverable files. The server cannot protect an edit it never received; the client has to do that.
7:35 Now that local application is safe, consider a server-side loss window. Suppose all blocks reach the server, then the client dies before metadata commit. The bytes exist but no visible revision points at them. The journal can resume or retry; unattached blocks are collected later. Now invert the failure. The commit succeeds, but its acknowledgment disappears. A blind retry could create a duplicate revision. How does the client learn whether the first request landed? The client looks up the original operation key, or retries the same conditional commit with that key.
8:09 The server must return the original result rather than create a second revision. Our crash tests should exercise both windows. Now that the retry key settles the lost acknowledgment, deletion gives us a different replay problem: Mira removes the document while Leo is disconnected. Oh, a digital zombie: if you've ever watched a deleted file return from another device, this is how. A versioned tombstone tells Leo's late laptop the file was deleted. It still has to keep his unsent edit as a recoverable copy.
8:42 The tombstone prevents resurrection without burying Leo's work. A move uses that same stable file ID, so changing a path does not invent a new object. But now Leo's conflict copy exists in the feed. Suppose Mira revokes a collaborator's access while that person's device catches up. Can it still fetch Leo's copy just because it knows the old revision or a block hash? Authorize metadata events and every block read, including after revocation. The hash from our integrity check cannot serve as an access token.
9:13 I had pictured content storage as a dumb lookup by digest. That is unsafe if a caller can bypass workspace permissions. Integrity and authorization are different checks, even for identical bytes. And all those recoverable objects still take space. Which can we eventually collect? Keep Leo's copy; collect orphaned upload pieces after 7 days. I think a 30-day recovery window for prior revisions and tombstones is a defensible promise here. Before collecting a block, check whether a visible or recoverable version still points to it. If cleanup silently erases Leo's last copy, we've broken the promise to him.
9:50 With the retention policy set, how do we prove the no-loss promise under bad schedules? For example, run two offline edits, interrupt an upload, lose a commit acknowledgment, then combine a delete with a pending edit. For each failure schedule, assert that every edit the system promised to keep remains recoverable. Test rename races, stale cursors, and missed pushes too. Dropbox describes deterministic simulation of its sync engine; that is the right spirit here. A green "up to date" badge alone is not the invariant.
10:22 Two devices can converge on the same wrong version. The test has to find the edit that disappeared. Let's run the opening through that test. Mira's edit becomes revision 8. Leo's upload finishes, but his revision-7 condition fails. His local save remains in the journal while Mira's update arrives; his edit is not yet a durable server revision. Leo commits a distinct conflict version and waits for its acknowledgment before clearing that journal entry. Now both versions are visible for resolution.
10:52 A chosen result becomes another conditional commit, and other devices catch up from their cursors. Nobody has to hunt for a vanished paragraph. With the replay complete, return to the map: Leo's journal kept his offline edit, immutable blocks held its bytes, the metadata revision check stopped the overwrite, and the change feed brought both devices to the same visible history. Each did a different job in that one incident. I would test the losing edit's path first, including a crash after its blocks upload.
11:22 If that work survives, performance tuning has a foundation. If it does not, fast sync merely loses data faster. The next time someone says "just sync the folder," ask what happens when both people come back online. Thanks for listening to Learning Podcasts.