Experiments · 2022 · Experiment
CloudSync
Distributed, conflict-free multi-cloud synchronization engine.
- Role
- Solo Engineer
- Duration
- 3 months
- Team size
- Solo
Project overview
What this product is
CloudSync is an experiment in treating S3 buckets, WebDAV servers and local folders as one logical filesystem — a sync engine with a small UI, built to answer one question honestly: can multi-cloud sync be conflict-free by design rather than by apology?
I built it over three months as a deliberate systems-design exercise: no framework, no commercial goal — just a real engine, tests that try to corrupt it, and a UI honest enough to show the engine's decisions.
The problem
Why it existed
Files live everywhere now, and every provider promises sync — until you use two at once, work offline, or watch a conflict file appear named 'final(3).docx'.
Who experienced it
Anyone whose work spans providers: photographers shuttling to S3, developers with local + cloud notes, small teams mixing storage they already pay for.
Why current solutions weren't enough
Vendor tools sync vendor storage well and everything else badly; multi-provider tools paper over conflicts with automatic copies that users discover too late.
The idea
Vision and constraints
Original vision
Content-defined chunking plus a Merkle tree per folder: sync decisions are computed from content hashes, not timestamps — the classic source of silent data loss.
Conflicts never resolve silently. When two divergent versions exist, both are preserved and the UI asks one clear question. The engine's job is to make that moment rare and legible.
Development process
01
Research
Studied rsync, Syncthing and rclone — not to copy, but to catalogue every way sync goes wrong.
02
Architecture
Chose content-defined chunking and Merkle diffing; wrote the conflict model before any code.
03
Development
Engine first with a CLI only; the UI was added once the engine survived a chaos suite.
04
Testing
A chaos suite that kills processes, corrupts chunks, rewinds clocks and disconnects networks mid-transfer.
05
Wrap-up
Documented the design openly and archived it as a reference — the experiment did its job.
My role
Exactly what I worked on
Solo project — engine, protocol design, tests and UI.
Sync Engine
Content-defined chunking, Merkle-tree diffing, and an operation log that replays deterministically.
Provider Adapters
S3, WebDAV and local-disk adapters behind one storage interface with capability flags.
Conflict Model
Explicit conflict objects with three-way merge for text and side-by-side keep-both for binary files.
UI
A minimal React dashboard: transfer graph, live queue, and the conflict-resolution flow.
Testing
A chaos suite that kills processes, corrupts chunks and disconnects networks mid-transfer.
Tech stack
The tools that shipped it
Key features
What makes it useful
Content-Addressed Sync
Files are chunked and addressed by content hash — renames and moves are detected for free, timestamps are never trusted.
- Deduplication across providers
- Moves and renames survive any order of operations
- Partial-file resume at chunk granularity
Explicit Conflicts
When versions diverge, both are kept and the user sees a single clear choice — never a silent overwrite, never a hidden copy.
- Three-way merge for plain text
- Keep-both with clear naming for binaries
- Conflict history with full context
One Logical Filesystem
Providers mount as branches of one tree; a file can live anywhere and remain findable everywhere.
- Capability-aware adapters (S3, WebDAV, local)
- Per-branch bandwidth and retention rules
- Transfer graph that shows what moves and why
Design & UX
How it feels to use
The UI's job was legibility: show the engine's reasoning, not just its progress.
The transfer graph
Every sync decision is drawn as a graph of causes — why this file moved, what triggered the check.
Conflicts as first-class citizens
The conflict view is the most-designed screen in the app; it's the moment trust is won or lost.
Engine-grade honesty
The UI never shows optimistic state the engine hasn't committed — no fictional progress.
Technical challenges
And how they were solved
Clocks lie, hashes don't
Timestamp-based sync fails across providers with skewed clocks. Moving all decisions to content hashes eliminated the entire class of 'newer version lost' bugs.
Chaos testing a sync engine
Writing the chaos suite — kill -9 mid-write, truncated chunks, rewound clocks — found more real bugs in a week than months of happy-path use. It remains my favourite artifact of the project.
Provider capability gaps
WebDAV servers disagree about etags, listings and partial writes. Adapters declare capabilities and the engine degrades gracefully — verified integrity where possible, explicit uncertainty where not.
What I learned
Honest takeaways
CloudSync was the deepest systems exercise I've set myself, and it changed how I design everything else.
- Never trust clocks or floating timestamps — content addressing removes a whole category of failure.
- Write the chaos tests first; they are the spec for distributed behaviour.
- Conflict UX is where sync products are actually won or lost.
- Some projects are worth doing for the education alone — this one pays dividends in every system I design.
Results
Outcomes where available
CloudSync is an archived experiment, shared as a design reference. No usage metrics exist — by design.
3 mo
Build duration
Evenings and weekends
0
Silent data-loss paths
Verified by the chaos suite
3
Provider adapters
S3, WebDAV, local disk
—
Usage metrics
None tracked; the experiment is its own result
Final takeaway
Building CloudSync taught me…
CloudSync taught me that correctness is a feature you can architect — and that the hardest, most instructive projects are the ones nobody is paying for.
Gallery
The tree — one logical filesystem, many providers
Conflict view — one clear question
Transfer graph — why files move
Queue — engine-grade honesty
Keep exploring
