Experiments · 2022 · Experiment

CloudSync

Distributed, conflict-free multi-cloud synchronization engine.

Role
Solo Engineer
Duration
3 months
Team size
Solo
cloudsync - fs mapMERKLE ROOTS3WEBDAVLOCALCONFLICT - TWO VERSIONS
01

Project overview

What this product is

CloudSync is an experiment in treating S3 buckets, WebDAV servers and local folders as one logical filesystem — a sync engine with a small UI, built to answer one question honestly: can multi-cloud sync be conflict-free by design rather than by apology?

I built it over three months as a deliberate systems-design exercise: no framework, no commercial goal — just a real engine, tests that try to corrupt it, and a UI honest enough to show the engine's decisions.

02

The problem

Why it existed

Files live everywhere now, and every provider promises sync — until you use two at once, work offline, or watch a conflict file appear named 'final(3).docx'.

Who experienced it

Anyone whose work spans providers: photographers shuttling to S3, developers with local + cloud notes, small teams mixing storage they already pay for.

Why current solutions weren't enough

Vendor tools sync vendor storage well and everything else badly; multi-provider tools paper over conflicts with automatic copies that users discover too late.

03

The idea

Vision and constraints

Original vision

Content-defined chunking plus a Merkle tree per folder: sync decisions are computed from content hashes, not timestamps — the classic source of silent data loss.

Conflicts never resolve silently. When two divergent versions exist, both are preserved and the UI asks one clear question. The engine's job is to make that moment rare and legible.

Development process

  1. 01

    Research

    Studied rsync, Syncthing and rclone — not to copy, but to catalogue every way sync goes wrong.

  2. 02

    Architecture

    Chose content-defined chunking and Merkle diffing; wrote the conflict model before any code.

  3. 03

    Development

    Engine first with a CLI only; the UI was added once the engine survived a chaos suite.

  4. 04

    Testing

    A chaos suite that kills processes, corrupts chunks, rewinds clocks and disconnects networks mid-transfer.

  5. 05

    Wrap-up

    Documented the design openly and archived it as a reference — the experiment did its job.

04

My role

Exactly what I worked on

Solo project — engine, protocol design, tests and UI.

Sync Engine

Content-defined chunking, Merkle-tree diffing, and an operation log that replays deterministically.

Provider Adapters

S3, WebDAV and local-disk adapters behind one storage interface with capability flags.

Conflict Model

Explicit conflict objects with three-way merge for text and side-by-side keep-both for binary files.

UI

A minimal React dashboard: transfer graph, live queue, and the conflict-resolution flow.

Testing

A chaos suite that kills processes, corrupts chunks and disconnects networks mid-transfer.

05

Tech stack

The tools that shipped it

Node.jsTypeScriptAWS S3WebDAVReact
06

Key features

What makes it useful

Content-Addressed Sync

Files are chunked and addressed by content hash — renames and moves are detected for free, timestamps are never trusted.

  • Deduplication across providers
  • Moves and renames survive any order of operations
  • Partial-file resume at chunk granularity

Explicit Conflicts

When versions diverge, both are kept and the user sees a single clear choice — never a silent overwrite, never a hidden copy.

  • Three-way merge for plain text
  • Keep-both with clear naming for binaries
  • Conflict history with full context

One Logical Filesystem

Providers mount as branches of one tree; a file can live anywhere and remain findable everywhere.

  • Capability-aware adapters (S3, WebDAV, local)
  • Per-branch bandwidth and retention rules
  • Transfer graph that shows what moves and why
07

Design & UX

How it feels to use

The UI's job was legibility: show the engine's reasoning, not just its progress.

The transfer graph

Every sync decision is drawn as a graph of causes — why this file moved, what triggered the check.

Conflicts as first-class citizens

The conflict view is the most-designed screen in the app; it's the moment trust is won or lost.

Engine-grade honesty

The UI never shows optimistic state the engine hasn't committed — no fictional progress.

08

Technical challenges

And how they were solved

Clocks lie, hashes don't

Timestamp-based sync fails across providers with skewed clocks. Moving all decisions to content hashes eliminated the entire class of 'newer version lost' bugs.

Chaos testing a sync engine

Writing the chaos suite — kill -9 mid-write, truncated chunks, rewound clocks — found more real bugs in a week than months of happy-path use. It remains my favourite artifact of the project.

Provider capability gaps

WebDAV servers disagree about etags, listings and partial writes. Adapters declare capabilities and the engine degrades gracefully — verified integrity where possible, explicit uncertainty where not.

09

What I learned

Honest takeaways

CloudSync was the deepest systems exercise I've set myself, and it changed how I design everything else.

  • Never trust clocks or floating timestamps — content addressing removes a whole category of failure.
  • Write the chaos tests first; they are the spec for distributed behaviour.
  • Conflict UX is where sync products are actually won or lost.
  • Some projects are worth doing for the education alone — this one pays dividends in every system I design.
10

Results

Outcomes where available

CloudSync is an archived experiment, shared as a design reference. No usage metrics exist — by design.

3 mo

Build duration

Evenings and weekends

0

Silent data-loss paths

Verified by the chaos suite

3

Provider adapters

S3, WebDAV, local disk

—

Usage metrics

None tracked; the experiment is its own result

Final takeaway

Building CloudSync taught me…

CloudSync taught me that correctness is a feature you can architect — and that the hardest, most instructive projects are the ones nobody is paying for.

11

Gallery

The tree — one logical filesystem, many providers

Conflict view — one clear question

Transfer graph — why files move

Queue — engine-grade honesty

Keep exploring

Related projects