Cloud-centric applications have dominated the last decade of software development. Every keystroke, every interaction, every piece of data flows through a remote server before users can see the result. While this model simplifies deployment, it introduces latency, creates single points of failure, and strips users of data ownership. Local-first software architecture inverts this relationship: the device is the source of truth, and the cloud becomes a convenience rather than a requirement.
The local-first movement gained momentum from research at Ink & Switch, but it has since evolved into a practical engineering discipline. Modern CRDT libraries, sync engines, and database tools make it feasible to build applications where data lives on the user's device first, syncs when connectivity allows, and resolves conflicts automatically. This guide explores the core patterns, data structures, and implementation strategies that make local-first architecture work in production.
Understanding the Local-First Philosophy
Local-first software satisfies seven ideals that distinguish it from traditional cloud applications. Data should be fast to access, available across multiple devices, functional without a network connection, supportive of real-time collaboration, durable across decades, secure with user-owned encryption keys, and compatible with existing workflows. These ideals translate directly into architectural decisions.
The fundamental shift is treating the local copy of data as authoritative. In a traditional application, the server holds the canonical state and clients are thin views. In a local-first system, each client holds a full or partial replica of the data it cares about. Changes happen locally and propagate outward. This means reads are always instant, writes never block on network availability, and the application remains fully functional offline.
This approach differs from simple offline caching. An offline cache is a temporary fallback that eventually reconciles with the server's truth. A local-first system treats the local state as truth and uses synchronization protocols to propagate changes between peers. The distinction is philosophical, but it has profound implications for how you design schemas, handle conflicts, and think about data ownership.
Consider what happens when two users edit the same document while offline. In a cache-based system, one user's changes typically overwrite the other's when they reconnect. In a local-first system, both sets of changes are preserved and merged automatically using conflict resolution algorithms. The expand/contract migration pattern shares a similar philosophy: never lose data, always maintain backward compatibility.
CRDTs: The Foundation of Conflict-Free Collaboration
Conflict-free Replicated Data Types are the mathematical foundation that makes local-first software possible. A CRDT is a data structure designed so that concurrent modifications by different users always converge to the same state, regardless of the order in which those modifications are applied. This property, called strong eventual consistency, eliminates the need for consensus protocols or central coordination.
CRDTs come in two fundamental varieties. State-based CRDTs (CvRDTs) synchronize by shipping their entire state to peers, who merge incoming state with their local copy. Operation-based CRDTs (CmRDTs) ship individual operations, which peers apply to their local state. Both approaches guarantee convergence, but they have different trade-offs in terms of bandwidth and complexity.
The simplest CRDT is the G-Counter, a grow-only counter where each node maintains its own count:
// G-Counter: each node tracks its own increment count
class GCounter {
private counts: Map<string, number> = new Map();
constructor(private nodeId: string) {}
increment(amount: number = 1): void {
const current = this.counts.get(this.nodeId) ?? 0;
this.counts.set(this.nodeId, current + amount);
}
value(): number {
let total = 0;
for (const count of this.counts.values()) {
total += count;
}
return total;
}
merge(other: GCounter): void {
for (const [nodeId, count] of other.counts) {
const current = this.counts.get(nodeId) ?? 0;
this.counts.set(nodeId, Math.max(current, count));
}
}
}
The merge function takes the maximum count per node, ensuring that no increments are lost and the result is the same regardless of merge order. This principle extends to more complex structures. An LWW-Register (Last-Writer-Wins Register) attaches a timestamp to each write and resolves conflicts by keeping the most recent value. A G-Set is a grow-only set where elements can be added but never removed, with union as the merge operation.
For real-world text editing, you need sequence CRDTs like RGA (Replicated Growable Array) or the Yjs YATA algorithm. These assign unique, ordered identifiers to each character or block, allowing concurrent insertions and deletions at arbitrary positions without conflicts:
// Simplified concept of a sequence CRDT position identifier
interface PositionId {
clock: number; // Lamport timestamp
nodeId: string; // Unique node identifier
offset: number; // Position within the sequence
}
// Each character in the document has a unique position
interface TextElement {
id: PositionId;
char: string;
deleted: boolean; // Tombstone for deletions
}
// Insertions create new positions between existing ones
// Deletions mark elements as tombstones rather than removing them
// Merge: union of all elements, ordered by position IDs
Sync Protocols and Replication Strategies
Having the right data structures is only half the battle. You also need protocols for efficiently propagating changes between nodes. The choice of sync protocol depends on your network topology, expected latency, and bandwidth constraints.
The simplest approach is full-state synchronization, where nodes periodically exchange their entire CRDT state. This works well for small data sets but becomes prohibitively expensive as data grows. A document with thousands of edits would need to ship the entire history on every sync.
Delta-state synchronization solves this by tracking which changes a peer has already seen. Each node maintains a version vector (a map of node IDs to logical timestamps) that summarizes its state. When syncing, a node sends only the deltas that the peer is missing:
// Version vector tracks what each node has seen
type VersionVector = Map<string, number>;
function computeDelta(
localState: Document,
peerVersion: VersionVector
): Delta[] {
const deltas: Delta[] = [];
for (const change of localState.changes) {
const peerClock = peerVersion.get(change.nodeId) ?? 0;
if (change.clock > peerClock) {
deltas.push(change);
}
}
return deltas;
}
// Sync handshake:
// 1. Peer sends its version vector
// 2. Local node computes and sends delta
// 3. Peer applies delta and updates its version vector
// 4. Reverse direction for bidirectional sync
For applications with real-time streaming requirements, you can layer live updates on top of delta sync. Peers maintain persistent connections (WebSocket, SSE, or WebRTC data channels) and push operations as they happen. The delta sync mechanism serves as a consistency fallback, catching any operations missed during disconnection.
More advanced sync protocols include Merkle-tree based comparison, where nodes exchange hashes of their data tree to identify differences efficiently. This is the approach used by systems like CouchDB and IPFS. The Merkle tree allows nodes to pinpoint exactly which subtrees differ and exchange only the relevant data.
Offline Support and Storage Architecture
True offline support requires more than caching responses. The entire application state must be stored locally in a durable, queryable format. The storage architecture for a local-first application typically involves three layers: an in-memory CRDT document for fast access, a persistent local database for durability, and a remote sync target for cross-device replication.
IndexedDB is the most common storage backend for browser-based local-first applications, though its API is notoriously cumbersome. Libraries like Dexie.js provide a much cleaner interface. For native applications, SQLite is the standard choice, and tools like cr-sqlite add CRDT capabilities directly to the database engine:
-- cr-sqlite extends SQLite with CRDT-aware tables
-- Changes to these tables can be merged without conflicts
-- Create a CRDT-enabled table
SELECT crsql_as_crr('documents');
-- The table works like normal SQLite
INSERT INTO documents (id, title, content)
VALUES ('doc-1', 'Meeting Notes', 'Initial content');
-- Track changes for sync
SELECT * FROM crsql_changes
WHERE db_version > ?last_sync_version;
-- Apply changes from a remote peer
INSERT INTO crsql_changes (table, pk, cid, val, col_version, db_version, site_id)
VALUES (?, ?, ?, ?, ?, ?, ?);
Service workers play a critical role in web-based local-first applications. They intercept network requests and serve cached responses when offline, but more importantly, they can run background sync operations. When the browser detects network availability, the service worker can trigger a sync without the user needing to have the tab open:
// service-worker.js
self.addEventListener('sync', (event) => {
if (event.tag === 'crdt-sync') {
event.waitUntil(syncWithServer());
}
});
async function syncWithServer() {
const localChanges = await getUnsynced();
if (localChanges.length === 0) return;
const response = await fetch('/api/sync', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
clientId: await getClientId(),
versionVector: await getVersionVector(),
changes: localChanges
})
});
const { remoteChanges, serverVersion } = await response.json();
await applyRemoteChanges(remoteChanges);
await updateSyncState(serverVersion);
}
Storage quotas are a practical concern for browser-based applications. Browsers typically allow 50MB to several gigabytes of IndexedDB storage, but the exact limits vary. Applications should implement storage monitoring and graceful degradation when approaching limits. For data-heavy applications, consider offloading older or less-accessed data to the sync server while keeping hot data locally.
Conflict Resolution Beyond CRDTs
While CRDTs handle many conflicts automatically, some application-level conflicts require domain-specific resolution strategies. CRDTs guarantee structural convergence, but they cannot always guarantee semantic correctness. Two users might make individually valid changes that create an invalid combined state from a business logic perspective.
Consider a collaborative spreadsheet where two users concurrently modify cells that feed into the same formula. The CRDT will correctly merge both cell changes, but the formula result might not match what either user expected. For these cases, you need application-level conflict detection and resolution:
interface ConflictResolver<T> {
// Detect if merged state has semantic conflicts
detectConflicts(merged: T, localVersion: T, remoteVersion: T): Conflict[];
// Strategy for resolving detected conflicts
resolve(conflict: Conflict): Resolution;
}
// Example: inventory management conflict resolution
class InventoryResolver implements ConflictResolver<InventoryItem> {
detectConflicts(merged, local, remote) {
const conflicts: Conflict[] = [];
// Both users reserved the last item
if (merged.reserved > merged.available) {
conflicts.push({
type: 'over-reservation',
localChange: local.reserved - merged.available,
remoteChange: remote.reserved - merged.available
});
}
return conflicts;
}
resolve(conflict: Conflict): Resolution {
// First-write-wins for reservations
// Or: queue the later reservation and notify the user
return {
action: 'queue-and-notify',
message: 'Item reserved by another user. Added to waitlist.'
};
}
}
Multi-value registers are another approach to conflict resolution. Instead of automatically picking a winner, the system preserves all concurrent values and presents them to the user for manual resolution. This is similar to how Git handles merge conflicts, letting the user decide which version to keep.
Schema conflicts present a unique challenge. When two peers modify the data schema concurrently, merging the schemas requires careful handling. The expand/contract pattern used in zero-downtime migrations applies here as well: always expand the schema to accommodate new fields without removing existing ones, and contract only after all nodes have upgraded.
Building a Local-First Application with Yjs
Yjs is one of the most mature CRDT libraries for JavaScript, powering collaborative features in applications like Notion, JupyterLab, and several VS Code extensions. It provides high-performance CRDT implementations for common data types and includes ready-made sync providers for various transports.
A typical Yjs setup involves creating a shared document, defining shared types, connecting a sync provider, and binding the data to your UI:
import * as Y from 'yjs';
import { WebsocketProvider } from 'y-websocket';
import { IndexeddbPersistence } from 'y-indexeddb';
// Create a shared document
const doc = new Y.Doc();
// Define shared data structures
const sharedText = doc.getText('main-content');
const sharedMap = doc.getMap('metadata');
const sharedArray = doc.getArray('comments');
// Local persistence — survives page reloads and offline periods
const indexeddbProvider = new IndexeddbPersistence('my-doc-room', doc);
indexeddbProvider.on('synced', () => {
console.log('Local data loaded from IndexedDB');
});
// Remote sync — connects to peers via WebSocket
const wsProvider = new WebsocketProvider(
'wss://sync.example.com',
'my-doc-room',
doc
);
wsProvider.on('status', ({ status }) => {
console.log(`Sync status: ${status}`); // connected | disconnected
});
// Observe changes from any source (local or remote)
sharedText.observe((event) => {
console.log('Text changed:', event.changes.delta);
renderDocument(sharedText.toString());
});
// Make local edits — automatically synced to all peers
function insertText(position: number, text: string) {
sharedText.insert(position, text);
}
// Metadata updates with LWW semantics
sharedMap.set('title', 'My Document');
sharedMap.set('lastEditedBy', currentUser.id);
The power of this approach is that the sync provider handles all the complexity of delta computation, conflict resolution, and reconnection. The application code simply reads from and writes to the shared types, and Yjs ensures convergence across all connected peers.
For applications that need server-side authority, you can run a Yjs document on the server as well. This server-side instance can validate changes, enforce permissions, and serve as a persistent sync hub. The WebAssembly Component Model opens interesting possibilities here, allowing you to run the same CRDT logic on both client and server regardless of the implementation language.
Production Considerations and Trade-Offs
Local-first architecture introduces trade-offs that you need to understand before committing to the approach. Memory consumption is the most immediate concern. CRDTs maintain history metadata (tombstones for deleted elements, version vectors, logical timestamps) that grows over time. A text document with 100,000 edits can consume significantly more memory than the visible text content.
Garbage collection is essential for long-lived documents. Yjs implements a garbage collection mechanism that removes tombstones and compacts the internal representation when all peers have observed the relevant changes. You can trigger this manually or let it run periodically:
// Yjs document compaction
const doc = new Y.Doc({ gc: true }); // Enable GC (default)
// For fine-grained control, snapshot and compact
const snapshot = Y.snapshot(doc);
// Later, compute state at the snapshot point
const restoredDoc = Y.createDocFromSnapshot(doc, snapshot);
// Monitor document size
function getDocStats(doc: Y.Doc) {
const state = Y.encodeStateAsUpdate(doc);
return {
encodedSize: state.byteLength,
deleteSet: Y.encodeStateVector(doc).byteLength,
// Warn if document size exceeds threshold
needsCompaction: state.byteLength > 5 * 1024 * 1024 // 5MB
};
}
Authorization in local-first systems is fundamentally different from server-centric models. Since data lives on the client, you cannot rely on the server to enforce read permissions. Encryption becomes the primary access control mechanism: data is encrypted with keys that are shared only with authorized users. Key rotation and revocation add complexity but are necessary for production systems.
Testing local-first applications requires simulating network conditions that are rare in development but common in production. You need to test concurrent edits from multiple nodes, network partitions of varying duration, slow and unreliable connections, and nodes that rejoin after extended offline periods. Property-based testing is particularly effective for verifying CRDT convergence properties:
// Property-based test for CRDT convergence
// Using fast-check for property generation
import fc from 'fast-check';
fc.assert(
fc.property(
fc.array(fc.oneof(
fc.record({ type: fc.constant('insert'), pos: fc.nat(), char: fc.char() }),
fc.record({ type: fc.constant('delete'), pos: fc.nat() })
)),
(operations) => {
const doc1 = new CRDTDocument('node-1');
const doc2 = new CRDTDocument('node-2');
// Apply operations in different orders
const shuffled = shuffle(operations);
operations.forEach(op => doc1.apply(op));
shuffled.forEach(op => doc2.apply(op));
// Merge in both directions
doc1.merge(doc2);
doc2.merge(doc1);
// Convergence: both documents must be identical
return doc1.toString() === doc2.toString();
}
)
);
Performance monitoring should track metrics specific to local-first systems: sync latency (time from local edit to peer receipt), document size growth rate, merge operation duration, and conflict frequency. These metrics help you identify when a document needs compaction, when sync is falling behind, or when your conflict resolution strategy needs adjustment.
The local-first architecture is not the right choice for every application. Systems that require strong consistency guarantees (financial transactions, inventory management with strict counts) are better served by traditional server-authoritative designs. But for collaborative tools, note-taking applications, creative software, and any system where user ownership of data matters, local-first architecture delivers a fundamentally better experience. The tooling has matured to the point where the implementation cost is reasonable, and the user experience benefits are substantial.