Two designers can move shapes, change colours and edit text in the same Figma file while watching each other's work. The file feels immediate and shared, yet it must also survive disconnected browsers, server restarts and mistakes that require an older version.
That combination creates several kinds of state: what each client currently sees, the server's accepted document, durable recovery data and the history exposed to users. They are related, but they are not interchangeable.
Introduction
Figma has published explanations of its multiplayer model, autosave and the journal added to improve reliability. These are valuable architectural case studies, although they describe particular versions of the system rather than a complete inventory of today's infrastructure.
We will connect those documented components with a simple design-file example. The example operations, sequence numbers and recovery walkthroughs explain the principles; they are not a claim to reproduce Figma's exact current protocol.
The central idea is that collaborative editing needs both a way to agree on changes and a way to retain accepted work. Solving one does not automatically solve the other.
A Design File Is Structured Data
Imagine a small mock-up containing a frame, a button and a text label. The editor needs to know each object's identity, properties and place in the hierarchy. A screenshot cannot preserve that editable structure.
Our teaching model might give the button an ID, a parent frame ID, position, size and fill colour. The label has its own identity and text properties. Moving the button changes a property rather than replacing the whole document with a new image.
Stable identities make collaboration more precise. If one designer changes the button's colour while another edits its label, the system can recognise that they targeted different properties or objects.
The hierarchy matters too. Moving an object into another frame changes a relationship, which may affect layout and rendering. Collaborative storage therefore includes both values and structural invariants.
The Server Orders Shared Changes
Figma's 2019 multiplayer explanation describes a client-server design using WebSockets, with document state synchronised through a server. Its model handles object properties and structural relationships, drawing inspiration from CRDT ideas without simply using a generic document CRDT. Figma's multiplayer design
In our example, Aisha changes the button colour while Tom changes its width. Independent property changes can coexist. If both change the same colour property, the system needs a deterministic rule about which accepted update wins.
A central ordering point makes that rule easier to express, but clients still need to reconcile their optimistic local views with the accepted sequence. Network arrival order and user intention are not always the same thing.
This is why collaboration algorithms should be judged by the product behaviour they produce. Eventual agreement is necessary, but users also need predictable results when their edits overlap.
Local Responsiveness Is Not Durable Storage
The editor can apply a local change immediately so dragging an object feels smooth. Waiting for a complete server round trip for every pointer movement would make interaction sluggish.
However, an object moving on screen does not prove the change has been stored durably. The update may still be waiting to leave the browser or waiting for the server to persist it.
Our conceptual interface therefore has at least three states: a local edit, an accepted shared edit and a durably retained edit. A production product may hide much of that complexity, but its saved-state indication must correspond to a meaningful guarantee.
If the browser closes while work is pending, recovery depends on whether those pending operations were retained locally. If the server restarts, recovery depends on what it wrote outside its own memory. These are separate failure domains.
Checkpoints Capture a Whole Document State
Figma's reliability article describes periodic checkpoints: compressed binary document state stored in S3. It explains the weakness of relying only on periodic checkpoints, because changes accepted since the last checkpoint could be vulnerable to a server crash. Figma's multiplayer reliability architecture
To understand a checkpoint, imagine taking a complete recoverable snapshot after operation 500. Loading that snapshot should reproduce the document as it stood at that point without replaying its entire lifetime.
Checkpoints make startup and recovery bounded, but creating them costs work. Serialising and writing a large design file after every tiny movement would be wasteful.
The interval is therefore a tradeoff between checkpoint overhead and the amount of newer work that must be recovered through another mechanism. A reliable design should make that tradeoff explicit rather than assuming frequent snapshots alone eliminate all data-loss windows.
A Journal Records Changes Between Checkpoints
The same Figma article describes adding a durable journal backed by DynamoDB, with sequence numbers linking changes to checkpoints. Recovery can load a checkpoint and apply later journal entries. The publication also describes batching changes before writing them.
In our example, the checkpoint represents sequence 500. The journal contains 501 through 540. After a restart, the service loads the checkpoint and replays those later operations in the defined order.
The sequence boundary prevents double application. Replaying operation 499 would repeat a change already included in the checkpoint. Skipping 501 could omit accepted work.
The journal is therefore more than an unstructured log message saying “button moved.” It must contain enough information, with the right ordering and identity, to reconstruct valid state. Diagnostic logs and recovery journals serve different purposes and should not be confused.
Recovery Needs Rules for Incomplete Work
Suppose the service writes a checkpoint file but crashes before recording that the checkpoint is ready. Or it records a new checkpoint reference before the file is fully available. Those two orderings create different recovery risks.
A teaching design can publish a checkpoint reference only after the snapshot is durably written and verified. Until then, the previous checkpoint remains the recovery starting point.
Journal cleanup must follow the same discipline. Entries should not be deleted merely because a checkpoint attempt started. They can be retired only when a suitable durable checkpoint safely covers them and the retention policy allows it.
These principles apply beyond design tools. Any system using snapshots plus a log must define the handover between them. The recovery path should remain correct even when a crash occurs at the least convenient moment in that handover.
Batching Balances Throughput and Durability Delay
Interactive editing can produce many small operations. Writing each separately to durable storage can add substantial overhead, while batching combines work into fewer writes.
The tradeoff is delay. An operation waiting in an in-memory batch is not protected by the destination store yet. The system must decide what it can acknowledge during that interval and how other copies, such as clients, contribute to recovery.
For our example service, monitor the oldest unpersisted operation and batch-write latency, not simply the number of successful requests. A low average can hide occasional long pauses during which more work is exposed.
Batching should also preserve useful boundaries. Combining operations must not erase the information needed for ordering, deduplication or reconstruction. Saving bandwidth is valuable only if the resulting representation still supports reliable recovery.
Version History Is a User-Facing Promise
A recovery checkpoint and a version shown to a user have different purposes. The first helps infrastructure reconstruct state. The second helps a person understand and restore meaningful stages of their work.
Figma's autosave article describes creating version-history checkpoints around applying recovered differences, providing a way to revert an undesirable result. That is a concrete example of connecting recovery behaviour to a usable product feature. Figma's autosave engineering
In our teaching editor, a named version such as “Approved checkout” should remain identifiable even as background checkpoints continue. It may refer to a specific document state plus author and timestamp metadata.
Restoring it also needs defined semantics. A safe conceptual approach is to make restoration a new change in the ongoing history, so the system preserves what happened rather than pretending later edits never existed.
Offline Edits Need Reconciliation, Not Blind Replacement
Imagine Aisha disconnects and changes several screens while Tom continues online. When Aisha reconnects, uploading her entire old document as the new truth would erase Tom's accepted work.
Figma's multiplayer account describes reconnecting by obtaining current document state and reapplying offline edits. The general challenge is to combine local intent with changes that occurred elsewhere.
Some conflicts are easy to identify but hard to interpret. If one person deleted a frame while another added content inside it, both actions may be locally reasonable. The merge policy must decide how to represent the result and when users need a way to recover alternatives.
This is why “the algorithm converges” is not the whole user experience. A mathematically consistent result can still surprise a designer. Version history and clear recovery affordances help people resolve situations where software cannot infer their intentions reliably.
Document Data and Account Metadata Differ
Figma's multiplayer publication distinguishes collaborative document state from other information such as users, teams and projects, which it described as stored in PostgreSQL through a separate system.
That separation follows the workload. Moving a shape repeatedly is different from changing who belongs to a team. The former needs responsive collaborative updates; the latter needs dependable access and account rules.
Our example editor would check permissions before admitting a client to a document session and again when relevant access changes. A long-lived connection should not imply permanent authority after access has been revoked.
Metadata can also influence storage routing and ownership. A file moving between projects should not require pretending that every shape was newly created. Stable document identity lets those administrative relationships change independently of the editable content.
Assets Have Their Own Lifecycle
A design file may refer to images or other binary assets. In a conceptual architecture, the structured document holds references and placement information, while a separate object store holds the bytes.
That avoids embedding repeated copies of a large image in every operation that changes its position. It also allows clients to fetch assets independently of the document's structural state.
The complication is reachability. An image removed from the current canvas may still be required by a retained historical version. Deleting it immediately could make version restoration incomplete.
Cleanup must therefore understand the retention model across current and historical states. This is a recurring pattern in versioned systems: data that appears unused in today's view may still be required to honour yesterday's saved version.
Scaling Active Documents Is Different from Storing Old Files
A rarely opened file mainly consumes durable storage. A file with many active collaborators consumes memory, connections and processing capacity for each stream of updates.
Those dimensions should be measured separately. Total stored bytes do not predict the load of a particularly active document, and total user accounts do not describe how many people are editing the same object at once.
For our teaching service, useful limits include maximum operation size, bounded work per update and safeguards against pathological document structures. A single malformed or extremely expensive operation should not freeze every collaborator's session.
Recovery capacity matters too. Restarting many active documents simultaneously can create a surge in checkpoint reads and journal replay. A system that works during steady state still needs a plan for that concentrated recovery workload.
Testing a Collaborative Storage Design
Test a server crash just before and just after a journal write, a checkpoint interrupted during upload and a client that resends operations after reconnecting. Stable operation identities should prevent duplicate application.
Also test two people moving the same object, deleting a parent during a child edit and restoring an old version that references assets no longer visible in the current file.
These scenarios examine both durable correctness and user intent. A file that reloads without errors may still have lost a meaningful edit or restored the wrong ordering.
For a smaller editor, begin with explicit object identities and a simple server-authoritative model. Add complex offline behaviour only when you can describe how its pending operations interact with current permissions, deletions and version history.
The Big Picture
Figma's published designs show how collaborative state, checkpoints, journals and user-facing history can support one editing experience. Each answers a different question: what is accepted now, what survives a failure, and what can a person restore?
For engineers, the central lesson is to separate responsiveness from durability without allowing them to drift apart. Fast local interaction needs a clear route into accepted shared state, and accepted state needs a tested route into recovery.
When those boundaries are explicit, collaboration can remain fluid even though the storage system is handling ordering, persistence and failure behind the scenes.
