Saving a Pin looks like a small action: choose a board, tap save, and the item appears in your collection. Behind that action are several relationships between people, boards, Pins and media.
The storage design must make those relationships easy to retrieve without duplicating every large image whenever someone organises an idea. It also needs to handle popular content, growing collections and changes that reach several storage locations at different times.
Introduction
Pinterest's engineering team has published accounts of its MySQL sharding architecture and database operations. These are historical case studies, particularly the widely cited 2015 sharding article, rather than a full inventory of Pinterest's current infrastructure.
They provide a useful foundation for understanding a collection-based product. We will connect the documented decisions with an illustrative home-renovation board. The sample records and workflows explain the design space; they are not a reproduction of Pinterest's current private schema.
The central idea is to distinguish an object from the relationships that organise it. A Pin's data, a board's membership and the underlying image do not need to be the same storage record.
Objects Describe Things, Mappings Describe Relationships
Pinterest's published sharding design distinguishes object tables from mapping tables. A mapping can connect a board to its Pins, while object records contain the data needed to describe the entities. Pinterest's MySQL sharding case study
Imagine a board called Kitchen Ideas. The board has its own identity, name and owner. Its saved items refer to content the user wants to collect. The image bytes are a further concern again.
Separating these responsibilities prevents a board rename from requiring changes to every image record. It also lets the application retrieve membership without loading every full object immediately.
The design question becomes explicit: which relationships must be queried in each direction? “Which Pins are on this board?” and “Which boards reference this item?” are different access patterns and may need different supporting indexes or mappings.
IDs Can Help Route Requests to Storage
The historical Pinterest design uses identifiers containing routing information, alongside configuration mapping logical shards to database hosts. It also describes storing flexible object data in JSON within MySQL.
The important distinction is between logical identity and physical placement. A logical shard can move to another machine while its records retain their identities, provided the routing layer knows the new location.
For our teaching system, a request to retrieve Pin P should reach the relevant storage location directly rather than asking every database whether it owns P. That keeps lookup work predictable as the system grows.
Encoding routing information is one approach, not a universal requirement. A directory service or another partitioning scheme can serve the same purpose. Each introduces tradeoffs around identifier design, migration and the amount of routing state the application must maintain.
Loading a Board Is a Two-Stage Read
Pinterest's article describes retrieving Pin IDs from a board mapping and then loading the Pin objects. It also describes caching mappings separately from objects, using different cache technologies for the two roles.
Our Kitchen Ideas page can follow the same conceptual shape: read a bounded page of membership IDs, batch-fetch the corresponding objects and assemble the result in the requested order.
Batching matters. If the application makes one sequential network request per saved item, a board with fifty visible items may spend most of its time waiting. Grouping lookups by destination reduces unnecessary round trips.
The result also needs to tolerate missing objects. A Pin may have been removed or made unavailable after the mapping was read. The service should handle that state explicitly rather than failing the entire board because one referenced item cannot be displayed.
Ordering Needs Its Own Stable Meaning
A board is more useful when its items appear in an intentional order. The order might reflect when items were saved or a user-defined arrangement. It should not depend accidentally on which database host returned first.
Our teaching model could attach an ordering value and stable entry identity to each membership record. Pagination then uses those values rather than relying on arbitrary storage order.
Concurrent changes make this interesting. If a user saves a new item while another device requests the next page, the service should minimise duplicates and unexpected omissions. A cursor anchored to stable ordering fields often behaves more predictably than a large numeric offset.
Manual reordering introduces another tradeoff. Renumbering every item after each move is simple but expensive for large collections. Alternative ordering representations can reduce writes, provided they include a plan for collisions and eventual rebalancing.
Saving an Item Can Touch Several Records
Suppose the user saves Pin P to Kitchen Ideas. A conceptual service checks the user's authority, records the membership and updates any derived counts or views.
If the relationship and a displayed count live separately, one may update before the other. The product should know which record is authoritative. A temporarily stale count is usually easier to tolerate than a missing saved item.
Retries create another issue. The service may save successfully while the response is lost. Repeating the request should not create an unintended duplicate membership if the product defines saving the same item to the same board as one relationship.
A uniqueness constraint or idempotency key can enforce that rule. The choice must follow product semantics: some collection types allow repeated entries, while others do not. Storage should implement the intended behaviour rather than assuming all lists work alike.
Cross-Shard Changes Need Recovery Rules
An object and a relationship may not share the same physical database. A workflow that updates both can therefore fail halfway through.
For our example, the service might create a membership but fail to update a secondary reverse mapping. Later, one query direction sees the relationship while the other does not.
There are several possible responses: co-locate the records, use a supported distributed transaction, or make one representation authoritative and repair the other through durable events. Each trades simplicity, coordination cost and temporary inconsistency differently.
The essential requirement is a known recovery path. A background task cannot repair a missing relationship if there is no durable record of what was intended. “We will retry later” is a design only when the pending work survives the failure that made the retry necessary.
Image Storage Is Separate from Collection Metadata
Pinterest's database-management publication identifies Pins, boards and image metadata among the important data historically stored in MySQL. Metadata is different from the complete image bytes. Pinterest's MySQL operations account
In our teaching architecture, object storage holds image files while the Pin record refers to them and contains useful display metadata. Different image sizes can support previews and detailed views without sending the largest file to every small card.
That separation keeps collection queries compact. Opening Kitchen Ideas should retrieve enough information to lay out the page while image delivery happens through an appropriate media path.
The media lifecycle remains important. An uploaded file may never become part of a completed save, while an old file may still be referenced by valid content. Cleanup must understand references rather than deleting anything that has not been viewed recently.
Caching Objects and Collections Solves Different Problems
An object cache helps when many readers request the same Pin. A membership cache helps when readers repeatedly open the same board. These caches have different invalidation triggers.
Changing a Pin's description affects its object representation. Adding an item to a board affects the board's membership. Invalidating every related board whenever one shared object changes could produce unnecessary work.
Our design can cache the membership IDs independently and fetch current object data when assembling the page. This increases reuse and narrows the effect of edits, although it still requires efficient batched reads.
Cache keys need enough context. A private board must not share an unrestricted rendered response with public readers. Separating public object data from viewer-specific access decisions reduces the chance of a performance optimisation becoming an information leak.
Popular Content Creates Uneven Traffic
One widely shared image can receive far more reads than ordinary Pins. At the same time, a large account or collaborative board can concentrate relationship updates.
Those are different hotspots. Object caching can help repeated reads of one Pin, while it may do little for contention around a heavily edited collection.
For our example service, measure request distribution by object and collection, not only by server average. A shard may look lightly loaded overall while one partition repeatedly reaches its limits.
Scaling decisions should follow that evidence. Adding storage capacity does not automatically improve a narrow read hotspot, and adding more application workers can worsen pressure on a constrained database. The goal is to reduce or distribute the specific work that is concentrated.
Search and Recommendations Use Derived Data
Opening a known board is a direct retrieval problem. Discovering relevant Pins from a phrase or a browsing context needs representations organised for discovery.
In a conceptual system, search can index descriptions and other permitted attributes, while recommendation pipelines produce candidate relationships or features. Neither should become the only durable record that an item exists.
This separation allows the product to rebuild discovery indexes or change ranking logic without rewriting every user's saved collection. It also introduces lag: newly saved or edited content may become discoverable after the primary operation completes.
The service should monitor that freshness and propagate removals. A deleted or restricted item must not remain indefinitely visible because a derived index has not caught up. Candidate retrieval and current eligibility checks are separate steps for a reason.
Private and Collaborative Boards Add Permission State
A collection can have an owner, collaborators and readers with different capabilities. A user who can view a board may not be allowed to add items, reorder them or change its visibility.
Our teaching design should model these actions explicitly rather than using one vague “has access” flag. Permission checks belong on the server for every relevant operation.
Changing a board from public to private affects cached listings, search and previously issued links. The system needs to decide which information can still be served and ensure the final read path respects the new state.
Collaboration also introduces races. Someone may lose edit permission while their browser still has an unsent operation. The server must evaluate current authority when accepting the change rather than relying solely on the permission state the client saw earlier.
Deleting a Save Is Different from Deleting an Object
Removing an item from Kitchen Ideas changes that board's membership. It does not automatically mean every underlying object or image can be deleted from the platform.
Other valid references may remain, depending on the product's content model. A cleanup system must distinguish removing one relationship from removing the shared entity itself.
Conversely, removing an object may require many collection views to stop displaying it. Updating every relationship synchronously could be expensive, so a design may combine background cleanup with serving-time validation.
This is a general lesson about shared data: reference counting, reachability and deletion policy are part of storage correctness. A storage bill can be reduced unsafely if cleanup treats “not visible here” as equivalent to “not needed anywhere.”
What to Test Before Adding More Infrastructure
For our smaller collection service, test a repeated save after a timeout, a deleted object still present in a membership cache and a board becoming private while another device is reading it.
Also test a partial failure between forward and reverse relationship updates. The recovery process should converge without duplicating memberships or losing valid saves.
Performance tests should include a very large board, a heavily read Pin and a cold cache. Uniform random traffic will not reveal the same bottlenecks as concentrated interest in one item.
Begin with clear relational records, indexes and bounded reads. Add sharding only when the workload or isolation requirements justify operating it. Pinterest's historical architecture is most useful as an explanation of tradeoffs, not as a checklist every new product must implement immediately.
The Big Picture
Pinterest's published storage design demonstrates the value of separating objects, relationship mappings and cached representations. A collection is an organised set of references, while media and discovery systems serve additional responsibilities.
For engineers, the most useful questions are which direction a relationship must be queried, how results stay ordered and what happens when an update reaches only some representations.
Answering those questions produces a design that can preserve a user's saved ideas through retries, edits and growth. The storage engine supports that design, but the relationships and lifecycle rules determine whether the product remains reliable.
