A playlist and a listening history can both look like lists of songs, but they represent different things. A playlist describes a collection someone intends to play. Listening history records events that have already happened.
That difference changes how the data should be written, queried and corrected. Moving a song within a playlist is an edit to current state. Playing the same song again is another event, not an update that erases the previous listen.
Introduction
Spotify's public playlist API documents versioned playlist behaviour, while its engineering posts describe event delivery and personalisation systems. Together, they provide a useful view of the separate storage responsibilities behind the product.
The public API does not reveal Spotify's complete internal playlist database. We will therefore distinguish documented behaviour from a simplified design that explains how such a system can work. Historical engineering examples are labelled by their publication period rather than presented as a definitive account of today's entire platform.
The central question is not which single database stores Spotify. It is which representation best serves a playlist edit, a recent-history request or an analytical calculation.
A Playlist Stores References and Order
Spotify's playlist documentation describes playlist items and snapshot identifiers associated with versions of a playlist. It also explains how those versions support changes to a collection. Spotify's playlist concepts
At a conceptual level, a playlist contains references to catalogue items, together with an order. Copying a track into a playlist does not require copying the audio file into that playlist's record.
Imagine a playlist called Sunday Run containing tracks A, B and C. The playlist has its own identity and metadata, while the catalogue holds information about the tracks. Changing the playlist name should not rewrite the track catalogue, and changing an album image should not require editing every playlist that references it.
This separation keeps ownership clear. The playlist controls membership and order; the catalogue controls the referenced content's identity and metadata.
An Item Is More Than a Track ID
Our teaching model could give each playlist entry an entry ID, a playlist ID, a track reference, an ordering value and information about when it was added.
Why distinguish the entry from the track? Because the same track can appear more than once in a collection. An operation to remove one occurrence should not necessarily remove every occurrence.
Position is also fragile. Suppose one collaborator inserts a track at the beginning while another removes the item they saw at position five. Without a shared version or stable entry identity, “position five” may now describe a different song.
A reliable editing interface needs to express both intent and context. It should know whether the user is removing a specific entry, removing all occurrences of a track, or changing the order of a particular version. These are different operations even when their buttons look similar.
Versions Help Coordinate Concurrent Edits
Spotify exposes snapshot IDs through its playlist API. These identify playlist versions and provide context for certain update operations. That public behaviour is useful without assuming how the underlying snapshots are physically stored.
In our example, Maya opens version V10 while Leo adds a track, creating V11. Maya's client must not blindly replace the entire playlist with its old copy plus one local change.
One possible service design accepts an operation together with the version it was based on. The server can apply it safely, transform it according to defined rules or reject it with enough information for the client to refresh.
The precise conflict policy is a product decision. Preserving two independent additions is usually desirable. Two people simultaneously moving the same entry may require a deterministic winner. Versioning detects the situation; it does not automatically decide what the user meant.
Reading a Playlist Is a Join Across Responsibilities
To display Sunday Run, our conceptual service retrieves playlist metadata and a bounded page of entries, then resolves the referenced catalogue information.
It should batch those lookups. Performing one network request per track can turn a hundred-item playlist into a hundred sequential waits. A batched request or an appropriate cache reduces repeated work.
Not every reference will necessarily be playable for every listener at every moment. Availability can differ from membership: the playlist may still contain an entry even when playback cannot proceed. The interface should represent that state instead of silently treating a missing playable file as a corrupted playlist.
This illustrates a broader storage principle. A reference expresses a relationship, while the target service determines the current state of the referenced object. Distributed reads must be prepared for those states to change independently.
Listening History Is an Event Stream
Spotify's 2019 event-delivery account describes a Cloud Pub/Sub-based pipeline and EndSong events used for important downstream calculations. It explains separating event types into their own delivery paths and storage destinations. Spotify's cloud event delivery
Conceptually, a playback event might describe a listener, a track, a session, an event time and relevant playback information. Replaying a track produces another event, even though the track identity is unchanged.
This model preserves history. Overwriting one “last played track” field would answer what happened most recently, but could not reconstruct the previous month.
An event stream can feed several views: recent listening, aggregate trends and recommendation features. Those consumers may need different storage layouts and freshness. They should not all force the playback application to wait for their processing before continuing.
Reliable Delivery Must Handle Duplicate Events
Suppose a mobile client sends a playback event and loses the acknowledgement. It retries because it does not know whether the first attempt arrived.
In our example pipeline, both copies could reach the consumer. Counting them independently would inflate listening statistics. A stable event ID lets consumers recognise the same logical event across delivery attempts.
Deduplication needs a defined window and storage strategy. Keeping every event ID in memory forever is not practical, but forgetting IDs too quickly can allow delayed retries to count twice. The acceptable approach depends on retention, delivery delay and the importance of the calculation.
This is distinct from listening to the same track twice. Those should be separate events with separate identities. As with playlist entries, identity should represent the actual operation rather than being guessed from similar content.
Event Time and Arrival Time Answer Different Questions
A listener may play music while disconnected, then reconnect later. The time an event arrives at a server can differ substantially from when the listening happened.
For a history screen, the user generally expects the listening sequence. For pipeline monitoring, the operator needs to know when the infrastructure accepted and processed the event. Keeping both concepts makes those questions answerable.
Clock accuracy complicates the first timestamp. A device clock may be wrong, so the system needs validation and clear rules for implausible values. It should not assume every timestamp is equally trustworthy simply because it uses the same format.
Late events also affect summaries. If a daily aggregate has already been produced, the pipeline must decide whether to revise it, append a correction or define a cutoff. These are correctness decisions, not merely choices about which streaming framework to install.
Recent History and Long-Term Analytics Need Different Views
Fetching one person's recent fifty plays should be a small, targeted read. Calculating an aggregate across a large population is a very different workload.
A teaching design might maintain a recent-history view keyed by listener and time, while storing events in date-partitioned analytical storage for broader processing. Both derive from the same accepted events but organise them differently.
The recent view can be updated incrementally. An analytical job may scan a whole time range and produce a compact result. Forcing both workloads through the same query path can make a large report compete with interactive requests.
Derived views need provenance: which events and versions produced this result? That information helps explain a discrepancy and supports rebuilding the view after a bug. A number that cannot be traced back to its inputs is difficult to trust, even if it loads quickly.
Personalisation Stores Derived Features
Spotify's 2015 personalisation article describes Cassandra stores for user-profile attributes and entity metadata, fed by stream and batch processing. That historical example concerns derived personalisation data; it is not evidence that every playlist or raw listening event lives in Cassandra. Spotify's personalisation storage case study
In a simplified recommendation service, a feature might summarise recent interest in an artist or a tendency to listen at a particular time. Such a feature is smaller and faster to retrieve than recalculating the full history on every request.
Features can have different lifetimes. A recent-session preference should fade, while a longer-term preference may remain useful. Storing both without considering expiry can leave stale signals influencing the experience long after the behaviour changed.
This introduces another distinction: the historical events record what happened, while the feature store holds an interpretation designed for serving a particular model or product.
Deletion and Privacy Cross Several Data Stores
Once listening data feeds several representations, removing it becomes a coordinated lifecycle problem. Deleting one row does not automatically remove cached history or a derived dataset.
For our hypothetical system, deletion handling should identify which stores contain personal records, which derived outputs need updating and how future reprocessing avoids restoring removed data from older inputs.
Access control should also reflect the sensitivity of the information. A public playlist and private listening behaviour are not interchangeable simply because both reference songs. Services should request only the data they need, and operational logs should avoid casually copying full histories.
These are engineering principles for the example, not a description of Spotify's private compliance implementation. The useful lesson is that data classification belongs in the architecture before datasets spread into many consumers.
Recovery and Reprocessing Need Explicit Boundaries
Suppose a bug miscalculates a recommendation feature for a week. Retained source events can allow the team to rebuild the affected feature without asking clients to replay their listening.
But replay is not automatically harmless. A consumer that also sends notifications or updates external systems could repeat those side effects. Reprocessing needs a mode or design that separates recalculating state from performing user-visible actions.
Versioning transformations helps too. A result built with algorithm A should be distinguishable from one built with algorithm B. Otherwise, an engineer comparing two datasets may mistake a deliberate logic change for missing input data.
The same principle applies to playlist restoration. A historical version is useful only if restoring it has a clear meaning for subsequent edits and for references whose catalogue state has changed.
What to Build in a Smaller Music Application
Begin with a relational model for playlists and entries, clear ordering rules and bounded reads. Add version checks before introducing collaborative editing. Test duplicate tracks, simultaneous moves and a lost acknowledgement after a successful edit.
For listening history, give events stable identities and distinguish event time from receipt time. Build a simple recent-history query before adding recommendation infrastructure.
When analytics becomes expensive, move it behind a durable processing boundary rather than making playback wait for reports. Monitor freshness and completeness alongside throughput. A pipeline delivering many events quickly can still be missing an important category.
This approach grows from the product's needs. It avoids copying a large platform's historical stack while overlooking the smaller correctness problems that appear with the very first unreliable mobile connection.
Define What Completeness Means
For a playlist read, completeness may mean returning the requested page and a reliable continuation. For a daily listening report, it may mean including all accepted events before a stated cutoff. Those are different guarantees.
Write the guarantee down before choosing monitoring thresholds. Otherwise, a team can celebrate fast processing while users see missing history, or delay every report indefinitely while waiting for events that may never arrive. A clear boundary makes both incidents and corrections easier to explain.
The Big Picture
Spotify's public material illustrates several distinct kinds of data: ordered playlist state, playback events, analytical history and derived personalisation features. They relate to the same listening experience but require different storage and update rules.
The strongest lesson is to preserve that distinction. An edit changes a collection, an event records an occurrence, and a feature interprets a set of occurrences. Once those responsibilities are clear, versions, queues, indexes and analytical stores each have a concrete purpose.
