When you change one line and push it to GitHub, the result looks like a new version of a folder. Underneath, Git represents that history through connected objects, while GitHub adds hosting, collaboration, permissions and reliability around it.
Understanding those two layers avoids a common misconception: a GitHub repository is not simply a row containing a ZIP file, and every feature visible beside the code is not automatically part of the Git repository.
Introduction
GitHub's engineering publications explain Git's object storage, repository infrastructure and relational databases. These sources describe different layers and different points in the platform's evolution.
We will begin with Git's durable data model, then use a small fictional repository to explain the hosting problems it creates. Where the discussion introduces a possible service design, it is an illustrative model rather than an assertion about GitHub's undisclosed current implementation.
The useful distinction is between immutable repository objects, mutable names that point to them, and application metadata such as issues or access rules. Each has a different lifecycle and requires different consistency decisions.
Git Stores Objects Rather Than Independent Folder Copies
Git's core object types include blobs for file contents, trees for directory entries, commits that connect a tree to history and metadata, and annotated tags. Objects are addressed through identifiers derived from their contents. GitHub's explanation of Git's object store
Consider a repository with README.md and app.py. A tree describes the entries and points to their content objects. A commit points to the root tree and records its relationship to earlier commits.
Changing app.py produces new content and a new tree describing the resulting snapshot. The unchanged README can still be referenced through its existing object. Conceptually, a commit describes a snapshot without requiring a separate physical copy of every unchanged file.
This sharing is possible because existing objects are treated as immutable. A reference to an object keeps the same meaning rather than silently changing when someone edits a file later.
Content Identity Is Different from a File Name
A file name describes where content appears in a tree. The blob stores the content itself. Moving a file can therefore change the directory structure without necessarily creating different file-content bytes.
That distinction is useful when reasoning about deduplication. Two paths can refer to identical content, but their meaning in a project still depends on where they appear. Content identity does not replace the directory model.
It also explains why a commit identifier changes when commit metadata or history changes, even if the visible files look the same. The commit is an object containing more than the source-code snapshot.
For an application built on similar principles, immutable object IDs make caching easier: a cached object's content does not need ordinary in-place invalidation. Mutable names still need freshness rules, however, because the name may later point to a different object.
Branches Are Mutable References into History
In our example, main points to commit C3. Creating feature-login creates another reference, initially pointing to the same commit. It does not require copying the complete repository into a new storage area.
A later commit C4 on the feature branch points back into the existing history. Moving the branch reference makes C4 the branch's new tip. The old objects remain part of the connected history where reachable.
This is why branch updates deserve special attention. Immutable objects can be uploaded safely before the branch becomes visible, but changing the branch determines which history users see under that name.
If two people push competing updates, the host must enforce the rules for the reference transition. A normal push should not silently discard someone else's accepted work. The validity of that transition depends on the previous tip and the requested new tip, not merely on whether the new commit object exists.
Packfiles Reduce the Cost of Many Objects
GitHub's object-store article explains packfiles and delta compression. Related objects can be represented more compactly than a large collection of separate loose files, with indexes helping locate them.
Conceptually, two versions of a large text file may share most of their content. Storing a compact difference against a suitable base can save space. That is a physical representation choice; the logical object still has its own content identity.
Compression creates a read tradeoff. Reconstructing an object may require reading a base and applying a delta. A storage system must balance compactness against the work needed for common reads and transfers.
This distinction between logical and physical storage is broadly useful. Applications should not assume that because a version looks like a complete snapshot, every version occupies a complete independent copy on disk.
Following a Push Through a Hosting Service
Imagine a simplified hosting service receiving a push. It authenticates the caller, checks repository permissions, receives the required objects and validates the requested reference update.
Only after the objects are available should the service expose a branch tip that depends on them. Otherwise, readers could observe a commit whose tree or file contents cannot be retrieved.
The success response needs a clear durability promise. Accepting bytes into one process's memory is different from preserving them across a machine failure. The hosting layer is responsible for turning Git operations into a reliable service.
Background tasks can then update indexes, trigger workflows or notify collaborators. Those derived activities should have their own retry rules. A failed notification must not imply that the code push failed, while a successful notification must not be the only evidence that the repository data was stored.
Hosting Requires More Than Git's Local Data Model
GitHub's historical DGit publication describes a distributed approach to repository storage and replication. It illustrates the additional machinery required to host repositories reliably across machines. It should be read as a historical architecture account, not a complete description of today's deployment. Introducing DGit
In any hosting design, multiple copies introduce coordination questions. Which copy can accept an update? What happens when a replica falls behind? Can a read return an older branch tip, and how is that detected?
Immutable objects simplify parts of replication because the same object should have the same contents everywhere. Mutable references remain the coordination boundary: replicas must agree sufficiently about which commit a branch currently names.
Replication also needs repair. A copy that misses an object or becomes unavailable cannot be assumed healthy merely because it once joined the replica set. Recovery and verification are ongoing responsibilities.
Issues and Pull Requests Are Application Data
A repository's code history and its surrounding collaboration features are related but distinct. Issue text, user permissions and many aspects of pull-request discussions belong to the hosting application's data model.
GitHub's relational-database publication describes partitioning MySQL-backed application data to handle growth. It is evidence that the platform includes substantial relational storage alongside Git repositories. GitHub's relational database partitioning
This explains why cloning a repository should not be treated as a complete export of everything visible on its GitHub pages. The clone follows Git's repository model, while collaboration records have separate interfaces and lifecycles.
Our teaching platform might connect a pull request to source and target repository references, comments and review state. Those records can refer to commits without being stored inside the commit objects themselves.
Searching Code Creates Another Representation
Opening a known file at a known commit is different from finding every occurrence of a symbol across many repositories. Search benefits from indexes organised around text or code features rather than only object identity.
In a conceptual hosting design, an accepted push produces work for an indexing service. The index may briefly lag behind the branch tip, so the product should distinguish repository freshness from search freshness.
Private repositories add an access boundary. Search must not reveal matching lines, filenames or result counts to someone without permission. Checking access only after a result is clicked is too late if the snippet already exposes source code.
When access changes, cached search results also need appropriate handling. This is the same general challenge found in document and messaging systems: derived representations must preserve the permissions of their source data.
Garbage Collection Must Respect Reachability
Objects can become unreachable after operations such as deleting references or rewriting history. Keeping every unreachable object forever wastes space, but removing objects too aggressively can interfere with concurrent work or recovery expectations.
GitHub's garbage-collection article describes improvements involving cruft packs for unreachable objects. This is a storage-management concern beyond the ordinary act of creating a commit. GitHub's garbage-collection engineering
In our example, a maintenance task should understand which objects remain reachable from protected roots and which may still need a grace period. It should not infer “unused” merely because one branch no longer points to an object.
This generalises to many immutable-data systems. Sharing saves space only if cleanup understands references. Deleting an object that another version still needs turns a local cleanup decision into data loss elsewhere.
Large Files Expose the Limits of Ordinary History
Version control works well for many source-code changes, but large binary assets can produce very different storage and transfer costs. A tiny visual change may correspond to a largely different binary file.
For a teaching system, separate large-object storage can keep repository operations manageable by storing references in the ordinary history and transferring the large bytes through another path. That adds a dependency: the reference and the external object must remain available together.
Backup and migration procedures need to include both. A repository containing valid pointers is incomplete for practical use if the referenced assets were not copied.
This is another reason to define the unit of recovery carefully. “We backed up the Git repository” is a narrower statement than “we can restore the complete project experience,” which may also include large objects, releases, issues and configuration.
Reliability Tests Should Focus on Reference Transitions
For our fictional host, useful tests include interrupting a push after objects arrive but before the branch moves, losing the response after the branch update and attempting two competing updates simultaneously.
The first case can leave unused objects that cleanup later handles. The second creates uncertainty for the client, which needs to inspect the remote state before assuming failure. The third tests whether the service enforces its update rules atomically.
Also test a replica that has the new branch reference but lacks a required object. The system should prevent or detect that inconsistency rather than presenting a broken repository to readers.
These cases are more informative than measuring upload throughput alone. They reveal whether the hosting service preserves a coherent graph of history through partial failures.
What Smaller Systems Can Learn
The most reusable idea is separating immutable data from mutable pointers. Immutable records are easier to cache, compare and replicate because their meaning does not change. A small set of mutable references defines the current view.
That pattern can help document versions, build artefacts and configuration histories. It does not eliminate coordination: moving the pointer still needs a clear rule, and garbage collection still needs to understand what remains reachable.
Start with explicit identities and lifecycle rules. Decide what a version includes, which references expose it and what “saved” promises. Then add compression, replication or search according to measured needs.
Copying Git's entire implementation is rarely necessary. Understanding why its object graph behaves predictably is the more valuable lesson.
Separate Integrity from Availability
Content-derived identifiers help detect whether retrieved bytes match the expected object, but they do not guarantee that the object is available when needed. A perfectly valid identifier can still point to data missing from a damaged replica.
Our hosting design therefore needs both integrity checks and recovery mechanisms. These answer different questions: whether the returned content is correct, and whether the service can retrieve or repair it at all. Monitoring only one leaves a significant part of reliability unexamined.
The Big Picture
GitHub combines Git's object-based history with the infrastructure and application data needed for collaborative hosting. Code objects, branch references, relational metadata and search indexes have different responsibilities.
Once those boundaries are clear, familiar behaviours make sense: branches are cheap to create, unchanged content can be shared, pushes need coordinated reference updates and cloning does not capture every collaboration feature.
Reliable storage is not just retaining bytes. It is preserving the relationships that give those bytes meaning, and ensuring readers see a coherent version even when uploads, replicas or background jobs fail.
