Field guide · Working draft

Databases
and storage

Modern software asks an enormous amount of the systems that hold its data. We want to keep entire libraries of video, search years of activity in a moment, and coordinate purchases made on opposite sides of the world. The data can grow far beyond a single machine. A quiet service can face a flood of requests within seconds. Through all of this, we expect it to stay fast and available, even when machines fail or a whole data centre goes offline.

The information itself takes very different shapes. A photograph is a file we want to retrieve intact. A payment may change several records that must be updated together. A stream of sensor readings grows continuously; a sales report may need to examine years of history. A social network holds connections we want to follow, while a search engine must find useful matches among everything it knows. These jobs put pressure on different parts of a storage system.

Our expectations can be uncompromising. A confirmed payment must not disappear. Two people must not buy the same last seat. Someone far from the nearest server still wants an immediate response. We'd like every copy of the data to be current and every request to succeed, even when those copies can't communicate. We can't always meet all of these demands together: checking with another machine takes time, keeping extra copies creates more work, and a network failure can force a choice between waiting and proceeding without agreement.

Databases, filesystems, and object stores have developed around different parts of this problem. Their designs favour particular shapes of data, ways of accessing it, and promises about what happens when it changes. This guide maps those approaches to the jobs they suit, from storing files and running an application's day-to-day transactions to analysing vast histories and following networks of relationships.

Object storage

A collection of photographs, videos, backups, and datasets can grow far larger than the application that manages it. An object store keeps each item under a name, so applications can upload it and retrieve it later. Services such as Amazon S3 let that collection grow independently of the application's servers, with storage classes that trade retrieval speed and access costs against the cost of keeping the data.

The objects can contain almost anything: the store doesn't need to understand a photograph or a backup to keep its contents. This makes it useful for large, varied collections that you retrieve by name. How often you read them matters, though. Cheap storage can come with retrieval charges or a wait for archived content, making a class suited to old backups a poor choice for photographs that must appear immediately on a web page.

The application still needs somewhere to keep information about those objects. An online shop might keep a mug's price and stock count in a database, its photograph in an object store, and the photograph's name in the product record. The application uses both to build the product page.

Illustration

The same product, two things to keep

A product record points to a separate photograph of a blue mug. Database record Object store
Name
Blue mug
Price
20 credits
In stock
4
Photo object name
mug.jpg
Names
the photo
mug.jpg The photograph’s bytes
The database can answer “which products are in stock?” The object store can return the named photograph. This is one common arrangement; databases can store file contents too.

Go deeper: naming objects and publishing files together →

Filesystems

Keeping a recording for download is one job; editing it is another. A program processing a large recording may need to open it, jump to a particular position, and change a small part without replacing the whole recording. A filesystem gives programs files and directories with operations like these. It's a natural fit when your applications and existing tools expect to work with file paths, read and write parts of files, or rename them.

The files can live on one machine, or a shared filesystem can let several machines work with them. Shared access adds questions about latency and what happens when two programs change the same file. Object stores can also return a range of bytes; the distinction is the set of file operations your software relies on, rather than simply whether it reads an entire file.

Illustration

Find a file, then change part of its contents

A directory branches into recordings and exports. A recording inside the left folder is enlarged into a strip with two portions marked for change. media/ recordings/ exports/ interview.wav Two portions edited
Directories organise files by path. A program can open media/recordings/interview.wav, seek to a position, and change part of the recording. The strip shows logical portions of that file, not physical disk blocks.

You may also encounter block storage: a disk-like volume that a filesystem or database uses underneath. It supplies storage capacity rather than an application-level query interface.

Further reading: Amazon EFS as an example of a shared filesystem ↗

Relational databases

The shop needs more from its order records than somewhere to keep them. Each order belongs to a customer, and each line in the order refers to a product. A relational database represents these records in tables: one row per customer or product, with columns for details such as a name or price. A query can combine records from different tables to answer a question, such as which products a customer ordered.

It can also help enforce rules as the data changes. A constraint can reject an order line that refers to a nonexistent product. A transaction lets the application record an order and reduce its stock together: if part of that work fails, the database can undo the unfinished transaction. These facilities matter when many people are changing shared information and a half-finished change would leave it wrong.

For the shop's orders and inventory, I'd start here. PostgreSQL, MySQL, SQL Server, and SQLite are familiar examples. They let the application ask new questions about its records without first arranging a separate copy for every question. That flexibility doesn't make every query cheap; a report examining years of orders can still compete with today's checkouts for resources.

Illustration

Separate records, connected by IDs

Four table records. The order refers to Ada’s customer record; the order line refers to that order and to the blue mug product. Customers tableOrders tableAdaC4O12 · C4C4O12P7Products tableOrder linesP7Blue mugO12 · P7Quantity: 2
Ada (C4) placed order O12. Its order line refers to product P7, the blue mug. Shared IDs let a query connect these records without copying the customer and product details into every order. The sheets represent tables; only one record from each is drawn.

An introduction to relational databases →

Document databases

Some information is useful as a bundle. A mug's product record might contain a description, two available sizes, and a colour; a lamp's might contain a fitting type and wattage. A document database can keep these details inside a single record, including lists or smaller records nested within it. The application can retrieve the mug and its available sizes together.

MongoDB is one example. This arrangement is worth considering when the application usually reads and changes a bundle together, and those bundles have different shapes. The awkward part is shared information. Copy a manufacturer's address into every product and changing that address becomes many updates. Keeping a reference to one manufacturer record avoids those copies, but the application now needs to retrieve that record as well.

Relational databases can also store nested data, often as JSON, the format many applications use to exchange records. If a few parts of your data fit better this way, your existing database may already be enough.

Illustration

A product with its details inside

One product document encloses two mug sizes and a blue colour swatch. Blue mugP7 Sizes250 ml350 ml ColourBlue
The blue mug is one document containing a name, two sizes, and a colour. The enclosing shape shows which details belong to that record. Relational databases can store nested JSON too; this is a representation, not an exclusive capability.

Go deeper: what belongs inside a document? →

Key-value and wide-column stores

Sometimes a request already knows exactly what it needs: the saved session for a signed-in user, or the recent readings from one device. A key-value store associates a key, such as a session ID, with a saved value. Supplying that key retrieves the value.

Some systems also organise groups of records this way. In DynamoDB or the wide-column database Cassandra, a design for sensor readings might group them by device and order each device's readings by time. The device ID takes a request to the right group; the time range selects the readings within it. Different devices' groups can be spread across machines, so requests for different devices can spread the work too.

That is useful when the main queries are predictable and work within these groups. Asking which devices exceeded a temperature yesterday is a different problem: the records needed are now scattered across many groups. You may need another arrangement of the data to answer it efficiently. One very busy device can also concentrate work despite having many machines available. The choice of keys becomes part of the application's design.

“Wide-column” is a confusing name. Here it refers to a family of stores organised around keyed groups; it doesn't mean the analytical column layout below.

Illustration

A known key takes you to its data

A session ID points to one saved session. Below it, the middle two readings in an ordered device record are marked as a time range. Session IDs42 Saved sessionAda · 2 blue mugs Ordered keys: device + time Device D809:01–09:02 Read this
time range
D8 · 09:00 18 °CD8 · 09:01 19 °CD8 · 09:02 18 °CD8 · 09:03 20 °C
A session key retrieves one saved value. A device key can instead locate a group whose readings are ordered by time. Range reads within that group require an ordered-key design; a basic key-value store does not necessarily provide them.

Go deeper: choosing keys for a device's history →

Analytical databases and column stores

A sales report may need the country and total from millions of orders, without their delivery instructions or other details. A column-oriented layout keeps values from each field together, letting that report leave unrelated fields unread. ClickHouse is an example; analytical engines and warehouses use this idea widely.

Illustration

Read the fields the report needs

Three separate strips store country, total, and delivery notes. Only country and total feed the report below; delivery is left unread. Stored together by field CountryTotalDelivery SwedenUKSweden201520Not readSweden: 40 · UK: 15
Entries at the same height in the three strips belong to the same order. The report reads country and total, leaving delivery notes unread. It still includes every order: Sweden totals 40 and the UK totals 15. The drawing shows logical column layout, not measured disk pages.

This suits dashboards and reports that count, total, or compare large portions of the data. The design that helps read a few fields across millions of orders isn't necessarily efficient for changing one complete order at a time. Many applications keep their everyday transactions in one database and send a copy to an analytical system. Reports gain a layout suited to their work and their own resources, at the cost of another copy to maintain. A report may lag behind the latest purchase while that copy catches up.

Go deeper: how column storage reduces the work of a query →

Graph databases

Connections can be the information we most need to explore. Suppose Checkout calls Orders, and Orders calls Inventory. If Inventory fails, Checkout may be affected even though it never calls Inventory directly. To discover that dependency, we have to follow the chain.

A graph database represents things as nodes and their connections as edges. In this case the nodes are services, and each edge means “depends on”. Neo4j is one example. This representation is useful when following paths through a changing network is central to the application: tracing dependencies, investigating connected accounts, or finding routes.

Simply having relationships doesn't require a graph database; relational systems handle them too. The stronger reason is the kind of questions you keep asking about those relationships. Even in a graph database, a densely connected network can make a query explore many paths. A convenient representation doesn't remove that cost.

Illustration

Who relies on Inventory?

Four service nodes are connected by three arrows. Checkout points to Orders. Orders and Catalogue each point to Inventory. Every arrow means “depends on”.
Each arrow means “depends on”. Orders and Catalogue depend directly on Inventory. Checkout depends on Orders, so it also relies on Inventory indirectly. Following incoming edges from Inventory uncovers those dependencies; reversing the question changes the path.
Text description of the illustration

Four service nodes are connected by three arrows. Checkout points to Orders. Orders and Catalogue each point to Inventory. Every arrow means “depends on”.

Go deeper: following dependencies without getting lost in cycles →

Caches and in-memory stores

If thousands of requests need the same product details, rebuilding the answer each time may be wasteful. A cache keeps a copy ready to reuse. Redis and Valkey are in-memory stores often used for this job, though they have other uses too.

In a common arrangement, the application checks the cache first. If the answer is there, it can return it immediately. Otherwise it reads the database, builds the answer, and saves a copy for later requests. This helps when many requests can reuse an answer that is expensive to produce.

The saved answer can become out of date when the source changes. You need to decide when to expire or replace it, and how the application behaves if the cache disappears. A product page might tolerate briefly showing an old price; completing a purchase needs a price the sale can actually honour. Caching adds this question even when looking up the copy is very fast.

Illustration

A saved answer makes a shorter journey

Hit: the application checks a cache holding an answer for P7, and the cache returns that answer. Miss: step 1, the application queries the database; step 2, the database returns data to the application; step 3, the application saves its built answer in the cache. The application then returns the answer to its caller. There is no direct connection between cache and database.
Enlarge illustration

Scroll sideways to see the full drawing.

Hit: the application checks a cache holding an answer for P7, and the cache returns that answer. Miss: step 1, the application queries the database; step 2, the database returns data to the application; step 3, the application saves its built answer in the cache. The application then returns the answer to its caller. There is no direct connection between cache and database.
The application first checks for a saved answer. A hit returns it from the cache; a miss takes the longer route to the database, then saves a copy for later requests. The shorter path avoids another database read, but the copy needs to be expired or replaced when it becomes stale.
Text description of the illustration

Hit: the application checks a cache holding an answer for P7, and the cache returns that answer. Miss: step 1, the application queries the database; step 2, the database returns data to the application; step 3, the application saves its built answer in the cache. The application then returns the answer to its caller. There is no direct connection between cache and database.

Go deeper: how long does an old price stay in the cache? →

How these choices fit together

These groups overlap because they describe different aspects of a system. Relational tables and nested documents are ways to represent information. Column storage is a way to lay it out for reading. A cache is a role a store plays in an application. One product can combine several of these ideas: a relational database may also hold documents and provide text or vector search.

Having a capability and being a good fit for a particular workload are different things. For our shop, a relational database and somewhere to keep photographs may be enough. Before adding a search service, I'd try the database's own search on the queries customers actually make. If reports start delaying checkouts, I'd investigate an analytical copy. Each addition has to earn the work of keeping another system running and its data up to date.

The demands at the start of this guide still matter whichever family you choose. Keeping copies near readers around the world can shorten the journey for a read, but those copies must receive changes. If two sites sell the same last seat, they need a way to coordinate that decision. When a connection fails, the application may have to wait or decline a sale to preserve its promise. A family name doesn't tell you how a particular system handles this.

Once you've found a promising approach, ask what a successful write guarantees, how old a read may be, and which failures the deployment is meant to survive. “Distributed” tells you that data or work spans machines; “managed” tells you someone else operates part of the service. Neither answers those questions on its own. Our introduction to distributed SQL follows one way of keeping transactions correct across machines.

Start with the work your application needs to do. The deeper articles linked here explain the designs behind these choices: how data is arranged, which work that arrangement saves, and where the cost moves instead. Those are ideas you can carry with you when the next database comes along.

Further reading

The links within each section lead to product documentation or a longer explanation on this site. These references offer another useful way into the subject:

Working draft, revised 4 October 2026. Product names are examples, not rankings; suggested starting points are our judgement about the workloads described. Current feature readiness is a separate question, explored in the Postgres vector-search assessment.