Field guide · Working draft
Databases
and storage
Modern software asks an enormous amount of the systems that hold its data. We want to keep entire libraries of video, search years of activity in a moment, and coordinate purchases made on opposite sides of the world. The data can grow far beyond a single machine. A quiet service can face a flood of requests within seconds. Through all of this, we expect it to stay fast and available, even when machines fail or a whole data centre goes offline.
The information itself takes very different shapes. A photograph is a file we want to retrieve intact. A payment may change several records that must be updated together. A stream of sensor readings grows continuously; a sales report may need to examine years of history. A social network holds connections we want to follow, while a search engine must find useful matches among everything it knows. These jobs put pressure on different parts of a storage system.
Our expectations can be uncompromising. A confirmed payment must not disappear. Two people must not buy the same last seat. Someone far from the nearest server still wants an immediate response. We'd like every copy of the data to be current and every request to succeed, even when those copies can't communicate. We can't always meet all of these demands together: checking with another machine takes time, keeping extra copies creates more work, and a network failure can force a choice between waiting and proceeding without agreement.
Databases, filesystems, and object stores have developed around different parts of this problem. Their designs favour particular shapes of data, ways of accessing it, and promises about what happens when it changes. This guide maps those approaches to the jobs they suit, from storing files and running an application's day-to-day transactions to analysing vast histories and following networks of relationships.
Object storage
A collection of photographs, videos, backups, and datasets can grow far larger than the application that manages it. An object store keeps each item under a name, so applications can upload it and retrieve it later. Services such as Amazon S3 let that collection grow independently of the application's servers, with storage classes that trade retrieval speed and access costs against the cost of keeping the data.
The objects can contain almost anything: the store doesn't need to understand a photograph or a backup to keep its contents. This makes it useful for large, varied collections that you retrieve by name. How often you read them matters, though. Cheap storage can come with retrieval charges or a wait for archived content, making a class suited to old backups a poor choice for photographs that must appear immediately on a web page.
The application still needs somewhere to keep information about those objects. An online shop might keep a mug's price and stock count in a database, its photograph in an object store, and the photograph's name in the product record. The application uses both to build the product page.
The same product, two things to keep

- Name
- Blue mug
- Price
- 20 credits
- In stock
- 4
- Photo object name
- mug.jpg
the photo mug.jpg The photograph’s bytes
Filesystems
Keeping a recording for download is one job; editing it is another. A program processing a large recording may need to open it, jump to a particular position, and change a small part without replacing the whole recording. A filesystem gives programs files and directories with operations like these. It's a natural fit when your applications and existing tools expect to work with file paths, read and write parts of files, or rename them.
The files can live on one machine, or a shared filesystem can let several machines work with them. Shared access adds questions about latency and what happens when two programs change the same file. Object stores can also return a range of bytes; the distinction is the set of file operations your software relies on, rather than simply whether it reads an entire file.
Find a file, then change part of its contents

You may also encounter block storage: a disk-like volume that a filesystem or database uses underneath. It supplies storage capacity rather than an application-level query interface.
Further reading: Amazon EFS as an example of a shared filesystem ↗
Relational databases
The shop needs more from its order records than somewhere to keep them. Each order belongs to a customer, and each line in the order refers to a product. A relational database represents these records in tables: one row per customer or product, with columns for details such as a name or price. A query can combine records from different tables to answer a question, such as which products a customer ordered.
It can also help enforce rules as the data changes. A constraint can reject an order line that refers to a nonexistent product. A transaction lets the application record an order and reduce its stock together: if part of that work fails, the database can undo the unfinished transaction. These facilities matter when many people are changing shared information and a half-finished change would leave it wrong.
For the shop's orders and inventory, I'd start here. PostgreSQL, MySQL, SQL Server, and SQLite are familiar examples. They let the application ask new questions about its records without first arranging a separate copy for every question. That flexibility doesn't make every query cheap; a report examining years of orders can still compete with today's checkouts for resources.
Separate records, connected by IDs

Document databases
Some information is useful as a bundle. A mug's product record might contain a description, two available sizes, and a colour; a lamp's might contain a fitting type and wattage. A document database can keep these details inside a single record, including lists or smaller records nested within it. The application can retrieve the mug and its available sizes together.
MongoDB is one example. This arrangement is worth considering when the application usually reads and changes a bundle together, and those bundles have different shapes. The awkward part is shared information. Copy a manufacturer's address into every product and changing that address becomes many updates. Keeping a reference to one manufacturer record avoids those copies, but the application now needs to retrieve that record as well.
Relational databases can also store nested data, often as JSON, the format many applications use to exchange records. If a few parts of your data fit better this way, your existing database may already be enough.
A product with its details inside

Key-value and wide-column stores
Sometimes a request already knows exactly what it needs: the saved session for a signed-in user, or the recent readings from one device. A key-value store associates a key, such as a session ID, with a saved value. Supplying that key retrieves the value.
Some systems also organise groups of records this way. In DynamoDB or the wide-column database Cassandra, a design for sensor readings might group them by device and order each device's readings by time. The device ID takes a request to the right group; the time range selects the readings within it. Different devices' groups can be spread across machines, so requests for different devices can spread the work too.
That is useful when the main queries are predictable and work within these groups. Asking which devices exceeded a temperature yesterday is a different problem: the records needed are now scattered across many groups. You may need another arrangement of the data to answer it efficiently. One very busy device can also concentrate work despite having many machines available. The choice of keys becomes part of the application's design.
“Wide-column” is a confusing name. Here it refers to a family of stores organised around keyed groups; it doesn't mean the analytical column layout below.
A known key takes you to its data

time range D8 · 09:00 18 °CD8 · 09:01 19 °CD8 · 09:02 18 °CD8 · 09:03 20 °C
Analytical databases and column stores
A sales report may need the country and total from millions of orders, without their delivery instructions or other details. A column-oriented layout keeps values from each field together, letting that report leave unrelated fields unread. ClickHouse is an example; analytical engines and warehouses use this idea widely.
Read the fields the report needs

This suits dashboards and reports that count, total, or compare large portions of the data. The design that helps read a few fields across millions of orders isn't necessarily efficient for changing one complete order at a time. Many applications keep their everyday transactions in one database and send a copy to an analytical system. Reports gain a layout suited to their work and their own resources, at the cost of another copy to maintain. A report may lag behind the latest purchase while that copy catches up.
Search and vector retrieval
A customer searching for “blue mug” hasn't supplied a product ID. The system has to find likely products, then decide which to show first. A word index can keep track of which products mention “blue” and which mention “mug”, giving it a route to candidates without reading every product description for each search. Search systems such as Elasticsearch and OpenSearch build on this idea, with facilities for interpreting search terms and ranking results.
Words don't always line up so neatly. Someone asking for a “cup for coffee” may still want the mug. Vector retrieval uses a model to turn a query and each item into lists of numbers, called vectors. It compares those vectors to find items close to the query in that representation. This can find matches beyond the exact words, including similar images, but closeness is no guarantee that a result is useful. You still need to try searches your customers would make and judge the results.
These capabilities can live in a search engine, a specialist service, or an existing database. A separate search system becomes worth considering when finding and ranking useful results is substantial work that the existing database handles poorly. If it holds another copy of the catalogue, product changes and deletions need to reach it too.
Two ways to find a match
Enlarge illustration
Scroll sideways to see the full drawing.

Text description of the illustration
A word index links blue to P9, a blue plate, and P7, a blue mug. It links mug to P7 and P8, a red mug. P7 contains both words in the query blue AND mug. Below, a diamond for the query “cup for coffee” has shorter distance guides to two mug points than to a plate point. All three candidate points use the same colour; no result boundary is drawn.
Go deeper: term indexes, similarity, and which results get returned →
Graph databases
Connections can be the information we most need to explore. Suppose Checkout calls Orders, and Orders calls Inventory. If Inventory fails, Checkout may be affected even though it never calls Inventory directly. To discover that dependency, we have to follow the chain.
A graph database represents things as nodes and their connections as edges. In this case the nodes are services, and each edge means “depends on”. Neo4j is one example. This representation is useful when following paths through a changing network is central to the application: tracing dependencies, investigating connected accounts, or finding routes.
Simply having relationships doesn't require a graph database; relational systems handle them too. The stronger reason is the kind of questions you keep asking about those relationships. Even in a graph database, a densely connected network can make a query explore many paths. A convenient representation doesn't remove that cost.
Who relies on Inventory?
Text description of the illustration
Four service nodes are connected by three arrows. Checkout points to Orders. Orders and Catalogue each point to Inventory. Every arrow means “depends on”.
Go deeper: following dependencies without getting lost in cycles →
Caches and in-memory stores
If thousands of requests need the same product details, rebuilding the answer each time may be wasteful. A cache keeps a copy ready to reuse. Redis and Valkey are in-memory stores often used for this job, though they have other uses too.
In a common arrangement, the application checks the cache first. If the answer is there, it can return it immediately. Otherwise it reads the database, builds the answer, and saves a copy for later requests. This helps when many requests can reuse an answer that is expensive to produce.
The saved answer can become out of date when the source changes. You need to decide when to expire or replace it, and how the application behaves if the cache disappears. A product page might tolerate briefly showing an old price; completing a purchase needs a price the sale can actually honour. Caching adds this question even when looking up the copy is very fast.
A saved answer makes a shorter journey
Enlarge illustration
Scroll sideways to see the full drawing.

Text description of the illustration
Hit: the application checks a cache holding an answer for P7, and the cache returns that answer. Miss: step 1, the application queries the database; step 2, the database returns data to the application; step 3, the application saves its built answer in the cache. The application then returns the answer to its caller. There is no direct connection between cache and database.
How these choices fit together
These groups overlap because they describe different aspects of a system. Relational tables and nested documents are ways to represent information. Column storage is a way to lay it out for reading. A cache is a role a store plays in an application. One product can combine several of these ideas: a relational database may also hold documents and provide text or vector search.
Having a capability and being a good fit for a particular workload are different things. For our shop, a relational database and somewhere to keep photographs may be enough. Before adding a search service, I'd try the database's own search on the queries customers actually make. If reports start delaying checkouts, I'd investigate an analytical copy. Each addition has to earn the work of keeping another system running and its data up to date.
The demands at the start of this guide still matter whichever family you choose. Keeping copies near readers around the world can shorten the journey for a read, but those copies must receive changes. If two sites sell the same last seat, they need a way to coordinate that decision. When a connection fails, the application may have to wait or decline a sale to preserve its promise. A family name doesn't tell you how a particular system handles this.
Once you've found a promising approach, ask what a successful write guarantees, how old a read may be, and which failures the deployment is meant to survive. “Distributed” tells you that data or work spans machines; “managed” tells you someone else operates part of the service. Neither answers those questions on its own. Our introduction to distributed SQL follows one way of keeping transactions correct across machines.
Start with the work your application needs to do. The deeper articles linked here explain the designs behind these choices: how data is arranged, which work that arrangement saves, and where the cost moves instead. Those are ideas you can carry with you when the next database comes along.
Further reading
The links within each section lead to product documentation or a longer explanation on this site. These references offer another useful way into the subject:
- The PostgreSQL tutorial introduces tables, queries, and transactions through a working database.
- MongoDB's data modelling guide discusses when to embed related information and when to keep references.
- Cassandra's data modelling introduction explains why the queries you need shape the way you store records.
Working draft, revised 4 October 2026. Product names are examples, not rankings; suggested starting points are our judgement about the workloads described. Current feature readiness is a separate question, explored in the Postgres vector-search assessment.