Databases / Working draft
What belongs inside
a document?
A product page is a convenient bundle to read. Keeping that bundle correct turns out to be the more interesting part.
A pottery marketplace sells mugs, plates, and bowls from independent studios. Opening the blue-mug page needs its name, description, dimensions, photographs, and available glazes. Those details usually travel together. A reader doesn't ask for the diameter of every mug just to open this one product.
We could store the product and its related records in tables, then assemble them. A document database offers another convenient starting point: put a nested product record in one document. A collection holds many such documents. MongoDB stores them as BSON, a binary representation with typed fields; the JSON below is a readable sketch of one.
{
"_id": "mug_42",
"name": "Blue mug",
"dimensions": { "height_cm": 9, "diameter_cm": 8 },
"glazes": ["blue", "white"],
"brand": { "id": "brand_7", "name": "Harbour Clay" }
} | ID | Name | Height | Diameter | Brand ID |
|---|---|---|---|---|
| mug_42 | Blue mug | 9 cm | 8 cm | brand_7 |
| Product ID | Glaze |
|---|---|
| mug_42 | blue |
| mug_42 | white |
| ID | Name |
|---|---|
| brand_7 | Harbour Clay |
The useful choice here is the boundary around the product. Nested fields and arrays let the stored shape resemble what the page needs. Mugs can have capacities while plates have diameters, without filling every product with irrelevant fields. That flexibility still needs an agreement about names and types. A page expecting a number cannot sensibly compare 9 with "nine centimetres".
A box that can be read in one visit
The dimensions and glazes are embedded: they live inside the product document. A lookup by mug_42 can return the whole bundle without finding a separate dimension record. A product edit can change several of those fields in one atomic operation. MongoDB's embedding guidance ties this design to data read and changed together.
Embedded brand
Product · mug_42
Dimensions and glazes
brand_7 · Harbour Clay
Referenced brand
Product · mug_42
Dimensions and glazes
brand_id: brand_7
Harbour Clay
Now the studio changes its name to Harbour Studio. If its name is copied into every product, we have to find and change every copy. Renaming the mug alone leaves the tea bowl displaying Harbour Clay. Each product is internally valid, yet the catalogue disagrees with itself.
A reference stores brand_7 in each product and keeps the name in a separate brand document. Rename that one document and subsequent reads can resolve the new name. The page now needs information from two documents. The application can perform the lookups, or MongoDB can combine collections with $lookup. One request can still involve more than one stored record. References are especially useful when shared data changes independently of its parents.
Neither layout makes the relationship disappear. Copying pays for it during updates. Referencing pays for it when reads assemble the result. A cache can reduce repeated lookups, but introduces its own freshness rule. For a rarely renamed brand on heavily visited pages, copying may be reasonable. We would need an explicit way to propagate the rename and tolerate or prevent the temporary disagreement. MongoDB's modeling guidance discusses that read-versus-update tradeoff.
Where does a brand rename have to go?
A pottery catalogue has three products made by Harbour Clay. Each product either embeds a copy of the current brand name or references a separate brand document. The catalogue page must display the brand's current name.
Choose Rename brand to change it to Harbour Studio. With embedded copies, use Update next copy and watch which products still have the old name. Switch to references and rename again: one document changes every resolved product view. Increase the product count to see the extra writes that copying requires.
Controls Changing the layout or product count resets the rename
Result Stored documents and the current product-page answers
Brand document: Harbour Clay. All three embedded product copies also say Harbour Clay. Each product page needs one product document.
All products show Harbour Clay.
What this model leaves out
The source brand document exists in both layouts. A rename writes it once; each propagation writes one product. Result labels resolve the brand afresh from the displayed documents after every step. Reference reads count a product and a brand document, which a join could fetch within one request. These are logical document counts, not disk reads or latency. The model has no cache, replica delay, concurrent renames, failed writes, or multi-document transaction. Embedded copies deliberately become consistent one at a time.
Sometimes a copy is supposed to stay old. An order should keep the product name and price as purchased, even if the catalogue changes later. That's a historical snapshot. The embedded brand on a live catalogue page has a different contract: it is supposed to describe the current studio. Before designing a propagation job, decide whether the copied field is history or a current fact.
Nested data still needs a way to find it
Fetching a known product is straightforward. Browsing all blue-glazed products asks a different question. Having a glazes array in each document doesn't tell the database which documents contain blue. Without a useful index it must inspect the collection.
MongoDB's multikey index gives array elements separate index entries pointing back to their document. The blue mug contributes an entry under blue and another under white. A query for blue can find candidate products through the blue entries, then fetch the needed document fields. The multikey documentation explains this mapping and its restrictions.
Adding amber to the mug's glazes also adds an index entry. Removing blue removes its entry. Changing the product document and maintaining its indexes are both part of the write path. More indexes mean more arrangements to maintain. Nested data makes some reads convenient; it doesn't make arbitrary filters free. MongoDB's index overview covers the ordering and write cost.
The query's meaning needs care too. If variants are embedded as objects with glaze and price, “a blue variant costing less than €20” must test both conditions on the same variant. A blue variant and a cheap white variant don't satisfy it together. In MongoDB, $elemMatch expresses that same-element requirement; its query documentation shows the distinction.
The document is an atomic boundary
An editor changes the mug's height and diameter in one MongoDB update. A reader doesn't see just half that update. Single-document writes are atomic even when they change several fields. But a bulk update of brand copies doesn't make the entire catalogue change atomically: individual documents can become visible at different points. MongoDB's atomicity rules distinguish those cases.
Atomicity doesn't stop a stale editor from overwriting someone else's work. Two whole-document replacements can undo one another. Updating just the height preserves an unrelated description edit, but two stale height updates still conflict. Include an expected revision or current value in the filter to detect that conflict. A failed check asks the editor to reconcile.
For a counter, an atomic increment applies arithmetic to the stored value instead of writing a number computed from an earlier read. Reserving stock needs a condition too: decrement only while stock is positive. Targeting one field reduces the scope of an edit; the operation and its filter determine whether a same-field decision remains safe.
CouchDB makes revisions explicit in its document update protocol. Saving an edit based on an outdated revision can return a conflict; applications need a policy for resolving it. It also supports replication between peers, where competing revisions can coexist. That is a different arrangement from assuming every reader goes through one primary server. The document shape doesn't determine replication or conflict semantics. CouchDB's technical overview describes these choices.
MongoDB supports multi-document transactions when a rule must span several documents. If stock is in a product and the purchase is in an order, reserving stock and creating the order may need such a transaction. As with relational checkout, the update must enforce the stock condition under concurrency. Making the order atomic with the reservation doesn't by itself make an earlier stock read safe. Choosing document boundaries can reduce coordination, but it cannot remove a real cross-document rule.
The storage engine has recovery work underneath those updates. MongoDB with WiredTiger uses checkpoints and a journal to recover changes after a crash. Acknowledgement and replication settings determine what durability and visibility the client has waited for. A document-level atomic update answers which fields change together; it does not answer how many machines have a durable copy. Journaling explains local recovery.
A product document cannot grow forever
Reviews look tempting to embed. The first three are easy to show beside the mug. But every new review extends the same array, and the application may eventually carry a long history just to render one page. Engines impose document-size limits; resource costs can become troublesome well before a hard limit. Embedding every review gives an open-ended collection the lifetime of one product.
A useful split is a bounded review preview inside the product, with the full reviews in another collection indexed by product ID. The page gets its usual preview in one read, while “all reviews” uses a separate query. The preview is a maintained copy, so edits and removals need a propagation policy. MongoDB's unbounded-array guidance describes this subset approach and its costs.
A busy product can also become a write hotspot. If every purchase, review, and page view updates the same document, writers keep meeting at that one boundary. WiredTiger can allow updates to different documents concurrently, but competing writes to one document can conflict and require retries. Splitting independent activity lets the boundary follow the work. WiredTiger's concurrency description explains that document-level behavior.
Flexible fields need care as the application changes. Require stable IDs and valid dimensions, and distinguish old and new document versions when migrating a field. MongoDB's schema validation can enforce types and ranges. Readers need a transition plan too: allowing different shapes in storage doesn't automatically teach old application code how to interpret them.
Choose the boundary from the reads and edits
I'd consider a document database for this catalogue when product pages read bounded bundles and most edits stay inside one product. I'd reference the shared studio identity, keep purchase snapshots, and store the review history separately.
Nested attributes alone don't justify a second database service. PostgreSQL can store and index JSON inside a relational row, alongside ordinary columns. Compare the catalogue's actual nested queries and updates with that arrangement. Include the constraints needed across products, the indexes serving browse pages, cross-product reports, and the operational cost of backups, migrations, and another service. Choose a document engine when its handling of that work earns the extra boundary.
If almost every request crosses products, studios, orders, and inventory, the benefit of assembling one document shrinks. A relational design may express those changing relationships and integrity rules more directly. If the real requirement is ranked text discovery, see search indexes; if photographs dominate the bytes, store the files in object storage and retain their references here. The catalogue doesn't need one storage mechanism for every kind of data.
Suppose the marketplace next adds a studio-wide sale. Should its discount be copied into every product? If every page must reflect the change immediately, a shared discount record avoids updating all those copies. If the page shows a precomputed price, updating that price becomes part of the sale's publication process. The useful question is now familiar: which reads do we want to simplify, and who will keep their stored answer correct?
Sources and model notes
Working draft; researched on 2026-10-02. Linked primary documentation was accessed on that date. MongoDB illustrates embedding, references, multikey indexes, and transactions; CouchDB illustrates a different revision and replication design. These choices aren't universal guarantees of the document family.
The catalogue experiment uses invented products and sequential logical operations. Both layouts retain a source brand document. Copied names are intended to stay current, rather than serve as purchase history. Document counts describe the selected design, not physical page reads, network requests, or measured latency.