Databases / Working draft

Choosing keys for
partitioned data

DynamoDB, Cassandra and ScyllaDBTwo interactive modelsThe database field guide ↗

Birch maintains cold rooms for supermarkets. Each room has a temperature sensor, and each sensor sends a reading every minute. When a customer phones about a warm room, support opens that customer's recent history. A device engineer has a different screen: the history of one sensor, including readings from before it moved to a new customer.

A reading contains a customer, a device identifier, a time, and a temperature. We could put all the readings in one table and search it whenever someone opens a screen. But most screens need a tiny part of the history. We'd like the request to name that part directly.

A partition key gives related records an address. With customer as the key, Birch's readings form one logical group. A sort key orders records inside that group. Time is useful here: after finding Birch, the database can seek to yesterday morning and read forward, leaving older readings behind.

Illustration · Routing a history request
Support requestBirch, since 10:00
Partition key: BirchLocate this logical group
Sort key: timeSeek to 10:00, read forward
The first key identifies a group; the second narrows the rows inside it. A logical group can have replicas and share physical storage with other groups. These three steps are not three machines.
01 / READS

The key starts with the question

Suppose Birch has devices B1 and B2. At 10:00 they report 21 and 22 degrees. The customer-keyed layout puts both readings under Birch. A query for Birch since 10:00 returns those two records without opening Alder's or Cedar's history.

Now use device as the partition key. B1 and B2 have separate addresses. The device engineer can open B1's history directly, but the support request must first know which devices belong to Birch, query each, and combine the results. If that membership list isn't available, finding Birch's records requires searching the data more broadly.

Sorting cannot repair a missing address. Time order helps once we've found a group; it doesn't tell us which device groups contain Birch. Conversely, a good partition key doesn't make every filter cheap. If Birch's records are ordered by temperature, a time filter can require checking records throughout the group. The database may return only two rows after reading many more.

Interactive · Partitioned data

Where does Birch's history live?

Alder, Birch and Cedar each have two devices. Each device sent a temperature at 09:00 and 10:00: twelve readings in total. Both queries ask for readings from 10:00 onward.

Change the partition key while keeping Birch's devices as the query. Notice which groups must be opened. With device keys, enable Known Birch device list to open just B1 and B2. Then change the sort key or add an alternate index.

Controls

Result Logical groups, not server assignments

What this model leaves out

The counts describe examined records and logical requests, not disk reads or latency. An ordered range can seek directly to 10:00. The optional membership list is assumed complete for this history interval, with B1 and B2 belonging to Birch. Without that list or a matching partition key, we model a fanout scan; real APIs may require a Scan operation or reject that query. The extra index projects all twelve readings and is fully caught up. We omit pagination, replication, indexes within files, and filters that run after reading.

DynamoDB makes this distinction explicit. A Query supplies one partition-key value, with an optional sort-key condition. A filter expression runs after that read; it doesn't reduce the items read. Searching without the partition key uses a different operation, Scan. Its pagination and capacity consumption still matter when very few items survive the filter. Key conditions and query and scan guidance describe the boundary.

Cassandra expresses the same modelling idea with a partition key and clustering columns, which order rows inside the partition. A table might use ((customer, day), time, device) as its primary key. Customer and day together identify a group; time and device distinguish its readings. Device breaks ties when two sensors report at the same time. Without enough key fields to identify a unique reading, an insert could replace a previous value. Cassandra's modelling guide follows queries into these choices.

02 / ANOTHER QUESTION

Keeping a second route into the history

Both screens are useful. We could store a customer-keyed history and another arranged by device. The 10:00 reading from B1 then appears in two places. Each layout answers its own question without making the reader discover the other key first.

That second route costs work on writes. An arriving reading has to update both representations. A corrected customer identifier must leave the old customer group and enter the new one. If the application maintains these copies and crashes after updating only one, the screens can disagree. Retrying safely needs a stable event identifier, and recovery needs a way to detect and repair missing copies.

In DynamoDB, a global secondary index can provide the alternate partition and sort keys. The service maintains it from the base table, with a chosen projection of attributes. That saves application bookkeeping but adds storage and write work. Global secondary index reads are eventually consistent, so a successful table write need not appear there immediately. A local secondary index retains the base partition key; it supplies another order within that key, rather than the device route we've been looking for. Secondary index documentation explains both forms.

For a DynamoDB base-table read, an application can request strong consistency. The default is eventual consistency. Cassandra and ScyllaDB instead let a request choose a consistency level specifying how many replicas must respond. These are different contracts, even when the key layout looks similar. In neither case does the word “partitioned” tell us what a reader sees after a write. DynamoDB read consistency and ScyllaDB replica acknowledgements make the product boundaries concrete.

03 / UNEVEN TRAFFIC

A large customer can keep one key busy

Grouping by customer looked convenient when every customer had two sensors. Now Birch has thousands. Their incoming readings all carry the same customer key. Other customers may be quiet while this group is busy. Adding total capacity doesn't ensure that the busiest key can use it.

A day bucket bounds the history in each group: Birch plus 2 October is separate from Birch plus 1 October. This helps keep groups from growing forever and makes expiry easier to organise. It doesn't spread writes that all arrive on 2 October. The current day can still be hot.

We can add a suffix: Birch, today, bucket 0; Birch, today, bucket 1; and so on. The writer chooses a suffix for each event. That produces several addresses which the storage system can distribute. To read all of today's history, support queries every suffix and merges the results by time. Choosing more addresses has made a write problem smaller by making the read procedure larger. DynamoDB documents random and calculated suffixes; the same tradeoff arises in application-defined bucket keys elsewhere.

Splitting a large customer by device can spread its many sensors across groups. But that still leaves each individual device with one current-day address. The next model follows that remaining case: B1 becomes noisy, so we add suffixes within its device-and-day group. The read fanout is for one device on that day.

Interactive · Partitioned data

One device sends most of the readings

A customer-and-day group can be hot because many devices share it. Even after switching to device-and-day keys, one noisy device can keep its own group busy. Here all 100 writes arrive on the same day: B1 sends 70 and the other five devices send six each. A suffix spreads each device's writes evenly across independently addressable logical buckets. The shared day is omitted from the bucket labels.

Choose Four buckets and compare the busiest bucket with the number of requests needed to read all of B1's history. Try an even workload too.

Controls

Result Logical groups, not server assignments

What this model leaves out

A bucket is a logical key, not a dedicated machine or a throughput unit. Real hashing can place several keys together; placement, replication and service limits still matter. Equal suffix distribution is assumed. Reading a device’s history for this day queries every suffix and merges its time order. The counts are invented, not measured capacity.

Hashing spreads distinct keys, not the traffic inside an individual key. And logical buckets aren't promises of separate machines: several may share physical capacity. Bucket design, actual placement, and the hottest customer's behaviour all need checking. A sensor that suddenly sends thousands of readings per second can expose a problem that an average writes-per-second figure hides.

04 / WRITES

Following a reading into Cassandra

The partition key explains where a reading belongs. It hasn't explained how a replica stores it. Cassandra's ordinary write path records a mutation in a commit log and an in-memory memtable. The log supports recovery; the memtable keeps recent data available to reads. When the memtable is flushed, it becomes an immutable sorted file called an SSTable.

A later correction doesn't edit an old SSTable in place. It adds another version through the write path. A read can consult memory and several files, reconciling applicable versions. Background compaction combines files and removes obsolete data where safe. This trades repeated file rewriting for keeping foreground writes simple. Deletes need tombstones so an older copy doesn't reappear; safe cleanup also depends on replicas and repair. Cassandra's storage engine documents this machinery.

ScyllaDB belongs to the same broad CQL, partition-and-clustering-key family, with its own implementation and operational choices. DynamoDB offers similar application-level keys as a managed service; those keys don't establish that its internal write path matches Cassandra's. Buffered writes and compaction are storage-engine choices, not consequences of having a partition key.

Ordinary replicated Cassandra writes also aren't Raft consensus transactions. Requiring a quorum means waiting for enough replica responses. Lightweight transactions use a separate consensus mechanism for conditional operations. DynamoDB has its own transactional APIs with documented scope and restrictions. An invariant involving several records needs that explicit contract, rather than an assumption that nearby keys make the changes atomic. See CQL conditional writes and DynamoDB transactions.

05 / A DIFFERENT SCREEN

What happens when the query changes?

This design suits a stream of readings with known history queries. The writer can calculate an address without looking through existing data, and the reader can ask for a bounded interval in a known group. It is less comfortable for an analyst who invents new filters and joins every afternoon. Each useful new route can demand another index, copy, or expensive scan. A column store addresses a different part of that reporting problem.

Consider one last request: find every device across every customer that was above 25 degrees yesterday. Customer-and-day keys bound the scan to yesterday, but still require visiting all customers. A device index doesn't fix it either: we don't yet know the devices. A separate daily alert stream, maintained as readings arrive, could answer it cheaply, but now corrections and threshold changes need defined behaviour. The right layout follows the actual question, including how often it changes and how fresh its answer must be.

← Return to the database field guide

Sources and model notes

Working draft. Primary documentation linked at the relevant claims was accessed on 2 October 2026. The datasets, bucket counts and projected copies are schematic. They establish no throughput, latency, pricing or production-readiness claim. The history model assumes a caught-up alternate layout and a range seek; it does not simulate either product's API or storage engine.

The teaching references are Sam Who's Load Balancing, Bartosz Ciechanowski's Gears, and Julia Evans on concrete technical scenes. Prose, diagrams and models here are original.