Databases / Working draft

Who depends on
Inventory?

A graph makes connections convenient to follow. The number of connections we follow still matters.

Using Neo4j, with relational and JanusGraph comparisonsDatabase field guide ↗

The Inventory service has stopped responding. Checkout calls it to reserve stock. Storefront calls Checkout to place an order. Even though Storefront never calls Inventory directly, its checkout page may now fail. We want to find the services whose work depends on Inventory, including those several calls away.

A list of service names won't answer that. We need the connections: who calls whom. A graph represents each service as a node, and each call dependency as a relationship between two nodes. Following those relationships gives us a way to discover indirect dependencies without having written down every indirect pair.

The call is a record too

In a property graph, the relationship carries a type and may have properties of its own. Our relationship type is CALLS. An arrow from Checkout to Inventory says that Checkout is the caller. A property such as required: true says that this operation cannot finish without the dependency.

Illustration · An indirect dependency
Storefront
Order page
CALLS
required: true
Checkout
Place order
CALLS
required: true
Inventory
Reserve stock
Storefront is two calls away from Inventory. The outage query starts at Inventory and follows incoming calls, against the arrows, to reach Checkout and then Storefront. The arrows record dependency direction, not the query's travel direction.

Required belongs on the connection because a service may have both required and optional dependencies. Email might try to fetch order details, but still send a generic message if that lookup fails. Labelling the entire Email service “optional” would lose which call has the fallback.

Neo4j stores this kind of property graph: nodes have labels and properties; relationships have a direction, a type, and properties. A query can choose outgoing or incoming relationships, or ignore their direction when that suits the question. Its graph concepts guide defines these pieces. Other graph models, including RDF triples, describe connections differently; this essay follows the property-graph family.

The records describe a potential effect of the outage. They don't prove every service is failing right now. An idle service, a cached answer, or a circuit breaker may change what users observe. Our graph answers “who is connected by these dependencies?” Actual incident impact also needs evidence about the running system.

First, find Inventory

We begin with a known service identifier. A property index can find the Inventory node without scanning all service nodes. Once there, the query needs Inventory's incoming CALLS relationships. Following them reaches Checkout and Warehouse, its direct callers.

That local list of connections is called adjacency. Think of each node having a way to locate its neighboring relationships, with each relationship leading to another node. The exact storage varies by engine. The useful distinction is between finding a starting node by a property and expanding from a node through its connections. An index on service ID helps the first operation; it doesn't make the second operation cost nothing.

Neo4j query plans expose this distinction with operators such as NodeIndexSeek and Expand(All). A plan may locate an indexed service, then expand its incoming calls. The operator descriptions show expansion from a known node. They also help us avoid a misleading promise: using a graph database doesn't guarantee that a query visits only a tiny part of the graph.

A service with a few incoming calls is a cheap place to expand. A shared authentication service might have thousands. Finding its node quickly still leaves all those connections to consider. JanusGraph distinguishes graph-wide indexes from vertex-centric indexes, which help narrow the relationships examined around a busy node. Its indexing explanation shows why global lookup and local traversal need separate attention.

The frontier moves out one hop

After finding Checkout and Warehouse, we can look for their callers. Storefront calls Checkout. Returns calls Warehouse. Those services are two hops from Inventory: two relationships along a dependency path.

The nodes still to expand form a frontier. A breadth-first traversal starts with Inventory, expands its direct callers, then expands that new frontier. It reaches services in order of their shortest distance from the starting point. To find dependencies within three hops, it stops expanding after the third layer.

The query must specify what counts as a usable connection. If we only follow required calls, an optional Email dependency doesn't extend the frontier. Rejecting an unsuitable relationship during expansion can avoid exploring the nodes beyond it. Cypher's variable-length paths support bounded patterns. This query binds a list of relationships to calls and tests each one:

MATCH (failed:Service {id: 'inventory'})
      <-[calls:CALLS*1..3]-(caller:Service)
WHERE caller <> failed
  AND all(call IN calls WHERE call.required = true)
RETURN DISTINCT caller.id;

all() requires every call on a path to be required. caller <> failed excludes the starting service even if a cycle leads back to it. Relationship-list binding and the all() predicate are defined in the Cypher reference. This query asks for the same distinct endpoints as our model; its execution plan need not use the model's breadth-first algorithm.

A limit is part of the question. “Within three calls” deliberately leaves out a service four calls away. It also bounds how far the engine searches. If we need every indirect dependent, a small fixed bound gives an incomplete answer, however quickly it runs.

Our example also has a cycle: Warehouse calls Returns, and Returns calls Warehouse. Starting from Inventory can discover Warehouse, then Returns, then Warehouse again. A reachability traversal uses a visited set to remember services already found. It can inspect that returning relationship without putting Warehouse back in the frontier. The cycle is present in the data; the query simply stops revisiting its nodes.

Interactive experiment

How far does the dependency reach?

Eight services call one another. An arrow points from the caller to its dependency. Inventory has failed; the query walks incoming calls to find services that depend on it, directly or indirectly. A solid call is required; a dashed call is optional and has a fallback.

Choose Advance one layer twice to find Inventory's direct callers, then their callers. Disable Required calls only to include Email and Reports. Add shared calls to see more relationships examined, including routes to services already found. Set the hop limit to 1: the result names the frontier left unexpanded. Choose Warehouse and four hops to see a traversal exhaust its frontier instead. Every query change starts a fresh traversal.

Controls

Result Relationships inspected and services found are counted separately

0Relationships examined
0Relationships followed
0Matched services

Inventory is the failed service. The start node is found; no relationships examined yet. Next frontier: Inventory.

UnexaminedFollowedExamined, optional: skippedOptional call (dashed)

Arrows point from caller to dependency. This query walks against the arrows to find callers. × marks the start; ✓ marks a matched service.

No callers found yet. Advance one layer.

Examined relationships in text

No relationships examined yet.

Inventory found. Ready to expand incoming calls.

What this model leaves out

This is breadth-first reachability on a fixed eight-node graph, using incoming adjacency lists and a global visited set. Each click expands one frontier, counts each inspected incoming relationship, applies the optional-call filter, and queues each newly found service once. The start node is excluded from matches. It retains one shortest path per endpoint and does not enumerate all paths or reproduce Cypher's default relationship-reuse semantics. The hop limit bounds the query. Calls are invented static dependencies, not observed failures; retries, caches, replicas, request traffic, and partial outages are absent. Counts show algorithmic work in this model, not measured engine performance.

Finding services is cheaper than listing every route

Suppose Storefront calls both Checkout and Warehouse. Both call Inventory. There are now two routes from Storefront to Inventory, but Storefront is still one affected service. A query for unique endpoints can discard duplicate discoveries. A query for every dependency path must keep both routes.

This changes the role of the visited set. Globally skipping a node already reached is useful for finding reachable services and one shortest witness path. It would discard valid alternatives if we were listing all paths. Then we need rules about which nodes or relationships a particular path may reuse, plus a stopping condition.

Cypher's default pattern matching disallows reusing a relationship in the same graph-pattern match, while nodes may recur. That differs from our experiment's global node visited set. The relationship-reuse documentation makes those semantics explicit. Choose the rule from the answer you need, rather than treating every graph traversal as the same algorithm.

Branching is the other source of work. In a simple tree with four callers at each step, three hops can generate 4, then 16, then 64 candidate services. Shared nodes and filters may cut that down. Asking for all distinct paths can expand the work again even when the final set of services is small. Those are illustrative counts, not a benchmark; they show why adding one hop can have a large consequence.

“Show me one shortest explanation for each service” asks less than “show me every route.” A shortest path also depends on what shortest means. Fewest calls isn't necessarily lowest latency or greatest reliability. A weighted query needs meaningful relationship weights and an algorithm matching that definition. Neo4j's shortest-path guide describes choices within path queries.

Connections have to stay correct as services change

A release removes Checkout's call to Inventory and adds a call to a new Reservation service. Writing only half that change would leave the graph describing a system that never existed. In Neo4j, the relevant node, relationship, and property changes can be grouped into a transaction, committed or rolled back together. Transaction management explains that boundary.

The write updates more than a line in a drawing. It must maintain relationship access structures, any affected property indexes, and recovery records. Neo4j records changes in a transaction log and uses checkpoints to limit recovery work. Transaction logging and checkpointing describe the durable write path.

Give service identifiers a uniqueness constraint, or the importer may create two Inventory nodes and attach callers to different copies. Constraints protect stored rules; they don't discover an omitted dependency. Neo4j's constraint documentation describes uniqueness and other supported checks.

A long query also needs an isolation promise. Neo4j's ordinary reads use Read Committed isolation; a transaction does not automatically receive a frozen graph for its entire lifetime. Concurrent edits can matter. For an incident report, we might query a versioned snapshot of deployment records, or accept a live view with its timestamp. Concurrent-access rules describe the engine's guarantees. The graph model doesn't make this decision for us.

How much of the application is traversal?

A relational database can store this graph with services and calls tables. An index on a call's target ID finds direct callers. Repeated joins follow a fixed number of hops; a recursive query follows a changing frontier. PostgreSQL has recursive WITH queries and cycle detection for exactly this kind of data. Its recursive-query documentation shows the machinery.

For a small service catalogue and an occasional two-hop report, I'd begin with the database already holding the service records. A graph database becomes more attractive when relationship traversal is central: varied path questions, frequent multi-hop queries, and a team that benefits from expressing those questions directly as graph patterns. We should compare actual plans and workloads before adding another system.

A graph engine still has awkward workloads. A broad query expanding every caller of a high-degree service may touch much of the graph. A calculation summing a few fields over every service is a scan, where analytical layouts may be more useful. Splitting a connected graph across machines can also put network work on traversal steps. Storage and deployment choices remain separate from the convenience of the graph model.

Now imagine asking which services depend on Inventory through required calls only, but have an optional direct call as well. The optional edge doesn't invalidate a required indirect path. We test each candidate path's relationships and return the endpoint if any eligible path reaches Inventory. That is a question about paths, rather than a single label on the service. The distinction is what the graph helps us express.

← Back to the database field guide · Relational transactions · Document boundaries

Sources and model notes

Working draft; researched on 2026-10-02. Linked primary documentation was accessed on that date. Neo4j illustrates property graphs, query expansion, and transactional updates. JanusGraph's versioned indexing explanation supplies a second implementation's distinction between global and local indexes. PostgreSQL supplies the relational comparison.

The service graph is invented. Its experiment implements breadth-first reachability with a global node visited set and one shortest witness per endpoint. It counts logical relationship examinations, not Cypher operations, disk reads, latency, or live incident impact. Endpoint discovery, path enumeration, and engine isolation are deliberately explained as separate choices.