Skip to content

Materialised Semantic Graph

create_semantic_graph() turns the vector index into relationships: every indexed node gets edges to its k nearest neighbours, so similarity becomes something you can traverse with Cypher and cluster with community detection.

report = db.create_semantic_graph(
    index="default",
    rel_type="SEMANTIC_SIMILAR",
    k=15,
    min_score=0.1,
)
print(report)  # SemanticGraphReport(742 edges, 100 nodes)

When to Use It — and When Not To

Materialised edges are a cache, and caches go stale

This writes k edges per node into the same table as your domain relationships. At 100k nodes with k=15 that is ~750k rows. Those edges:

  • dominate unqualified traversals. MATCH (a)-[]->(b) now sweeps similarity noise, and every centrality or community result is meaningless unless it excludes them.
  • go stale silently. They reflect the embeddings as of the build. Nothing updates them when a vector changes.
  • need maintenance forever. Every ingest leaves the graph a little more out of date until you rebuild.

Before reaching for this, check whether a query answers the same question:

-- No stored edges, never stale
CALL db.vector.search('default', 'transformers', 15) YIELD node, score
RETURN node, score

SIMILAR() and db.vector.search cover most of what materialised neighbours are used for, and semantic_subgraph() covers most of the rest.

The case that genuinely needs materialised edges is community detection over similarity. Modularity algorithms need edges to exist; there is no way to run Louvain against an ANN index:

db.create_semantic_graph(k=15, min_score=0.3)

for community in db.communities(
    "louvain", rel_types=["SEMANTIC_SIMILAR"], weight_property="score", seed=42
):
    print(community.size, [n.properties["title"] for n in community.nodes[:3]])

Multi-hop similarity patterns are the other — "things like this, and things like those", which an ANN query cannot express:

MATCH p=(a:Doc {id: 'd1'})-[:SEMANTIC_SIMILAR*1..2]-(b:Doc)
RETURN DISTINCT b

Controlling the Edge Count

min_score is the most effective lever — far more so than k, because it cuts the long tail of weak neighbours that inflate the graph without adding signal:

db.create_semantic_graph(k=15, min_score=0.5)   # far fewer edges than the 0.1 default

It defaults to 0.1. That is lower than the SIMILAR() default of 0.5, on purpose: SIMILAR() answers "is this similar?", where a permissive threshold gives wrong answers, whereas here k already bounds the result and min_score only trims the tail.

On an l2 index the default is refused rather than guessed — those scores are negated distances (<= 0), so any positive threshold would build an empty graph:

db.create_semantic_graph(k=15)                  # DatabaseError on an l2 index
db.create_semantic_graph(k=15, min_score=-0.5)  # explicit, and meaningful

symmetrize decides how each node's k-nearest list becomes edges:

Mode Keeps a pair when Relative size
"union" (default) either node chose the other baseline
"mutual" both nodes chose the other much smaller
"directed" one edge per node per neighbour ~2× union

union and mutual emit one edge per pair, so query them with an undirected pattern:

MATCH (a)-[:SEMANTIC_SIMILAR]-(b)   -- note: no arrow

(undirected=True/False is the older shorthand for union/directed and still works; an explicit symmetrize wins.)

Mutual k-NN

Reciprocity removes exactly the edges that do the most damage to clustering. A node sitting at the edge of a cluster gets pulled into distant nodes' neighbour lists without them appearing in its own — those one-sided links are what glue unrelated groups together.

db.create_semantic_graph(k=5, min_score=0.3, symmetrize="mutual")
db.communities("louvain", rel_types=["SEMANTIC_SIMILAR"], weight_property="score")

It trades coverage for precision — raise k with it

Measured on synthetic clusters with deliberate overlap, three planted clusters of six nodes:

k mode edges crossing clusters communities found
3 union 40 40% 3
3 mutual 14 14% 8
5 union 59 44% 3
5 mutual 31 39% 3

At k=3, mutual cuts cross-cluster edges from 40% to 14% — and shatters the three clusters into eight fragments, because too few pairs survive to keep each cluster connected. At k=5 the fragmentation is gone, and so is most of the advantage.

So mutual is not a free improvement: it needs a larger k than you would use with union, and the right value depends on your data. This is what the numbers looked like on one synthetic corpus — measure it on yours.

labels restricts which nodes participate, and max_edges caps the build:

db.create_semantic_graph(k=15, min_score=0.3, labels=["Article"], max_edges=100_000)

The cap is enforced between nodes, so it can overshoot by up to k-1. That is deliberate: a node cut off half way through its neighbours would still look processed — it is the source of an edge — and no later refresh would finish it. Because every node a capped build does process is complete, running refresh_semantic_graph() under the same cap makes progress each time and eventually produces the whole graph:

db.create_semantic_graph(k=15, min_score=0.3, max_edges=50_000)
while db.refresh_semantic_graph(k=15, min_score=0.3, max_edges=50_000).edges_created:
    pass   # each pass links another batch of nodes, completely

Loop on edges_created, not on nodes_processed. A node whose every edge was already contributed by its neighbours has no outgoing edge of its own, so it is re-searched on each pass and produces nothing — a few percent of the corpus, harmless but enough that nodes_processed never reaches zero.

Provenance and Rebuilding

Every generated edge carries score, index, generated_by, and generated_at. Those last two are what make rebuilds safe: replace=True (the default) deletes only edges this method created from the same index, so both a hand-made SEMANTIC_SIMILAR relationship and edges generated from a different vector index survive a rebuild.

db.create_semantic_graph(k=15, min_score=0.3)      # idempotent: re-running replaces
db.drop_semantic_graph()                           # generated edges only
db.drop_semantic_graph(index="papers_vec")         # ...from one index only

replace=False appends instead of rebuilding, and deduplicates against what is already stored — so it adds the edges a wider k discovers without stacking a second copy of the ones already there.

The rebuild is atomic

Neighbours are computed first — the slow part, minutes on a large corpus — and the old edges are swapped for the new ones inside a single short transaction. A failure mid-build leaves the previous graph intact, and concurrent readers never observe a window where the semantic graph is missing.

Incremental Updates

After adding documents, link the new ones without rebuilding everything:

db.index_documents(new_rows, label="Doc")
report = db.refresh_semantic_graph(k=15, min_score=0.3)
print(report)  # SemanticGraphReport(30 edges, 2 nodes, 100 skipped)

refresh_semantic_graph() only processes nodes whose neighbourhood has already been searched — that is, nodes that are the source of a generated edge for this rel_type and index. Three things deliberately do not count as done:

  • hand-made relationships of the same type;
  • edges generated from a different vector index;
  • edges a node merely received, which say nothing about whether its own neighbours were ever computed. A build cut short by max_edges leaves exactly such nodes, and counting them would strand them permanently.

Being strict here can re-search a node whose edges were all deduplicated away. That costs a lookup and creates nothing: the refresh deduplicates against edges already in the database, so running it repeatedly converges instead of accumulating.

It will not notice that an existing node's neighbourhood changed — new documents can be a better match for old ones than what is currently stored. Only a full rebuild fixes that, so schedule one periodically if your corpus keeps growing.

Keeping Analysis Honest

Once these edges exist, every graph analysis needs to say whether it wants them:

# Structure of the domain
db.centrality("pagerank", exclude_rel_types=["SEMANTIC_SIMILAR"])

# Structure of the embedding space
db.communities("louvain", rel_types=["SEMANTIC_SIMILAR"], weight_property="score")

The same applies to subgraph expansion — a single hop through similarity edges reaches a large part of the database:

db.semantic_subgraph("agents", k=20, expand=1, exclude_rel_types=["SEMANTIC_SIMILAR"])

API Reference

grafito.ingest_report.SemanticGraphReport dataclass

What :meth:~grafito.GrafitoDatabase.create_semantic_graph materialised.

edges_created is the number worth watching: it grows as k times the node count, and those edges live in the same table as the domain's own. A build that produced far more edges than expected is usually a min_score that is too permissive.

Source code in grafito/ingest_report.py
@dataclass
class SemanticGraphReport:
    """What :meth:`~grafito.GrafitoDatabase.create_semantic_graph` materialised.

    ``edges_created`` is the number worth watching: it grows as ``k`` times the
    node count, and those edges live in the same table as the domain's own. A
    build that produced far more edges than expected is usually a ``min_score``
    that is too permissive.
    """

    edges_created: int = 0
    edges_removed: int = 0
    #: Nodes whose neighbourhood was computed.
    nodes_processed: int = 0
    #: Nodes skipped because `approximate` found them already linked.
    nodes_skipped: int = 0
    #: True when `max_edges` stopped the build before every node was processed.
    truncated: bool = False

    def __str__(self) -> str:
        parts = [
            f"{self.edges_created} edges",
            f"{self.nodes_processed} nodes",
        ]
        if self.edges_removed:
            parts.append(f"{self.edges_removed} replaced")
        if self.nodes_skipped:
            parts.append(f"{self.nodes_skipped} skipped")
        if self.truncated:
            parts.append("truncated")
        return f"SemanticGraphReport({', '.join(parts)})"