Benchmarks

The numbers, and exactly how they were measured.

Measured on engine 0.9.15, on 2026-08-25, against a graph pinned to exactly 3,000,000 edges — the same graph for every engine on this page, which is what makes the comparison below mean anything. Every chart is drawn from that run's own output, and the raw per-engine JSON is public.

64,702/s
Read queries1
0.51 ms
2-hop p952
118,041/s
Ingest edges3
106 B
On-disk / edge4
  1. 1Read-only throughput: 40 concurrent clients replaying a seeded workload for 60 s. The mixed run (10% writes) sustained 49,777/s, and the write-only run 25,668/s. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.
  2. 2p95 latency for a two-hop, friends-of-friends traversal against the loaded graph (p50 0.25 ms). Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.
  3. 3Edges merged per second, sustained while loading the graph over Bolt in batches, alongside 224,519 nodes/s. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.
  4. 4On-disk store size divided by edges loaded (304 MB for 3,000,000 edges), read from the container's cgroup counters rather than self-reported. A compact store keeps the working set cached. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.

Latency

Per-query latency, not just an average.

A median hides the tail that users actually notice, so every query shape is reported at p50, p95 and p99 across the measured reps.

Point lookup · read mixlower is better · ms
p50
0.38 ms
p95
1.79 ms
p99
3.19 ms
MATCH (n:SocialUser {id:$x}) RETURN n.id
1-hop expand · read mixlower is better · ms
p50
0.39 ms
p95
1.83 ms
p99
3.24 ms
MATCH (:SocialUser {id:$x})-[:FRIEND]->(m) RETURN count(m)
2-hop expand · read mixlower is better · ms
p50
0.41 ms
p95
2.40 ms
p99
5.08 ms
MATCH (:SocialUser {id:$x})-[:FRIEND]->()-[:FRIEND]->(m) RETURN count(m)
Property write · write mixlower is better · ms
p50
2.10 ms
p95
2.81 ms
p99
5.04 ms
MATCH (n:SocialUser {id:$x}) SET n.v = $v
Edge merge · write mixlower is better · ms
p50
0.32 ms
p95
0.56 ms
p99
0.82 ms
MATCH (a {id:$x}) MATCH (b {id:$y}) MERGE (a)-[:FRIEND]->(b)

Throughput

What it sustains, and while writing.

Read-only is the easy number. The mixed and write-only mixes are the ones that show what contention costs.

Queries per second · 40 concurrent clientshigher is better · queries/s
Read-only
64,702
Mixed (10% writes)
49,777
Write-only
25,668

11 errors across the measured reps.

Sustained edge ingest, across the load windowedges/s
Sustained edge ingest, across the load window — edges per second against window038k75k113k150k123456window
Values
windowedges/s
1124,191
2120,439
3119,876
4114,885
5115,434
6114,088

Ingest rate per window while loading the graph over Bolt. A flat line is the point: the rate holds as the store grows rather than decaying, because folding runs alongside the load instead of stopping it.

Cross-engine

Five engines, one identical graph.

Every engine below loaded exactly 3,000,000 edges and 663,230 nodes on its own dedicated n2-standard-8, with the load generator on a second machine one hop away. Each ran its own stock configuration plus an index on the node id. Query equivalence is verified, not assumed: replaying the identical workload over the identical graph must return the same results, and every engine here agrees with the Neo4j reference to within 1.18%.

Fastest durable ingest

118,041/s

1.11× Memgraph, the next best

Fastest reads

64,702/s

1.43× Memgraph, the next best

Fastest mixed

49,777/s

1.14× Memgraph, the next best

Densest on disk

106 B/edge

4.65× Neo4j, the next best

Four of the six measured columns. The other two go to Memgraph — writes and peak memory — and both are charted below rather than left out.

Sustained edge ingest, edges/shigher is better · /s
CognoDB 0.9.15
118,041durable
NebulaGraph 3.8.0
117,426async — not comparable
Memgraph 3.12.0
106,561durable
ArangoDB 3.12.10
42,251durable
Neo4j 5.26.29
32,022durable

CognoDB is the fastest durable loader here. NebulaGraph's figure is within half a percent of it, but Nebula's storage layer defaults to wal_sync=false: it acknowledges a commit before the write reaches stable storage, so the two numbers are not measuring the same promise. Memgraph ships with its WAL and snapshots switched off entirely and was raised to --storage-wal-enabled for this run so its row means something.

Read requests/s at 40 clientshigher is better · /s
CognoDB 0.9.15
64,702
Memgraph 3.12.0
45,153
Neo4j 5.26.29
18,418
ArangoDB 3.12.10
10,922
NebulaGraph 3.8.0
6,494

A 35/35/30 mix of point lookup, 1-hop and 2-hop expansion. This is a latency-derived request rate, not a capacity ceiling: the client runs a fixed 40 workers closed-loop, so the figure is arithmetically 40 ÷ mean latency, and no engine here used more than about 1.3 of its 8 cores. Read it as “how fast does one request come back under this much concurrency”, not “how much can it take”.

Write requests/s at 40 clientshigher is better · /s
Memgraph 3.12.0
38,276
CognoDB 0.9.15
25,668
Neo4j 5.26.29
14,221
NebulaGraph 3.8.0
11,511
ArangoDB 3.12.10
7,318

Memgraph leads this column, at roughly 1.5× CognoDB. Two-thirds property writes, one-third edge merges over pairs already in the graph. Every engine except Neo4j returned a handful of retryable transaction conflicts at this concurrency — 13, 11, 9 and 7 out of some 1.5 million requests — which is optimistic concurrency control working, not an error rate worth charting.

Mixed read/write requests/s at 40 clientshigher is better · /s
CognoDB 0.9.15
49,777
Memgraph 3.12.0
43,701
Neo4j 5.26.29
18,051
ArangoDB 3.12.10
10,394
NebulaGraph 3.8.0
6,756

The mix an application actually issues: reads and writes interleaved against the same graph. Memgraph closes most of its read deficit here — 43,701 against CognoDB's 49,777 — because the write share plays to where it is strongest.

On-disk bytes per edgelower is better · B
CognoDB 0.9.15
106
Neo4j 5.26.29
494

Only two engines in this run can be measured on disk: the rest write their store somewhere the container's cgroup counters cannot size, so charting them would be inventing a number. On the pair that can be measured, CognoDB stored the same graph in 304 MB against Neo4j's 1.41 GB.

Where CognoDB does not lead: Memgraph is faster on writes, uses less CPU and less memory, and takes the tail on the deepest traversal. Every one of those is charted below rather than described here.

Latency by query shape

Per-query latency, at three percentiles.

Throughput is one number for a whole mix. This is the distribution underneath it: each of the five query shapes, measured separately on every engine, at the median, the 95th and the 99th percentile. Lower is better throughout.

Point lookup — p50, mslower is better · ms
CognoDB 0.9.15
0.38p95 1.789 · p99 3.194
Memgraph 3.12.0
0.65p95 2.245 · p99 3.354
Neo4j 5.26.29
1.98p95 3.582 · p99 5.44
ArangoDB 3.12.10
3.04p95 5.968 · p99 8.689
NebulaGraph 3.8.0
4.89p95 9.991 · p99 15.515
1-hop expand — p50, mslower is better · ms
CognoDB 0.9.15
0.39p95 1.828 · p99 3.243
Memgraph 3.12.0
0.66p95 2.248 · p99 3.342
Neo4j 5.26.29
2.00p95 3.598 · p99 5.475
ArangoDB 3.12.10
3.13p95 6.103 · p99 8.958
NebulaGraph 3.8.0
5.22p95 10.474 · p99 16.166
2-hop expand — p50, mslower is better · ms
CognoDB 0.9.15
0.41p95 2.396 · p99 5.081
Memgraph 3.12.0
0.70p95 2.356 · p99 3.545
Neo4j 5.26.29
2.04p95 3.884 · p99 6.194
ArangoDB 3.12.10
3.54p95 7.469 · p99 12.914
NebulaGraph 3.8.0
6.00p95 15.764 · p99 32.755
Latency in milliseconds by query shape and percentile, per engine
Query shapeCognoDBNebulaGraphMemgraphArangoDBNeo4j
Point lookup p500.3824.8920.6453.0431.982
Point lookup p951.7899.9912.2455.9683.582
Point lookup p993.19415.5153.3548.6895.440
1-hop expand p500.3915.2240.6593.1312.000
1-hop expand p951.82810.4742.2486.1033.598
1-hop expand p993.24316.1663.3428.9585.475
2-hop expand p500.4136.0020.6973.5442.041
2-hop expand p952.39615.7642.3567.4693.884
2-hop expand p995.08132.7553.54512.9146.194
Property write p502.0993.2230.6305.3483.118
Property write p952.8056.0023.5139.5895.890
Property write p995.0418.0534.51713.5397.236
Edge merge p500.3193.1990.6124.6931.847
Edge merge p950.5576.3473.4957.9663.349
Edge merge p990.8208.8614.49611.2034.248

Bold is the best value in that row, and it does not always land on CognoDB: Memgraph takes the 95th and 99th percentile on the two-hop expansion, and the median on property writes. CognoDB is fastest at all three percentiles on point lookup and one-hop expansion.

Sustained ingest

Does the load rate hold as the graph grows?

Edges per second per window while loading, one line per engine. A median tells you the level; only the shape tells you whether an engine slows down as its store fills.

Sustained edge ingest by windowCognoDB 0.9.15NebulaGraph 3.8.0Memgraph 3.12.0ArangoDB 3.12.10Neo4j 5.26.29
Sustained edge ingest by window — edges per second against window038k75k113k150k123456window
Values
windowCognoDB 0.9.15NebulaGraph 3.8.0Memgraph 3.12.0ArangoDB 3.12.10Neo4j 5.26.29
1124,191119,648119,63044,26732,770
2120,439122,565111,97541,81832,408
3119,876119,265110,48942,30632,262
4114,885116,191108,03041,14832,703
5115,434123,613107,29740,99332,352
6114,088109,669107,79342,11832,057

CognoDB and NebulaGraph run in the same band, and Memgraph just below. CognoDB drifts from 124,191 to 114,088 edges/s across the load — about 8% — which is the folding cost being paid alongside the write path rather than in a stop-the-world pause. Neo4j is flat and low; ArangoDB flat in between.

Resource footprint

What each engine cost the box.

Measured from each container's cgroup counters over the ingest phase — the same instrument for every engine, not a self-report.

Peak RAM during ingest, MBlower is better · MB
Memgraph 3.12.0
791
NebulaGraph 3.8.0
851
Neo4j 5.26.29
1,409
ArangoDB 3.12.10
1,706
CognoDB 0.9.15
2,683

CognoDB is last here, at 2,683 MB against Memgraph's 791. Worth one line of context rather than an excuse: every engine ran with no memory limit on a 32 GiB box, and CognoDB uses the RAM available to it as cache, so this is where it settled rather than what it requires. Under a limit it runs in far less — but that is a different measurement, and this page only publishes the one that was taken.

CPU consumed during ingest, core-secondslower is better · s
Memgraph 3.12.0
34.00.81 of 8 cores avg
CognoDB 0.9.15
54.01.29 of 8 cores avg
NebulaGraph 3.8.0
69.00.94 of 8 cores avg
ArangoDB 3.12.10
1031.09 of 8 cores avg
Neo4j 5.26.29
1261.1 of 8 cores avg

Memgraph is the most CPU-frugal loader; CognoDB second, at less than half of Neo4j's. The averages matter more than the ranking: no engine used more than about 1.3 of the 8 cores available, so nothing on this page was CPU-bound and none of the throughput figures is a saturation ceiling.

Resource footprint per engine
EngineDurabilityPeak RAMCPUDisk writtenStoreB/edge
CognoDB 0.9.15durable2,683 MB54 s222 MB304 MB106.1
NebulaGraph 3.8.0async851 MB69 s488 MBnot measurable—
Memgraph 3.12.0durable791 MB34 s201 MBnot measurable—
ArangoDB 3.12.10durable1,706 MB103 s2,354 MBnot measurable—
Neo4j 5.26.29durable1,409 MB126 s870 MB1,412 MB493.6

“Not measurable” is literal: three of these engines write their store somewhere the container's cgroup counters cannot size, so there is no honest number to put in the cell. Disk-written is what the engine actually pushed to the device during the load, and it tracks durability posture more than efficiency — ArangoDB wrote 2.35 GB, CognoDB 222 MB.

Correctness

Proof the engines answered the same question.

A throughput comparison is meaningless if one engine is doing less work. Replaying an identical workload over an identical graph must return identical results, so the mean result per query is compared against the Neo4j reference. This is a pass/fail gate, not a footnote: an engine that misses it has its serve numbers withheld from this page.

Correctness and graph-parity per engine
EngineNodes loadedEdges loadedRead deviationErrorsVerdict
CognoDB 0.9.15663,2303,000,0001.18%11verified
NebulaGraph 3.8.0663,2303,000,0000.38%13verified
Memgraph 3.12.0663,2303,000,0000.94%7verified
ArangoDB 3.12.10663,2303,000,0000.33%9verified
Neo4j 5.26.29663,2303,000,000reference0reference

Every engine loaded the same 3,000,000 edges and 663,230 nodes — not approximately, exactly — which is what makes the deviation column meaningful in the first place. The errors are retryable transaction conflicts on the write leg at 40 concurrent writers, a handful out of roughly 1.5 million requests each; that is optimistic concurrency control working, not a failure rate.

Methodology

One public dataset, one stock machine, no tuning.

Every figure above comes from the same run: the public SNAP soc-Pokec social graph loaded into CognoDB 0.9.15 on a single Google Cloud n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), with the load generator on a second machine one hop away so the client never competes with the database for CPU. The engine ran on its defaults — no configuration flags, no cache warming beyond a single discarded warm-up rep — and every reported number is the median of 3 measured repetitions.

The graph is pinned to exactly 3,000,000 edges over 663,230 nodes. That detail matters more than it sounds: the serve legs run against whatever the load left behind, so a build that ingests faster would otherwise be measured against a larger graph and score worse for it. Pinning the edge count is what makes one run comparable to another.

Queries were driven over Bolt with standard drivers — the same protocol and code path a production application uses. Storage figures are read from the container's cgroup counters rather than self-reported by the engine.

Engine build
0.9.15 (3106d98)
Measured
2026-08-25
Machine
n2-standard-8 · Intel Cascade Lake
Graph
3,000,000 edges / 663,230 nodes
Concurrency
40 clients, 60 s per leg
Reps
3, median reported

Workloads

What each number actually measures.

Read queries

64,702/s

Read-only throughput: 40 concurrent clients replaying a seeded workload for 60 s. The mixed run (10% writes) sustained 49,777/s, and the write-only run 25,668/s. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.

2-hop p95

0.51 ms

p95 latency for a two-hop, friends-of-friends traversal against the loaded graph (p50 0.25 ms). Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.

Ingest edges

118,041/s

Edges merged per second, sustained while loading the graph over Bolt in batches, alongside 224,519 nodes/s. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.

On-disk / edge

106 B

On-disk store size divided by edges loaded (304 MB for 3,000,000 edges), read from the container's cgroup counters rather than self-reported. A compact store keeps the working set cached. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.

Token economics

The benchmark that matters for AI workloads.

Token economics

Cut LLM token costs by ~98.7%.

Instead of dumping the whole knowledge base into context every turn, CognoDB retrieves only the connected neighbourhood relevant to the current query. At 3,700 entities, that drops from 202,285 tokens per query to 2,668, and the gap widens as the graph grows, because baseline tokens scale with corpus size while graph-retrieval tokens stay bounded by neighbourhood size.

Token counts measured with the GPT-4 BPE tokenizer (cl100k_base) over identical rendered text: the graph path fetches the 2-hop neighbourhood around the query entities, the baseline dumps every node and edge. Savings scale with knowledge-base size; retrieval stays bounded by neighbourhood, not corpus.

Whole knowledge base202,285 tokens
Connected neighbourhood2,668 tokens
Baseline tokens
202,285per query @ 3.7K entities
Graph retrieval
2,668per query @ 3.7K entities
Saved per query
~199,617tokens avoided
Cost saved / 1K queries
~$599at $3 / 1M input tokens
Token efficiency~98.7%
Tokens per query as the knowledge base growsDump everythingGraph retrieval
Tokens per query as the knowledge base grows — tokens per query against entities in the knowledge base063k125k188k250k1854629252k4kentities in the knowledge base
Values
entities in the knowledge baseDump everythingGraph retrieval
1859,4983,562
46223,7793,014
92547,8823,107
1,85095,5282,671
3,700202,2852,668

The saving is not one number — it is a widening gap. At 185 entities retrieval saves 62%; at 3,700 it saves 99%, because dumping the corpus scales with the corpus while retrieving a neighbourhood does not. On a small knowledge base the advantage is real but modest, and quoting only the largest row would overstate it.

What we don't publish

Four engines are missing, and here's why.

This page used to publish no competitor numbers at all. The reason given was a condition rather than a refusal — that cross-vendor benchmarks are only honest when the harness, dataset, hardware and configuration are all published and reproducible. That condition is now met, so the comparison is here. We are the vendor of one of these engines, which is exactly why the harness, the raw output and the disqualifications are all public rather than summarised by us.

Nine engines were measured. Four are not shown, because their numbers could not be verified rather than because of how they scored — every one of them loaded the graph more slowly than CognoDB, so leaving them out removes no result that would count against us:

  • ArcadeDB loaded 0 of 3,000,000 edges. It deadlocks on batched edge merges, so there was no graph to serve.
  • FalkorDB returned about 19% fewer 2-hop paths than the reference on a sample of 1.5 million queries. It shed 21.7% of its reads under this concurrency, and dropping the heaviest queries biases what survives — so the figures would describe an easier workload than everyone else answered.
  • Apache AGE completed 8,303 queries in the measurement window against Neo4j's 3.3 million. That sample is too small to confirm or deny query equivalence, so we do neither.
  • Dgraph sat at 3.08% deviation against a 2% tolerance — close enough that it needs a fixed-count replay to settle, which this run did not do.

The complete nine-engine output, including all four of these and the harness that produced them, is at github.com/wexaai/graph-db-benchmarks. The rest of our claims posture is on the same page it always was: see Why CognoDB.

Start now

~98.7%

token efficiency at 3,700 entities (see the footnotes above)

Run your own benchmark.

The most credible benchmark is yours: your queries, your data, your drivers. A free instance takes about a minute and no card.

First-graph path

Live
  1. Create a free instance

    No card. Ready in about a minute.

  2. Connect your driver

    bolt+s:// URI into the driver you already use.

  3. Write two MERGEs

    That's the entire shape of agent memory.

  4. Point an agent at it

    One MCP config block. No integration code.