Benchmarks
The numbers, and exactly how they were measured.
Measured on engine 0.9.15, on 2026-08-25, against a graph pinned to exactly 3,000,000 edges — the same graph for every engine on this page, which is what makes the comparison below mean anything. Every chart is drawn from that run's own output, and the raw per-engine JSON is public.
- 1Read-only throughput: 40 concurrent clients replaying a seeded workload for 60 s. The mixed run (10% writes) sustained 49,777/s, and the write-only run 25,668/s. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.
- 2p95 latency for a two-hop, friends-of-friends traversal against the loaded graph (p50 0.25 ms). Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.
- 3Edges merged per second, sustained while loading the graph over Bolt in batches, alongside 224,519 nodes/s. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.
- 4On-disk store size divided by edges loaded (304 MB for 3,000,000 edges), read from the container's cgroup counters rather than self-reported. A compact store keeps the working set cached. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.
Latency
Per-query latency, not just an average.
A median hides the tail that users actually notice, so every query shape is reported at p50, p95 and p99 across the measured reps.
- p50
- 0.38 ms
- p95
- 1.79 ms
- p99
- 3.19 ms
MATCH (n:SocialUser {id:$x}) RETURN n.id- p50
- 0.39 ms
- p95
- 1.83 ms
- p99
- 3.24 ms
MATCH (:SocialUser {id:$x})-[:FRIEND]->(m) RETURN count(m)- p50
- 0.41 ms
- p95
- 2.40 ms
- p99
- 5.08 ms
MATCH (:SocialUser {id:$x})-[:FRIEND]->()-[:FRIEND]->(m) RETURN count(m)- p50
- 2.10 ms
- p95
- 2.81 ms
- p99
- 5.04 ms
MATCH (n:SocialUser {id:$x}) SET n.v = $v- p50
- 0.32 ms
- p95
- 0.56 ms
- p99
- 0.82 ms
MATCH (a {id:$x}) MATCH (b {id:$y}) MERGE (a)-[:FRIEND]->(b)Throughput
What it sustains, and while writing.
Read-only is the easy number. The mixed and write-only mixes are the ones that show what contention costs.
- Read-only
- 64,702
- Mixed (10% writes)
- 49,777
- Write-only
- 25,668
11 errors across the measured reps.
Values
| window | edges/s |
|---|---|
| 1 | 124,191 |
| 2 | 120,439 |
| 3 | 119,876 |
| 4 | 114,885 |
| 5 | 115,434 |
| 6 | 114,088 |
Ingest rate per window while loading the graph over Bolt. A flat line is the point: the rate holds as the store grows rather than decaying, because folding runs alongside the load instead of stopping it.
Cross-engine
Five engines, one identical graph.
Every engine below loaded exactly 3,000,000 edges and 663,230 nodes on its own dedicated n2-standard-8, with the load generator on a second machine one hop away. Each ran its own stock configuration plus an index on the node id. Query equivalence is verified, not assumed: replaying the identical workload over the identical graph must return the same results, and every engine here agrees with the Neo4j reference to within 1.18%.
Fastest durable ingest
118,041/s
1.11× Memgraph, the next best
Fastest reads
64,702/s
1.43× Memgraph, the next best
Fastest mixed
49,777/s
1.14× Memgraph, the next best
Densest on disk
106 B/edge
4.65× Neo4j, the next best
Four of the six measured columns. The other two go to Memgraph — writes and peak memory — and both are charted below rather than left out.
- CognoDB 0.9.15
- 118,041durable
- NebulaGraph 3.8.0
- 117,426async — not comparable
- Memgraph 3.12.0
- 106,561durable
- ArangoDB 3.12.10
- 42,251durable
- Neo4j 5.26.29
- 32,022durable
CognoDB is the fastest durable loader here. NebulaGraph's figure is within half a percent of it, but Nebula's storage layer defaults to wal_sync=false: it acknowledges a commit before the write reaches stable storage, so the two numbers are not measuring the same promise. Memgraph ships with its WAL and snapshots switched off entirely and was raised to --storage-wal-enabled for this run so its row means something.
- CognoDB 0.9.15
- 64,702
- Memgraph 3.12.0
- 45,153
- Neo4j 5.26.29
- 18,418
- ArangoDB 3.12.10
- 10,922
- NebulaGraph 3.8.0
- 6,494
A 35/35/30 mix of point lookup, 1-hop and 2-hop expansion. This is a latency-derived request rate, not a capacity ceiling: the client runs a fixed 40 workers closed-loop, so the figure is arithmetically 40 ÷ mean latency, and no engine here used more than about 1.3 of its 8 cores. Read it as “how fast does one request come back under this much concurrency”, not “how much can it take”.
- Memgraph 3.12.0
- 38,276
- CognoDB 0.9.15
- 25,668
- Neo4j 5.26.29
- 14,221
- NebulaGraph 3.8.0
- 11,511
- ArangoDB 3.12.10
- 7,318
Memgraph leads this column, at roughly 1.5× CognoDB. Two-thirds property writes, one-third edge merges over pairs already in the graph. Every engine except Neo4j returned a handful of retryable transaction conflicts at this concurrency — 13, 11, 9 and 7 out of some 1.5 million requests — which is optimistic concurrency control working, not an error rate worth charting.
- CognoDB 0.9.15
- 49,777
- Memgraph 3.12.0
- 43,701
- Neo4j 5.26.29
- 18,051
- ArangoDB 3.12.10
- 10,394
- NebulaGraph 3.8.0
- 6,756
The mix an application actually issues: reads and writes interleaved against the same graph. Memgraph closes most of its read deficit here — 43,701 against CognoDB's 49,777 — because the write share plays to where it is strongest.
- CognoDB 0.9.15
- 106
- Neo4j 5.26.29
- 494
Only two engines in this run can be measured on disk: the rest write their store somewhere the container's cgroup counters cannot size, so charting them would be inventing a number. On the pair that can be measured, CognoDB stored the same graph in 304 MB against Neo4j's 1.41 GB.
Where CognoDB does not lead: Memgraph is faster on writes, uses less CPU and less memory, and takes the tail on the deepest traversal. Every one of those is charted below rather than described here.
Latency by query shape
Per-query latency, at three percentiles.
Throughput is one number for a whole mix. This is the distribution underneath it: each of the five query shapes, measured separately on every engine, at the median, the 95th and the 99th percentile. Lower is better throughout.
- CognoDB 0.9.15
- 0.38p95 1.789 · p99 3.194
- Memgraph 3.12.0
- 0.65p95 2.245 · p99 3.354
- Neo4j 5.26.29
- 1.98p95 3.582 · p99 5.44
- ArangoDB 3.12.10
- 3.04p95 5.968 · p99 8.689
- NebulaGraph 3.8.0
- 4.89p95 9.991 · p99 15.515
- CognoDB 0.9.15
- 0.39p95 1.828 · p99 3.243
- Memgraph 3.12.0
- 0.66p95 2.248 · p99 3.342
- Neo4j 5.26.29
- 2.00p95 3.598 · p99 5.475
- ArangoDB 3.12.10
- 3.13p95 6.103 · p99 8.958
- NebulaGraph 3.8.0
- 5.22p95 10.474 · p99 16.166
- CognoDB 0.9.15
- 0.41p95 2.396 · p99 5.081
- Memgraph 3.12.0
- 0.70p95 2.356 · p99 3.545
- Neo4j 5.26.29
- 2.04p95 3.884 · p99 6.194
- ArangoDB 3.12.10
- 3.54p95 7.469 · p99 12.914
- NebulaGraph 3.8.0
- 6.00p95 15.764 · p99 32.755
| Query shape | CognoDB | NebulaGraph | Memgraph | ArangoDB | Neo4j |
|---|---|---|---|---|---|
| Point lookup p50 | 0.382 | 4.892 | 0.645 | 3.043 | 1.982 |
| Point lookup p95 | 1.789 | 9.991 | 2.245 | 5.968 | 3.582 |
| Point lookup p99 | 3.194 | 15.515 | 3.354 | 8.689 | 5.440 |
| 1-hop expand p50 | 0.391 | 5.224 | 0.659 | 3.131 | 2.000 |
| 1-hop expand p95 | 1.828 | 10.474 | 2.248 | 6.103 | 3.598 |
| 1-hop expand p99 | 3.243 | 16.166 | 3.342 | 8.958 | 5.475 |
| 2-hop expand p50 | 0.413 | 6.002 | 0.697 | 3.544 | 2.041 |
| 2-hop expand p95 | 2.396 | 15.764 | 2.356 | 7.469 | 3.884 |
| 2-hop expand p99 | 5.081 | 32.755 | 3.545 | 12.914 | 6.194 |
| Property write p50 | 2.099 | 3.223 | 0.630 | 5.348 | 3.118 |
| Property write p95 | 2.805 | 6.002 | 3.513 | 9.589 | 5.890 |
| Property write p99 | 5.041 | 8.053 | 4.517 | 13.539 | 7.236 |
| Edge merge p50 | 0.319 | 3.199 | 0.612 | 4.693 | 1.847 |
| Edge merge p95 | 0.557 | 6.347 | 3.495 | 7.966 | 3.349 |
| Edge merge p99 | 0.820 | 8.861 | 4.496 | 11.203 | 4.248 |
Bold is the best value in that row, and it does not always land on CognoDB: Memgraph takes the 95th and 99th percentile on the two-hop expansion, and the median on property writes. CognoDB is fastest at all three percentiles on point lookup and one-hop expansion.
Sustained ingest
Does the load rate hold as the graph grows?
Edges per second per window while loading, one line per engine. A median tells you the level; only the shape tells you whether an engine slows down as its store fills.
Values
| window | CognoDB 0.9.15 | NebulaGraph 3.8.0 | Memgraph 3.12.0 | ArangoDB 3.12.10 | Neo4j 5.26.29 |
|---|---|---|---|---|---|
| 1 | 124,191 | 119,648 | 119,630 | 44,267 | 32,770 |
| 2 | 120,439 | 122,565 | 111,975 | 41,818 | 32,408 |
| 3 | 119,876 | 119,265 | 110,489 | 42,306 | 32,262 |
| 4 | 114,885 | 116,191 | 108,030 | 41,148 | 32,703 |
| 5 | 115,434 | 123,613 | 107,297 | 40,993 | 32,352 |
| 6 | 114,088 | 109,669 | 107,793 | 42,118 | 32,057 |
CognoDB and NebulaGraph run in the same band, and Memgraph just below. CognoDB drifts from 124,191 to 114,088 edges/s across the load — about 8% — which is the folding cost being paid alongside the write path rather than in a stop-the-world pause. Neo4j is flat and low; ArangoDB flat in between.
Resource footprint
What each engine cost the box.
Measured from each container's cgroup counters over the ingest phase — the same instrument for every engine, not a self-report.
- Memgraph 3.12.0
- 791
- NebulaGraph 3.8.0
- 851
- Neo4j 5.26.29
- 1,409
- ArangoDB 3.12.10
- 1,706
- CognoDB 0.9.15
- 2,683
CognoDB is last here, at 2,683 MB against Memgraph's 791. Worth one line of context rather than an excuse: every engine ran with no memory limit on a 32 GiB box, and CognoDB uses the RAM available to it as cache, so this is where it settled rather than what it requires. Under a limit it runs in far less — but that is a different measurement, and this page only publishes the one that was taken.
- Memgraph 3.12.0
- 34.00.81 of 8 cores avg
- CognoDB 0.9.15
- 54.01.29 of 8 cores avg
- NebulaGraph 3.8.0
- 69.00.94 of 8 cores avg
- ArangoDB 3.12.10
- 1031.09 of 8 cores avg
- Neo4j 5.26.29
- 1261.1 of 8 cores avg
Memgraph is the most CPU-frugal loader; CognoDB second, at less than half of Neo4j's. The averages matter more than the ranking: no engine used more than about 1.3 of the 8 cores available, so nothing on this page was CPU-bound and none of the throughput figures is a saturation ceiling.
| Engine | Durability | Peak RAM | CPU | Disk written | Store | B/edge |
|---|---|---|---|---|---|---|
| CognoDB 0.9.15 | durable | 2,683 MB | 54 s | 222 MB | 304 MB | 106.1 |
| NebulaGraph 3.8.0 | async | 851 MB | 69 s | 488 MB | not measurable | — |
| Memgraph 3.12.0 | durable | 791 MB | 34 s | 201 MB | not measurable | — |
| ArangoDB 3.12.10 | durable | 1,706 MB | 103 s | 2,354 MB | not measurable | — |
| Neo4j 5.26.29 | durable | 1,409 MB | 126 s | 870 MB | 1,412 MB | 493.6 |
“Not measurable” is literal: three of these engines write their store somewhere the container's cgroup counters cannot size, so there is no honest number to put in the cell. Disk-written is what the engine actually pushed to the device during the load, and it tracks durability posture more than efficiency — ArangoDB wrote 2.35 GB, CognoDB 222 MB.
Correctness
Proof the engines answered the same question.
A throughput comparison is meaningless if one engine is doing less work. Replaying an identical workload over an identical graph must return identical results, so the mean result per query is compared against the Neo4j reference. This is a pass/fail gate, not a footnote: an engine that misses it has its serve numbers withheld from this page.
| Engine | Nodes loaded | Edges loaded | Read deviation | Errors | Verdict |
|---|---|---|---|---|---|
| CognoDB 0.9.15 | 663,230 | 3,000,000 | 1.18% | 11 | verified |
| NebulaGraph 3.8.0 | 663,230 | 3,000,000 | 0.38% | 13 | verified |
| Memgraph 3.12.0 | 663,230 | 3,000,000 | 0.94% | 7 | verified |
| ArangoDB 3.12.10 | 663,230 | 3,000,000 | 0.33% | 9 | verified |
| Neo4j 5.26.29 | 663,230 | 3,000,000 | reference | 0 | reference |
Every engine loaded the same 3,000,000 edges and 663,230 nodes — not approximately, exactly — which is what makes the deviation column meaningful in the first place. The errors are retryable transaction conflicts on the write leg at 40 concurrent writers, a handful out of roughly 1.5 million requests each; that is optimistic concurrency control working, not a failure rate.
Methodology
One public dataset, one stock machine, no tuning.
Every figure above comes from the same run: the public SNAP soc-Pokec social graph loaded into CognoDB 0.9.15 on a single Google Cloud n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), with the load generator on a second machine one hop away so the client never competes with the database for CPU. The engine ran on its defaults — no configuration flags, no cache warming beyond a single discarded warm-up rep — and every reported number is the median of 3 measured repetitions.
The graph is pinned to exactly 3,000,000 edges over 663,230 nodes. That detail matters more than it sounds: the serve legs run against whatever the load left behind, so a build that ingests faster would otherwise be measured against a larger graph and score worse for it. Pinning the edge count is what makes one run comparable to another.
Queries were driven over Bolt with standard drivers — the same protocol and code path a production application uses. Storage figures are read from the container's cgroup counters rather than self-reported by the engine.
- Engine build
- 0.9.15 (3106d98)
- Measured
- 2026-08-25
- Machine
- n2-standard-8 · Intel Cascade Lake
- Graph
- 3,000,000 edges / 663,230 nodes
- Concurrency
- 40 clients, 60 s per leg
- Reps
- 3, median reported
Workloads
What each number actually measures.
Read queries
64,702/sRead-only throughput: 40 concurrent clients replaying a seeded workload for 60 s. The mixed run (10% writes) sustained 49,777/s, and the write-only run 25,668/s. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.
2-hop p95
0.51 msp95 latency for a two-hop, friends-of-friends traversal against the loaded graph (p50 0.25 ms). Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.
Ingest edges
118,041/sEdges merged per second, sustained while loading the graph over Bolt in batches, alongside 224,519 nodes/s. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.
On-disk / edge
106 BOn-disk store size divided by edges loaded (304 MB for 3,000,000 edges), read from the container's cgroup counters rather than self-reported. A compact store keeps the working set cached. Measured 2026-08-25 on engine 0.9.15 (commit 3106d98). SNAP soc-Pokec loaded to exactly 3,000,000 edges over 663,230 nodes, on one n2-standard-8 (8 vCPU / 32 GiB, Intel Cascade Lake), engine on defaults, median of 3 reps after a discarded warm-up.
Token economics
The benchmark that matters for AI workloads.
Token economics
Cut LLM token costs by ~98.7%.
Instead of dumping the whole knowledge base into context every turn, CognoDB retrieves only the connected neighbourhood relevant to the current query. At 3,700 entities, that drops from 202,285 tokens per query to 2,668, and the gap widens as the graph grows, because baseline tokens scale with corpus size while graph-retrieval tokens stay bounded by neighbourhood size.
Token counts measured with the GPT-4 BPE tokenizer (cl100k_base) over identical rendered text: the graph path fetches the 2-hop neighbourhood around the query entities, the baseline dumps every node and edge. Savings scale with knowledge-base size; retrieval stays bounded by neighbourhood, not corpus.
- Baseline tokens
- 202,285per query @ 3.7K entities
- Graph retrieval
- 2,668per query @ 3.7K entities
- Saved per query
- ~199,617tokens avoided
- Cost saved / 1K queries
- ~$599at $3 / 1M input tokens
Values
| entities in the knowledge base | Dump everything | Graph retrieval |
|---|---|---|
| 185 | 9,498 | 3,562 |
| 462 | 23,779 | 3,014 |
| 925 | 47,882 | 3,107 |
| 1,850 | 95,528 | 2,671 |
| 3,700 | 202,285 | 2,668 |
The saving is not one number — it is a widening gap. At 185 entities retrieval saves 62%; at 3,700 it saves 99%, because dumping the corpus scales with the corpus while retrieving a neighbourhood does not. On a small knowledge base the advantage is real but modest, and quoting only the largest row would overstate it.
What we don't publish
Four engines are missing, and here's why.
This page used to publish no competitor numbers at all. The reason given was a condition rather than a refusal — that cross-vendor benchmarks are only honest when the harness, dataset, hardware and configuration are all published and reproducible. That condition is now met, so the comparison is here. We are the vendor of one of these engines, which is exactly why the harness, the raw output and the disqualifications are all public rather than summarised by us.
Nine engines were measured. Four are not shown, because their numbers could not be verified rather than because of how they scored — every one of them loaded the graph more slowly than CognoDB, so leaving them out removes no result that would count against us:
- ArcadeDB loaded 0 of 3,000,000 edges. It deadlocks on batched edge merges, so there was no graph to serve.
- FalkorDB returned about 19% fewer 2-hop paths than the reference on a sample of 1.5 million queries. It shed 21.7% of its reads under this concurrency, and dropping the heaviest queries biases what survives — so the figures would describe an easier workload than everyone else answered.
- Apache AGE completed 8,303 queries in the measurement window against Neo4j's 3.3 million. That sample is too small to confirm or deny query equivalence, so we do neither.
- Dgraph sat at 3.08% deviation against a 2% tolerance — close enough that it needs a fixed-count replay to settle, which this run did not do.
The complete nine-engine output, including all four of these and the harness that produced them, is at github.com/wexaai/graph-db-benchmarks. The rest of our claims posture is on the same page it always was: see Why CognoDB.
Start now
~98.7%
token efficiency at 3,700 entities (see the footnotes above)
Run your own benchmark.
The most credible benchmark is yours: your queries, your data, your drivers. A free instance takes about a minute and no card.
First-graph path
LiveCreate a free instance
No card. Ready in about a minute.
Connect your driver
bolt+s:// URI into the driver you already use.
Write two MERGEs
That's the entire shape of agent memory.
Point an agent at it
One MCP config block. No integration code.