Atlas Scout
1.0.0-preview.21
Persistent structural index
Builds a local multi-language structural map before serving the measured result. No project build or compilation database was required.
Formal campaign + additive extension · July 29–30, 2026 · factual review open
The frozen 48-position campaign measured Atlas Scout, CodeGraphContext, Serena, and mcpls across Ollama, Kodi, Firefox, and Linux. A separate 12-position Kuzu campaign and 60 accepted new-host runs extend it without changing the original denominators.
Two product-readiness boundaries: a completed persistent index plus an asserted result, or a prepared cold language-server session returning the asserted same-file symbol. The original and extension agent-task studies are separate from both readiness boundaries.
Swipe charts and dense tables sideways to inspect every product and value.
Frozen versions · default primary configurations
This is not a moral ranking. Each project reflects a useful architectural choice, and the comparison became more informative precisely because those choices differ.
1.0.0-preview.21
Persistent structural index
Builds a local multi-language structural map before serving the measured result. No project build or compilation database was required.
0.5.2
Persistent graph index
Builds a persistent graph through its default FalkorDB Lite path. Its completed indexes were compact and its indexing stayed near one logical core. A separately frozen KuzuDB configuration is reported in the extension.
1.6.1
Language-server navigation
Uses language servers for symbol navigation and offers a broader agent toolkit. This campaign measured navigation only, with its documented C/C++ preparation.
0.3.8
LSP-to-MCP bridge
Exposes language-server navigation through MCP. Once clangd had the prepared project metadata, known-file symbol readiness was exceptionally fast.
Sourcegraph was reserved for a separate enterprise reference category and was not part of this local four-product campaign. Shell-only navigation appears in the separate agent task study, not the resource tables below.
The most important methodological decision
Atlas Scout and CodeGraphContext finish an explicit indexing command before the assertion is queried. The measured product state is a completed reusable structural or graph index.
Comparable within this lane.
Serena and mcpls start a cold language-server session and return an asserted symbol from a named file. Background language-server work may continue afterward.
Valuable, but not index completion.
The primary metric is therefore time to an asserted navigation result with the completion boundary attached. Empty product caches were used; the operating system page cache was uncontrolled and potentially warm.
Atlas Scout · CodeGraphContext
On both corpora where both products completed, Atlas Scout reached the declared finish line first. On Kodi the gap is felt rather than read: 11.6 seconds against eleven and a half minutes. The large-corpus outcome is equally important: Scout finished all Firefox and Linux repetitions; default CodeGraphContext reached the frozen Linux usefulness ceiling before a qualifying result.
| Corpus | Atlas Scout p50 | CodeGraphContext p50 | Category result |
|---|---|---|---|
| Ollama | 3.379 s | 47.514 s | Scout 14.1× faster |
| Kodi | 11.624 s | 688.114 s | Scout 59.2× faster |
| Firefox | 333.346 s | Skipped after Linux ceiling | No fabricated comparison |
| Linux | 296.549 s | DNF at 5,400.034 s | Index not ready at ceiling |
CodeGraphContext documents KuzuDB as a larger-capacity embedded alternative. It was not silently substituted for the frozen default FalkorDB Lite configuration; a KuzuDB run belongs in a separately labeled tuned campaign.
Serena · mcpls · prepared engine state
The LSP-backed products demonstrate a different—and genuinely compelling—experience. With the required project metadata prepared, mcpls returned the asserted Firefox symbol in under a second. Serena reached the same-file milestone in under a minute on every corpus.
| Corpus | Serena p50 | mcpls p50 | Boundary |
|---|---|---|---|
| Ollama | 29.762 s | 28.771 s | Prepared same-file symbol |
| Kodi | 1.977 s | 1.832 s | Prepared same-file symbol |
| Firefox | 55.145 s | 0.773 s | Prepared same-file symbol |
| Linux | 31.828 s | 1.105 s | Prepared same-file symbol |
| Corpus | First asserted cross-file result | Clean-corpus stage sum | Exhaustive completion |
|---|---|---|---|
| Kodi | DNF at 900 s | At least 912.209 s | Not exposed or proved |
| Firefox | 240.081 s | 356.026 s | Not exposed or proved |
| Linux | 94.523 s | 217.201 s | Not exposed or proved |
Equivalent strong cross-file pilots were not run for mcpls, so its values are not measured—not inferred from Serena and not treated as zero.
Setup is not query latency—but it is still user time
Serena’s official C/C++ guidance requires a correctly configured compile_commands.json. Because mcpls also launches clangd and recognizes that metadata, both products received the same hash-locked database and isolated clangd cache. Atlas Scout and CodeGraphContext required no corpus build preparation.
| Corpus | Passing preparation | Preparation tree | Compilation database |
|---|---|---|---|
| Kodi | 12.209 s | 174,286,263 B | 6,067,607 B · 1,857 entries |
| Firefox | 115.945 s | 2,524,298,342 B | 48,026,321 B · 11,812 entries |
| Linux | 122.678 s | 1,052,802,447 B | 9,033,532 B · 3,089 entries |
What preparation actually meant
Kodi needed an out-of-tree CMake configure and a retry with its documented TinyXML2 fallback after the default failed. Firefox configured isolated Mozilla build state and downloaded pinned toolchains and sysroots. Linux required an external x86-64 defconfig, a full make -j16 build, and the kernel’s compilation-database generator.
The Linux database represents one x86-64 configuration—not every architecture or conditional file. Serena’s 280,339,061-byte managed clangd tree is recorded separately; its acquisition time was not measured well enough to publish as zero.
Speed, memory, CPU, writes, and disk
Scout’s large-corpus completion came with substantial resource use. CodeGraphContext’s completed indexes were much smaller and its builds stayed near one logical core. The LSP products used little persistent cache by comparison, but their process trees continued working after the first result and reached multi-gigabyte memory peaks.
| Corpus / product | Index time | Peak memory | Average CPU | Physical write | Final cache |
|---|---|---|---|---|---|
| Ollama / Scout | 3.254 s | 0.299 GiB | 193.39% | 0.128 GiB | 0.060 GiB |
| Ollama / CGC | 46.271 s | 0.673 GiB | 100.94% | 0.024 GiB | 0.009 GiB |
| Kodi / Scout | 11.476 s | 0.975 GiB | 146.44% | 0.593 GiB | 0.275 GiB |
| Kodi / CGC | 685.555 s | 1.359 GiB | 100.62% | 0.108 GiB | 0.026 GiB |
| Firefox / Scout | 329.550 s | 12.309 GiB | 199.19% | 23.214 GiB | 7.724 GiB |
| Linux / Scout | 295.406 s | 11.563 GiB | 167.32% | 20.882 GiB | 7.016 GiB |
| Corpus | Serena peak memory | Serena average CPU | mcpls peak memory | mcpls average CPU |
|---|---|---|---|---|
| Ollama | 1.324 GiB | 164.26% | 1.871 GiB | 202.59% |
| Kodi | 1.727 GiB | 1,004.89% | 4.211 GiB | 883.25% |
| Firefox | 3.386 GiB | 614.16% | 3.677 GiB | 1,534.94% |
| Linux | 3.135 GiB | 774.69% | 2.856 GiB | 1,483.34% |
LSP memory can peak after the readiness milestone. Average CPU covers the complete measured phase, including a frozen 30-second live observation—not only the named-file query.
AT001 · 45 valid runs · five arms · three agent hosts
A fast index is worth nothing if the agent holding it still answers wrong. A separate campaign asked each agent to locate Ollama’s (*Server).ChatHandler, report its exact inclusive range, and explain the opening request-validation behavior from source. Every host ran shell-only navigation and all four MCP products three times in a frozen order.
Overall, 37 of 45 answers earned 5/5. Every answer earned full behavior and source evidence credit; the remaining eight missed only the exact-range points, and all eight came from Claude Code.
| Arm | Full-credit runs | Points | Mean score | Median elapsed |
|---|---|---|---|---|
| Atlas Scout | 9/9 | 45/45 | 5.000 | 28.889 s |
| CodeGraphContext | 8/9 | 43/45 | 4.778 | 44.851 s |
| mcpls | 7/9 | 41/45 | 4.556 | 45.336 s |
| Serena | 6/9 | 39/45 | 4.333 | 98.146 s |
| Shell-only | 7/9 | 41/45 | 4.556 | 45.381 s |
These are whole-agent totals: instructions, tool catalogs, cached context, navigation, source reads, retries, reasoning, and final answers. Each host uses a different model, tokenizer, cache scheme, and usage schema, so the only honest comparisons are within one host.
Claude Code · claude-sonnet-4-6
Scout was the only Claude cell with 3/3 full-credit answers. Against shell, its median used 66.8% less total input, 55.8% less output, and 73.2% less provider cost.
| Median | Atlas Scout | Shell-only |
|---|---|---|
| Total input | 49,588 | 149,390 |
| Uncached input | 1,563 | 9,100 |
| Cache-read input | 48,025 | 139,918 |
| Cache-write input | 969 | 8,467 |
| Output | 682 | 1,543 |
| Provider cost | $0.03114 | $0.11604 |
Fourteen Claude runs included a small auxiliary Haiku invocation. The final provider record includes both models; cache-write input is a subset of uncached input, not an extra category.
Codex · gpt-5.6-sol · medium effort
Shell used 7.0% fewer total input tokens. Scout still used 40.3% less uncached input and 35.7% less output, and had the lowest total and uncached input of the four Codex MCP arms. Every Codex cell earned full credit.
| Median | Atlas Scout | Shell-only |
|---|---|---|
| Total input | 83,783 | 77,928 |
| Uncached input | 21,319 | 35,688 |
| Cache-read input | 62,464 | 43,264 |
| Output | 595 | 925 |
| Reasoning output | 139 | 342 |
Codex reports total input with cached input already included. No cache-write input was reported in these runs.
Antigravity 1.1.8 · gemini-3.6-flash-medium
Scout used 49.8% less total input and 79.2% less uncached input than shell, while producing 28.9% more output. Serena’s uncached-input median was 7.8% lower than Scout’s. All Antigravity cells earned full credit.
| Median | Atlas Scout | Shell-only |
|---|---|---|
| Total input | 113,509 | 225,976 |
| Uncached input | 23,973 | 115,426 |
| Cache-read input | 89,536 | 138,511 |
| Output | 3,169 | 2,459 |
| Thinking output | 2,255 | 1,385 |
Formal values were recovered retrospectively from retained local provider metadata and validated against supported live JSON output. Future runs must archive stream-json directly.
A correct final answer does not prove the assigned MCP product helped. Agents can fall back to shell and source reads after an empty result, cancellation, or startup failure. The routing record keeps product contribution separate from final-answer correctness.
| Arm | Calls | Completed | Non-empty | Empty | Failed |
|---|---|---|---|---|---|
| Atlas Scout | 9 | 9 | 9 | 0 | 0 |
| CodeGraphContext | 12 | 6 | 3 | 3 | 6 |
| mcpls | 12 | 4 | 4 | 0 | 8 |
| Serena | 12 | 12 | 12 | 0 | 0 |
| Shell-only | 0 | 0 | 0 | 0 | 0 |
Serena also returned non-empty output on every call. CodeGraphContext contributed non-empty results in all three Claude runs, while its other host paths were empty or cancelled. mcpls contributed four non-empty Claude results; its Antigravity discovery incompatibility produced three valid startup-failure outcomes, and its Codex calls were cancelled. Those are host-product interoperability results—not discarded harness failures.
Interpretation boundary
This is one English exact-symbol task on one pinned Ollama commit, with three runs per cell and one partially blinded reviewer. It supports a transparent case study—not a general claim that Scout always improves every agent, repository, task, token category, or edit outcome.
Authoritative revision-4 lock SHA-256: 34e4696febd67c9f3df9bcf91be4cca53d1ea14b6b74e7ef7e39cad2c240cdceToken-augmented summary SHA-256: 8318ddadb2a22bc30ffb016ac600cf3c146734a0767de15cb202fd70d4d61d06
Additive extension · July 30, 2026 · 12 resource positions · 60 accepted agent runs
This extension does not rewrite the original 48-position resource campaign or 45-run agent study. It separately tested CodeGraphContext’s KuzuDB backend, added Claude Opus 5, and added three OpenCode model lanes through the operator-owned ZaguanAI gateway with Fireworks as their common downstream inference provider.
CodeGraphContext 0.5.2 with KuzuDB 0.11.3 did not finish indexing Ollama, the smallest frozen corpus. At 44 minutes 35.810 seconds, the resource guard fired as host-available memory crossed the frozen 12 GiB floor. This was not a query-quality failure: no ready index existed to query.
The first complete run was rejected
Sixty finished runs were discarded rather than scored against a contaminated host. Revision 2 completed all 60 agent runs, but postflight inspection found a CodeGraphContext lane had daemonized Redis and contaminated later host state. None of those results was scored. Cleanup was strengthened, the failure path was reproduced, and revision 3 restarted every ordinal. In the accepted campaign, all 12 CodeGraphContext lanes stopped their lane-local Redis daemon and left zero residual processes.
All 60 accepted runs answered the task, preserved the source worktree, and passed model, transcript, mutation, credential, and process-cleanup gates. Every answer correctly explained the request-validation behavior. The four non-perfect answers all contained exact-range errors; one also shifted the cited validation lines.
| Host / model | Atlas Scout | CodeGraphContext | mcpls | Serena | Shell | Total |
|---|---|---|---|---|---|---|
| Claude Opus 5 | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 | 75/75 |
| OpenCode / DeepSeek V4 Flash | 15/15 | 15/15 | 15/15 | 8/15 | 15/15 | 68/75 |
| OpenCode / GLM-5p2 | 15/15 | 15/15 | 15/15 | 15/15 | 13/15 | 73/75 |
| OpenCode / Kimi K3 | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 | 75/75 |
| Arm total | 60/60 | 60/60 | 60/60 | 53/60 | 58/60 | 291/300 |
| Median wall time | Atlas Scout | CodeGraphContext | mcpls | Serena | Shell |
|---|---|---|---|---|---|
| Claude Opus 5 | 32.110 s | 40.519 s | 46.981 s | 40.338 s | 37.979 s |
| OpenCode / DeepSeek V4 Flash | 9.214 s | 21.893 s | 40.481 s | 112.413 s | 25.656 s |
| OpenCode / GLM-5p2 | 15.722 s | 33.674 s | 49.337 s | 120.164 s | 28.842 s |
| OpenCode / Kimi K3 | 34.054 s | 42.744 s | 70.308 s | 147.755 s | 60.920 s |
Every observed MCP call in the accepted extension completed with non-empty output. Scout used one MCP call per run; the other products used between 17 and 32 calls across their 12 runs.
| Arm | Calls | Completed non-empty | Failed | Permission denied |
|---|---|---|---|---|
| Atlas Scout | 12 | 12 | 0 | 0 |
| CodeGraphContext | 17 | 17 | 0 | 0 |
| mcpls | 32 | 32 | 0 | 0 |
| Serena | 24 | 24 | 0 | 0 |
| Shell | 0 | 0 | 0 | 0 |
Hosts and models use different tokenizers, prompts, caches, and accounting schemas. These values therefore identify the lowest median inside each frozen lane; they are not pooled into a cross-model token leaderboard.
| Host / model | Lowest total input | Lowest uncached input | Lowest output |
|---|---|---|---|
| Claude Opus 5 | Scout · 56,983 | Shell · 6,410 | Scout · 1,395 |
| OpenCode / DeepSeek V4 Flash | Scout · 22,083 | Scout · 8,067 | Scout · 421 |
| OpenCode / GLM-5p2 | Scout · 20,863 | Scout · 7,705 | Scout · 496 |
| OpenCode / Kimi K3 | Scout · 19,640 | Scout · 8,376 | Scout · 1,085 |
Claude Opus shell’s uncached median was only 29 tokens below Scout’s 6,439. OpenCode reported zero cost because the custom provider configuration contained no pricing metadata; that is not evidence that ZaguanAI, Fireworks, or the model invocation was free.
cbccbc3a6f125626d1e73afa8d4ba81a6640b789eafb3bc7d520864575bdfbf7a1c4e030cae1d9742e63130efd8420ac89863d89108f787acc751f3e2f7ad10f3e63989d14a37414600dd0b792751f4369fc4ec9159aaa0e3a6f7918ce63415bbed1c017543d4942e022f385b2d4b89ddaa186018d16a1f8405f3eb1bb12d0e35c4f7990da4f34385def3dcb9f3cb2c4010a196ca06a48e03c500fcdbc858d03Interpretation boundary
This remains one English exact-symbol task on one pinned Ollama commit, with three runs per cell and one partially blinded reviewer. It supports the exact Kuzu safety, correctness, routing, wall-time, and within-lane token observations above—not a general product or model winner.
The formal report, frozen locks, invalidated revision, all 60 accepted reruns, deterministic analyses, and retained raw evidence are published in the benchmark evidence repository.
No grand score · no loser column
The lowest completed readiness time among the two tested persistent indexers, full large-corpus completion in minutes, and no build or compilation-database prerequisite.
Its cost is equally clear: large indexes, high peak memory, and heavy write amplification on Firefox and Linux.
Much smaller completed index footprints and near-single-core indexing in the completed lanes, with a documented alternative backend for larger-capacity use.
Its frozen default FalkorDB Lite path did not meet the Linux usefulness ceiling on this workstation. The separately frozen KuzuDB variant crossed the host-memory safety floor before Ollama became ready, so no Kuzu query quality was scored.
Useful cold, prepared language-server navigation plus demonstrated strong cross-file results on Firefox and Linux, inside a broader agent-oriented product.
C/C++ quality depends on real project metadata and backend preparation; the broader editing experience was deliberately outside this navigation-only comparison.
Remarkably fast prepared known-file navigation: sub-second on Firefox and near one second on Linux through a small MCP bridge to clangd.
Strong cross-file and exhaustive-completion milestones were not measured, so the page makes no claim about them.
Frozen protocol · deterministic aggregation · no outlier removal
| Corpus | Commit | Tracked files | Checkout bytes |
|---|---|---|---|
| Ollama | 83d4311ffecb79bf3a7a1f0341afa4eb069da853 | 1,267 | 53,619,446 |
| Kodi | bf8548300cac6c9d1400ad6284ab9909f8cdec25 | 10,087 | 105,628,337 |
| Firefox | f4e6e71e9c4deb3818880f2c8a22de09a92955ce | 469,749 | 3,524,879,410 |
| Linux | 248951ddc14de84de3910f9b13f51491a8cd91df | 94,843 | 1,615,205,959 |
Every product command and its descendants ran in an isolated cgroup. The harness measured complete process trees, retained every observation, and sampled host memory and swap. The campaign began after reboot with zero swap; the observed maximum stayed below 50 MiB and the frozen 256-MiB launch gate.
expanded-resource-001-preview21-r520260729cdd8598516f6b1c2f8f3929d0553cf81979133ab39ae301799bb00f7717f9004f8218f338ab938292a9c7a1bdc85785e4894c0ec9f047a550048a17d52b42d9592e396cfa00199d16f457e4bf7b2bff4dd78b88725b2e8169dfd846fa65e88bdb8dc8732f536ee27e3f486a11e8e64d2e30b603675de75ced6a5fc560b0ea510The 48 immutable lane manifests, machine-readable aggregation, included raw runs, and normalized analyses are published in the public evidence repository. The repository also records the precise boundary between published evidence and excluded reproducible or third-party artifacts.
What this page cannot honestly claim
The claim the evidence supports
Atlas Scout had the lowest complete-readiness time among the two tested persistent graph indexers on both corpora where both completed.
In the separate 60-run extension, Atlas Scout had the lowest median wall time, total input, and output in each of the four new host/model lanes.
The benchmark did what it was supposed to do: it found a real Atlas Scout advantage, several real competitor advantages, and the next engineering work Scout has not earned the right to avoid.
Public evidence repository