Preview 18 · released July 27 · measured July 26–27, 2026
A smaller tool catalog, tested like a product change
A 24-hour evidence campaign measured the MCP wire contract, validated every tool result, and ran 224 frozen real-agent captures across six host/model combinations.
The result in one screen
Substantially less static catalog, with stronger aggregate validity
A tool catalog is context an agent pays for before it asks its first question. The released treatment keeps every Free and Pro tool, every exact tool-specific data schema, runtime structuredContent, and text fallback. What changes is the catalog advertised during discovery: repeated schema and prose weight becomes a smaller, stable contract that still tells agents how to route safely.
Strict analyzer validity rose from 103/112 control captures to 107/112 treatment captures. Codex, Claude Code, and Antigravity were 70/70 valid in the treatment arm; aggregate OpenCode validity improved from 35/42 to 37/42.
Static MCP wire measurement
The complete catalog falls below half its previous size
Compact MCP catalog bytes
tools/list responses. Tool inventory remains six for Free and fourteen for Pro.| Edition | Tools | Compact control | Preview 18 treatment | Reduction | o200k estimate |
|---|---|---|---|---|---|
| Free | 6 | 51,605 B | 22,493 B | 56.41% | 12,511 → 5,297 (57.66%) |
| Pro | 14 | 115,889 B | 47,985 B | 58.59% | 28,350 → 11,489 (59.47%) |
Where the reduction came from
| Surface | Free control → treatment | Pro control → treatment |
|---|---|---|
| Output schemas | 43,214 → 13,988 B (−67.63%) | 101,736 → 33,542 B (−67.03%) |
Output $defs | 41,898 → 10,626 B (−74.64%) | 98,685 → 25,717 B (−73.94%) |
| Server instructions | 1,168 → 800 B (−31.51%) | 1,168 → 1,000 B (−14.38%) |
| Input schemas | 4,671 → 4,671 B (unchanged) | 7,223 → 7,455 B (+3.21%) |
The Pro input schema became slightly larger on purpose. Three measured guardrails state the real four-query orientation limit, that outlines are not cursor-paginated, and that relationship kinds belong on reference queries. The smallest catalog was rejected when real smaller-model trajectories showed weaker behavior.
Protocol and schema verification
The optimization changes advertising, not tool results
- All fourteen tool-specific data schemas remain exact.
- Tool-specific data remained byte-equivalent in validation, and every text fallback decoded to its corresponding structured result.
- The advertised metadata contract retains all six possible fields and the complete empty-result trust vocabulary.
- Tool order, annotations, entitlements, and vendor metadata remain unchanged for each edition.
- Direct and daemon-backed catalogs are byte-equivalent for the same edition.
- Daemon protocol 4 prevents a new client from silently accepting the older Preview 17 catalog through a stale local daemon.
112 control + 112 treatment captures
Real agents had to navigate real repositories
A smaller catalog would be worthless if agents navigated worse with it. The matrix froze models, tasks, permissions, repository commits, warm indexes, and two repetitions per arm. Ollama contributed 996 files, 32,981 symbols, and 110,960 relationships. Kodi contributed 6,337 files, 186,919 symbols, and 262,357 relationships.
Strictly valid captures by host and model
| Host/model | Strict validity | Input change | Scout-call change | Duration change |
|---|---|---|---|---|
| Codex / gpt-5.6-sol | 36/38 → 38/38 | -27.39% | -46.80% | -15.87% |
| Claude / Sonnet 4.6 | 20/20 → 20/20 | -40.08% | -30.83% | -20.83% |
| Antigravity | 12/12 → 12/12 | Unavailable | -24.19% | -13.41% |
| OpenCode / GLM 5P2 | 11/14 → 13/14 | -9.48% | -38.55% | -16.32% |
| OpenCode / Kimi K3 | 13/14 → 13/14 | -6.66% | -14.08% | +7.42% |
| OpenCode / DeepSeek V4 Flash | 11/14 → 11/14 | +1.60% | -8.05% | +2.22% |
Fewer Scout calls in every measured host/model aggregate
The inconvenient rows stayed in the result
Kimi was 7.42% slower in aggregate and DeepSeek was 2.22% slower with 1.60% more input. Individual tasks moved further in the wrong direction: Kimi’s Norwegian R07 used 37.51% more input, DeepSeek R09 used 35.01% more, and GLM R15 used 44.93% more. Those captures remained in the final matrix.
The treatment was approved because correctness and evidence were preserved, aggregate validity improved, every host/model aggregate used fewer Scout calls, and the remaining failures were visible routing variance - not a weakened result contract.
Measured, derived, and unavailable stay separate
This is not a universal “60% fewer model tokens” claim
None of the measured hosts exposed the exact catalog bytes inserted into a model prompt, whether the catalog was reinjected on later turns, or its context-window occupancy. The static token counts are tokenizer estimates over the MCP catalog, not observed model-visible tokens.
- Codex, Claude Code, and OpenCode exposed their own usage accounting, which supports within-host control/treatment comparisons only.
- Antigravity exposed a host cache projection that discards output schemas. Its Pro projection grew 0.96% because of the deliberate input guardrails, while its Scout calls fell 24.19% and duration fell 13.41%.
- Results across different hosts and models are not treated as directly comparable token totals.
- No treatment capture timed out, leaked compressed schema grammar into its answer, or failed its required response language.
The claim is narrower: Atlas Scout advertises the same tools and exact result data in a much smaller MCP catalog, and a fixed multi-host agent matrix found no aggregate quality regression.
Frozen method and provenance
The benchmark itself had failure conditions
The task corpus covered direct reads, known-file outlines, exact and ambiguous search, callers, JSX usage, Go implementations, edit impact, architecture, semantic anchors, literals, honest empty and partial results, truncation, Free entitlements, and broad orientation. Norwegian and Spanish tasks checked that the agent answered in the requested language.
An initial set of 84 OpenCode captures was rejected after the host kept the wrong project boundary despite its subprocess working directory. Those files were archived, the runner was corrected to pass the documented project argument, a regression test was added, and all 84 captures were rerun against the intended Ollama and Kodi roots. They were not silently overwritten into the final aggregate.
- Control
- Atlas Scout
1.0.0-preview.17 - Control source
141a1e28e01c07e669212146f62d5cebc5b10857- Control binary
963205067efba91925dea07bc0f0a73717ced54169d37ea9835d8502e88518a3- Measured treatment binary
a1dc841b63374975e54d463faf1ee8ad50091c2721703dddc903bfaecf979db3- Free treatment catalog
27ca36ef5885ad4edff1922881a514dc2633da17c55c23259f395b0c0e28568a- Pro treatment catalog
b011ac70a473840719ed044a26864691511eaa1fd1c4dfc76fc4320cf13d7183
Raw host transcripts and catalogs remain in the private ignored benchmark tree because they contain absolute local paths. The public report publishes the frozen method, aggregate measurements, corrections, digests, and claim boundaries without exposing those paths.