Skip to content
← all writing

Hermes Agent Cuts Tool-Definition Tokens by Half With Progressive Disclosure

  • hermes
  • mcp
  • tokens
  • optimization

Large MCP tool catalogs are expensive. Each server's full JSON schemas ship on every API call inside the system prompt. Anthropic's engineering data puts tool-definition overhead between 15,000 and 60,000 tokens per turn for typical multi-server deployments. A 34-tool setup can produce ~45,000-token prompts, with roughly half consumed by schemas alone.

Hermes Agent now ships a built-in answer: progressive tool disclosure via tool_search. When active, MCP and non-core plugin tools are removed from the model-visible tools array and replaced with three bridge tools. The model discovers tools by name, loads full schemas on demand, then invokes them. Core Hermes tools like read_file, terminal, and search_files are never deferred.

How it works

The bridge exposes three tools:

  • tool_search(queries, limit?) — search the deferred catalog
  • tool_describe(names) — load full schemas for specific tools
  • tool_call(name, arguments) — invoke a deferred tool

Discovery uses BM25 lexical retrieval over tokenized tool names, source server names, descriptions, and parameter names, with Snowball stemming and a literal substring fallback. An optional embedding reranker using nomic-embed-text-v2-moe adds semantic matching on top.

Auto mode

The feature runs in auto mode by default. It activates only when deferrable tool schemas would exceed 10% of the active model's context window. Below that threshold, the tools array passes through unchanged and you pay zero overhead.

tools:
  tool_search:
    enabled: auto        # auto (default), on, or off
    threshold_pct: 10    # % of context at which auto kicks in
    search_default_limit: 5
    max_search_limit: 20

Token savings

Internal benchmarks from the Hermes repo measured two configurations against a baseline with heavy tool schemas.

Configuration Tool tokens per turn Reduction
Arm A: eager loading ~15,000 —
Arm B: progressive disclosure ~7,200–8,600 ~43–52%
Savings ~6,400–7,800 tok ~43–52%

A separate internal test with 72 registered tools on a 32K-context model showed a sharper reduction: tool-definition overhead fell from ~19,210 tokens to ~2,198 tokens, an 89% decrease.

Accuracy gains

The same Hermes benchmarks and Anthropic's published evals both show that fewer tool definitions can improve task accuracy. Large tool catalogs create decision paralysis; the model picks irrelevant options or hallucinates capabilities that are not present.

Model Without tool_search With tool_search Delta
Claude Opus 4 49% 74% +25 pts
Claude Opus 4.5 79.5% 88.1% +8.6 pts

The accuracy improvement is not free. Live testing against Claude Haiku 4.5 showed extra round trips on cold tool access.

Scenario ON: elapsed OFF: elapsed Extra round trips
Single obvious tool 18.5s 9.7s +3
Multi-tool chain 20.3s 14.1s +4
Core plus deferred 33.1s 9.8s +3
Pure knowledge 8.2s 2.8s 0

Single-tool tasks ran roughly twice as slow with the bridge active. Multi-tool chains were about 1.4x slower. Pure-knowledge prompts had zero extra round trips because no tool search was triggered.

Tiers and listing modes

Tool Search uses tiered disclosure. The catalog listing budget is min(threshold_pct% of context, listing_max_tokens). At tier 1, a skills-style manifest of deferred tool names and short descriptions stays visible so the model can often skip tool_search and go straight to tool_describe. At tier 2, even names collapse to per-server summaries for catalogs too large to list.

Core tools are never deferred. Subagents, kanban workers, and gateway sessions see only the tools they were granted; the catalog is scoped to the session's own enabled toolsets.

Practical notes

  • The feature is already merged onto Hermes main and shipped.
  • It costs roughly 300 extra tokens per turn for the bridge tools themselves.
  • Deferred schemas loaded via tool_describe do not benefit from system-prompt cache prefix, though they do enter conversation history and cache on subsequent turns.
  • Adding or removing tools mid-session invalidates the prompt cache, same as any toolset edit.
  • Smaller models may need prompt guidance to write good search queries; the published accuracy numbers are with Opus-class models.

Sources

Termagotchi
_

Ryan Underdown

Autodidact. Rarely listens to advice.

Follow on X @catamarammed or GitHub @underdown