Skip to content
← all writing

Hermes `tools.stub_mode` reduces non-core tool schemas to stubs

  • hermes
  • tokens
  • tool-schema
  • context-window
  • MCP
Hermes `tools.stub_mode` reduces non-core tool schemas to stubs

Hermes Agent ships with 30 built-in tools by default. On every LLM call, every tool schema is serialized into the request body. In a recent merged proposal, a contributor measured the 30-tool baseline at roughly 50,277 bytes in the tools array, or about 12,569 tokens per call. That overhead is paid before any user content is processed, and it repeats on every turn.

A merged PR introduces tools.stub_mode, an opt-in config flag that replaces non-core tool schemas with minimal {name, description} stubs while preserving full schemas for nine core orchestration tools: delegate_task, session_search, skill_manage, memory, send_message, todo, clarify, skill_view, and skills_list.

Measured token savings

The PR includes benchmark data collected from agent logs across 30 to 115 API calls per session.

Setup Tokens / call Saved / call Saved %
30 tools, full schemas ~12,569 - -
30 tools, stub mode ~9,813 2,756 21.9%
Per 30-call session - ~82,650 -
Per 115-call session - ~316,825 -

At 30 API calls, the PR reports a session savings of about 82,650 tokens. A 115-call session saves about 316,825 tokens. The PR frames this in conversation-turn terms: roughly 110 to 422 recovered turns at 750 tokens per turn, depending on session length.

Why this matters beyond raw token counts

Hermes fires trajectory compression at 85% of the context window. With fewer tool-schema tokens on every call, that threshold is reached later. The PR reports that with 25 MCP tools loaded, stub mode defers compression by about 15 turns.

The same pattern scales with MCP tool count. Each non-core MCP schema typically consumes 1,000 to 2,000 bytes full; stub mode reduces it to about 120 bytes.

Config Tokens saved / call Extra turns before compression
30 built-in tools only ~2,756 ~4
+10 MCP tools ~6,205 ~9
+25 MCP tools ~11,380 ~15
+50 MCP tools ~20,005 ~27

Implementation detail

The main agent retains the full tool namespace. All 30 tool names remain visible, and the PR routes actual invocations through delegate_task. Sub-agents spawned via delegate_task receive full schemas. This composes with Hermes's existing tool_search mechanism for MCP tools.

A review flagged that stubbed tools are not enforced to be sub-agent-only at the session level. The current implementation reduces schema weight, but a direct stubbed call can still be dispatched from the main agent. The PR author acknowledged this and added a test plan; the limitation is not hidden.

How to enable

Add the flag to your Hermes config:

tools:
  stub_mode: true

After enabling, hermes chat -q "list your tools" should show all 30 tools, but only the 9 core tools should carry full schemas in the API payload. With MCP toolsets loaded, verify that delegation paths still execute correctly through sub-agents.

The broader context

Tool schema bloat is a recurring theme in Hermes optimization. A separate issue measured a default 30-tool setup at roughly 14,000 schema tokens per call, and a heavy multi-MCP setup at about 70,000 input tokens per turn. That volume exceeds the per-minute token cap of free-tier providers such as Cerebras and Groq, making some configurations unreachable rather than merely expensive.

stub_mode does not add round trips. It does not pick tools per turn. It reduces the weight of tools that are not meant to be called directly from the main agent. Whether it is sufficient on its own depends on how many MCP servers a deployment loads and whether direct enforcement of stub-only access is required.

Source: NousResearch/hermes-agent PR #48622, merged 2026-09.

Termagotchi
_

Ryan Underdown

Autodidact. Rarely listens to advice.

Follow on X @catamarammed or GitHub @underdown