Hermes `tools.stub_mode` reduces non-core tool schemas to stubs
- hermes
- tokens
- tool-schema
- context-window
- MCP

Hermes Agent ships with 30 built-in tools by default. On every LLM call, every tool schema is serialized into the request body. In a recent merged proposal, a contributor measured the 30-tool baseline at roughly 50,277 bytes in the tools array, or about 12,569 tokens per call. That overhead is paid before any user content is processed, and it repeats on every turn.
A merged PR introduces tools.stub_mode, an opt-in config flag that replaces non-core tool schemas with minimal {name, description} stubs while preserving full schemas for nine core orchestration tools: delegate_task, session_search, skill_manage, memory, send_message, todo, clarify, skill_view, and skills_list.
Measured token savings
The PR includes benchmark data collected from agent logs across 30 to 115 API calls per session.
| Setup | Tokens / call | Saved / call | Saved % |
|---|---|---|---|
| 30 tools, full schemas | ~12,569 | - | - |
| 30 tools, stub mode | ~9,813 | 2,756 | 21.9% |
| Per 30-call session | - | ~82,650 | - |
| Per 115-call session | - | ~316,825 | - |
At 30 API calls, the PR reports a session savings of about 82,650 tokens. A 115-call session saves about 316,825 tokens. The PR frames this in conversation-turn terms: roughly 110 to 422 recovered turns at 750 tokens per turn, depending on session length.
Why this matters beyond raw token counts
Hermes fires trajectory compression at 85% of the context window. With fewer tool-schema tokens on every call, that threshold is reached later. The PR reports that with 25 MCP tools loaded, stub mode defers compression by about 15 turns.
The same pattern scales with MCP tool count. Each non-core MCP schema typically consumes 1,000 to 2,000 bytes full; stub mode reduces it to about 120 bytes.
| Config | Tokens saved / call | Extra turns before compression |
|---|---|---|
| 30 built-in tools only | ~2,756 | ~4 |
| +10 MCP tools | ~6,205 | ~9 |
| +25 MCP tools | ~11,380 | ~15 |
| +50 MCP tools | ~20,005 | ~27 |
Implementation detail
The main agent retains the full tool namespace. All 30 tool names remain visible, and the PR routes actual invocations through delegate_task. Sub-agents spawned via delegate_task receive full schemas. This composes with Hermes's existing tool_search mechanism for MCP tools.
A review flagged that stubbed tools are not enforced to be sub-agent-only at the session level. The current implementation reduces schema weight, but a direct stubbed call can still be dispatched from the main agent. The PR author acknowledged this and added a test plan; the limitation is not hidden.
How to enable
Add the flag to your Hermes config:
tools:
stub_mode: true
After enabling, hermes chat -q "list your tools" should show all 30 tools, but only the 9 core tools should carry full schemas in the API payload. With MCP toolsets loaded, verify that delegation paths still execute correctly through sub-agents.
The broader context
Tool schema bloat is a recurring theme in Hermes optimization. A separate issue measured a default 30-tool setup at roughly 14,000 schema tokens per call, and a heavy multi-MCP setup at about 70,000 input tokens per turn. That volume exceeds the per-minute token cap of free-tier providers such as Cerebras and Groq, making some configurations unreachable rather than merely expensive.
stub_mode does not add round trips. It does not pick tools per turn. It reduces the weight of tools that are not meant to be called directly from the main agent. Whether it is sufficient on its own depends on how many MCP servers a deployment loads and whether direct enforcement of stub-only access is required.
Source: NousResearch/hermes-agent PR #48622, merged 2026-09.