Cutting Hermes Agent overhead with prompt-size and on-demand tools
- hermes
- tokens
- prompt-engineering
- optimization
A thread on X last week by IBuzovskyi laid out a practical Hermes Agent token-optimization path. The core observation is that Hermes sends a large fixed token payload on every turn: system prompt, skill descriptions, long-term memory, and tool/MCP schemas. That overhead is easy to ignore until you run hermes prompt-size and see 6,000–10,000+ tokens spent before your message even begins.
The thread's workflow is straightforward.
- Run
hermes prompt-sizeto get a baseline. - Disable unused skills and MCP servers.
- Clean
USER.mdandMEMORY.md, removing outdated context. - Switch
tool_searchto on-demand mode. - Re-run
hermes prompt-sizeand compare.
The highest-impact change in that list is on-demand tool loading. Instead of exposing every tool schema every turn, the agent loads schemas only when needed.
tools:
tool_search:
enabled: auto
The thread claims most users can cut overhead by 40–70% after one cleanup session. That claim is testable: measure before/after with hermes prompt-size.
There are also secondary guidelines for writing token-efficient skills. Short, high-signal descriptions beat long philosophical intros. Fewer broad skills beat many narrow ones. Removing redundant examples from skill files reduces repeated schema weight.
This is not a novel architectural idea, but it is one of the more concrete Hermes-specific optimization guides circulating this month. The actionable part is the measurement step. If you don't know your baseline, you can't verify improvement.
[^1]: IBuzovskyi. "Hermes Agent Token Optimization (Prompt + Fixed Overhead)" X post. September 2026.