Skip to content
← all writing

Hermes Agent Adds Completion Contracts - /goal Now Judges Done Against Evidence, Not Vibes

  • hermes
  • goals
  • verification
  • v0.18
  • judgement-release
  • completion-contracts

Hermes Agent v0.18.0 shipped on July 1, 2026 as "The Judgement Release." The headline is bug clearance - 692 P0/P1 items closed in twelve days. The behavioral change that matters more for operators is inside the Goals section: /goal now carries a structured completion contract, and the standing-goal judge marks a goal done only when verification evidence is concrete.

The change is in PR #50501 by @teknium1. The contract idea is adapted from OpenAI Codex's /goal use-case; the implementation is independent and layered on Hermes' existing GoalManager and SessionDB architecture.

What the standing-goal loop does

A bare /goal <text> sets a standing objective. Hermes works on it, then a judge model runs after each turn and decides done or continue. If continue, the agent takes another step automatically. The loop is bounded by a turn budget (default 20).

The PR frames the problem this way, per teknium1 in #50501:

This directly tightens the most common /goal failure mode: premature completion or endless over-continuation on an underspecified objective.

A vague goal makes for vague judging. The judge can only check what you told it to want. A goal like "fix the parser" leaves three options - mark it done when the model says it's done (premature), keep going because nothing looks conclusive (over-continuation), or guess. All three are wrong.

The contract fields

The contract is five optional fields:

Field Meaning
outcome The single end state that must be true when done.
verification The specific test, command, or artifact that proves the outcome.
constraints What must not change or regress.
boundaries Which files, dirs, tools, or systems are in scope.
stop_when The condition under which Hermes should stop and ask for input.

Two prompts change when the contract is set: the continuation prompt tells the agent to target the verification surface and respect the constraints. The judge prompt decides done only when the verification criterion is met with concrete evidence (a command result, file excerpt, or test output) - not a "looks done" claim.

Two ways to set a contract

Let Hermes draft it - adapted from Codex's "let the agent draft the goal" tip:

/goal draft Migrate the auth service from session cookies to JWT

Hermes expands the one-liner into a full contract via the goal_judge auxiliary model, sets it, and shows the result for review. If the aux model is unavailable, it falls back to a plain free-form goal. Drafting never blocks setting a goal.

Write it inline with field: value lines:

/goal Migrate auth to JWT
verify: pytest tests/auth passes
constraints: keep the /login response shape unchanged
boundaries: only touch services/auth and its tests
stop when: a DB schema migration is required

The first non-field line is the goal headline. Recognized field prefixes (verify:, verified by:, constraints:, preserve:, boundaries:, scope:, stop when:, blocked:) populate the contract. A plain goal with an incidental colon (Fix bug: the parser drops commas) is not mangled - only known field prefixes are pulled out.

What the judge looks for

The judge prompt is rebuilt to require concrete evidence. The goals docs frame the new policy this way:

The judge prompt decides done only when the verification criterion is met with concrete evidence (a command result, file excerpt, test output) - not a loose "looks done" claim.

Behind that policy sits a profile-scoped verification evidence ledger from PR #52285. The ledger records foreground terminal test, lint, typecheck, and build results as scoped evidence (full vs targeted, pass vs fail). It's intentionally passive: it records evidence, not guarantees. A targeted test stays targeted; lint, typecheck, and build are classified separately.

Evidence matching covers the common real-world command spellings - pnpm test, bash scripts/run_tests.sh, uv run pytest - while avoiding echoed command text false positives. Storage is bounded: output summaries are capped, changed-path lists are capped, old per-session and workspace events are pruned, and evidence naturally expires after 30 days.

The same release ships a pre_verify hook from PR #55413 for wiring in custom project checks and a coding guidance config from PR #52296 for ad-hoc verification scripts.

Test and validation numbers

The PR ran its own test surface:

Test surface Result
tests/hermes_cli/test_goals.py 73/73 (+18 new tests)
Broader goal surface 42/42, 0 regressions
test_commands.py 156/156
Live E2E Set → persist → reload → judge prompts contract-aware → legacy row clean

ruff was clean across all four Python files touched.

How contracts survive session resume

The contract persists in SessionDB.state_meta alongside the goal. /resume reloads both intact. Old goal rows from before the feature ship with no contract and load unchanged - the upgrade is fully backward compatible. Contracts compose with /subgoal criteria: subgoals fold into the contract as extra criteria the judge must also satisfy. The continuation prompt stays a plain user-role message, so the contract doesn't invalidate any prompt-cache prefix.

Defaults and surface gating

verify-on-stop defaults OFF with a one-time v32 migration (#53552). The behavior is surface-aware: messaging surfaces (Telegram, Discord, Slack, Matrix, Signal, WhatsApp, SMS, iMessage) skip verify-on-stop because the interruption model differs. Doc-only edits also skip. /goal wait <pid> from #50503 parks the standing-goal loop on a background process and auto-resumes when the process exits.

Why this matters

The release codename is The Judgement Release, and the completion-contract change gives that name operational weight. Before, the judge was the model's say-so. After, "done" is whatever you said it was, with whatever evidence you said would prove it, and the judge has a ledger to look at.

[^1]: Nous Research. "Hermes Agent v0.18.0 - The Judgement Release" X. July 1, 2026. [^2]: teknium1. "feat(goals): completion contracts for /goal - evidence-based judging (#50501)" GitHub. June 2026. [^3]: OutThisLife. "feat(agent): record coding verification evidence (#52285)" GitHub. June 2026. [^4]: Nous Research. "Goals documentation" GitHub. 2026. [^5]: Teknium. "How is everyone liking The Judgement release of Hermes Agent??" X. July 3, 2026.

Termagotchi
_

Ryan Underdown

Autodidact. Rarely listens to advice.

Follow on X @catamarammed or GitHub @underdown