The mode from ocw-context/docs/reviewed-auto-mode.md (rev. 4), v1 scope.
coworker/reviewer.py (new)
- The 8.3 prompt verbatim, cache-shaped: instructions + known world (folders
and remotes only) + user-message history in the stable prefix; this turn's
request and ONE action in the suffix.
- parse_verdict: any defect (empty, non-JSON, unknown verdict) -> unsure.
There is no parse path that results in execution (8.5).
- Reviewer.review never raises: provider errors and timeouts -> unsure.
Metering counters (checks / verdicts / tokens) for 1.7.
- AGENT_DENY_MESSAGE: the terse, non-diagnostic refusal the agent gets on a
deny; the full reason goes to the user only (8.4 asymmetry).
coworker/engine.py
- Reviewer consulted ONLY when: attached, mode is AUTO_APPROVE, session
explicitly attended (unset is_attended counts as NOT attended, so
automations can never be reviewed), fewer than two denials this turn.
- Consulted ONLY on decisions the gate marked needs_user - hard denies
never reach it, so it can only turn "ask" into "allow" (1.2).
- One action per request, fired concurrently for all of a turn's escalating
calls before the sequential authorize loop (8.6): a verdict cannot land
on the wrong action, and approval cards still reach the human one at a
time in call order.
- allow -> runs, audited with the reason. deny -> blocked; user event
carries the full reviewer reason + allow_anyway; agent message carries
only AGENT_DENY_MESSAGE. unsure -> today's card.
- Reviewer sees the user's words only, extracted mechanically from
role=user messages - never agent output, never tool results (4.4).
coworker/permissions.py
- Mode.AUTO renamed Mode.BYPASS_APPROVALS ("bypass-approvals"); legacy
"auto" still parses via _missing_ so configs, saved sessions, and the
golden decision table are untouched.
- Mode.AUTO_APPROVE ("auto-approve"): gate-identical to INTERACTIVE except
session grants ("always allow this ...") no longer auto-allow - they
route to the reviewer instead (1.5: out-of-band standing policy may skip
the judge; an in-flow click may not). Config allowlists still skip.
- _domain_allowed(include_session=False) checks the user-settings list only.
coworker/config.py: auto_approve flag, off by default, _GLOBAL_ONLY (a
cloned repo cannot hand itself a looser reviewer). agent.py attaches the
Reviewer only when the flag is on; without it AUTO_APPROVE behaves exactly
like INTERACTIVE.
server/manager.py: autonomy audit ranks auto-approve above interactive
(turning the reviewer on IS raising autonomy) and below bypass.
GUI: mode picker label "Full access" -> "Bypass approvals" (wire value
"auto" kept). Verified live against the real sidecar; e2e spec updated;
tsc and all 111 GUI unit tests pass.
Tests: tests/test_auto_approve.py (33) - gate behaviour per mode, fail-
closed parsing, prompt shape, deny asymmetry, retry guard, attended
gating, hard-deny isolation, per-action verdict landing, and that the
reviewer never sees agent prose. Permission suites + golden table: 146
passing unchanged.
in-project files that execute later
Two floors, both mode-independent.
1. The settings files (config.toml, risk_overrides.json, workspace_trust.json,
unattended.json, coworker.db which holds the saved grants, secrets.json,
inbox_routing.json) cannot be modified by any tool, in any mode. The check
runs BEFORE mode, allowlists and every auto-approve path, and returns a hard
refusal rather than an approvable prompt -- loosening requires editing the
files out-of-band.
The escalation this blocks: approve one ordinary-looking `python setup.py`,
it quietly appends to risk_overrides.json, and every future session is more
permissive. That happens in the DEFAULT interactive mode, so the protection
cannot be a property of a sandbox or of any single mode.
Covered for write tools (resolved path), patch/diff blobs (path extracted),
and run_shell (command text). Shell matching is deliberately full-path only:
an earlier draft also matched bare filenames, which hard-denied any command
merely mentioning `secrets.json` -- caught by
test_shell_commands_not_auto_allowed_by_default, which reads that path with
`cat`. Shell is parser depth: it stops accidents and casual attempts, not a
determined adversary. That needs the OS sandbox (OPE-14).
2. Files inside the workspace that run on a later, innocuous action --
.git/hooks/**, .github/workflows/**, .vscode/tasks.json, .coworker/** --
stay writable but never WITHOUT a human. Auto mode, custom-mode auto_allow
and session "always allow" all lose to this. Otherwise accept-edits is a
clean bypass: write .git/hooks/pre-commit, then run an innocuous git commit.
19 new tests, incl. every mode parametrized and a lookalike case
(docs/pre-commit.md stays ordinary). 167 permission-related tests green.
test_standing_approvals::test_blocked_run_does_not_stall_other_tasks is an
intermittent pre-existing scheduler timing flake (fails ~1 in 3 on the
unmodified tree).
Design of record: ocw-context/docs/reviewed-auto-mode.md Part 3.