- Icon.tsx: "warning" caution triangle (24px grid, 1.7 stroke, Lucide-style
rounded triangle + exclamation) matching the existing icon set.
- Composer.tsx: ModeOption extends Dropdown's Option with `caution` (warning
triangle before the label, themed via text-warnInk so it follows
light/dark) and `note` (a second, dimmer italic line under the
description). Bypass approvals carries the caution icon.
The Auto-Approve picker entry itself remains unshipped until the settings
pass gates it on the server-exposed auto_approve flag; its copy is decided
(owner, 2026-08-12): description "A reviewer clears routine actions;
doubtful ones still ask", note "Uses your session model for judgement - one
extra model call per check".
tsc clean; 111 GUI unit tests pass; rendered live and verified (note line
under Auto-Approve, warnInk triangle on Bypass).
The mode from ocw-context/docs/reviewed-auto-mode.md (rev. 4), v1 scope.
coworker/reviewer.py (new)
- The 8.3 prompt verbatim, cache-shaped: instructions + known world (folders
and remotes only) + user-message history in the stable prefix; this turn's
request and ONE action in the suffix.
- parse_verdict: any defect (empty, non-JSON, unknown verdict) -> unsure.
There is no parse path that results in execution (8.5).
- Reviewer.review never raises: provider errors and timeouts -> unsure.
Metering counters (checks / verdicts / tokens) for 1.7.
- AGENT_DENY_MESSAGE: the terse, non-diagnostic refusal the agent gets on a
deny; the full reason goes to the user only (8.4 asymmetry).
coworker/engine.py
- Reviewer consulted ONLY when: attached, mode is AUTO_APPROVE, session
explicitly attended (unset is_attended counts as NOT attended, so
automations can never be reviewed), fewer than two denials this turn.
- Consulted ONLY on decisions the gate marked needs_user - hard denies
never reach it, so it can only turn "ask" into "allow" (1.2).
- One action per request, fired concurrently for all of a turn's escalating
calls before the sequential authorize loop (8.6): a verdict cannot land
on the wrong action, and approval cards still reach the human one at a
time in call order.
- allow -> runs, audited with the reason. deny -> blocked; user event
carries the full reviewer reason + allow_anyway; agent message carries
only AGENT_DENY_MESSAGE. unsure -> today's card.
- Reviewer sees the user's words only, extracted mechanically from
role=user messages - never agent output, never tool results (4.4).
coworker/permissions.py
- Mode.AUTO renamed Mode.BYPASS_APPROVALS ("bypass-approvals"); legacy
"auto" still parses via _missing_ so configs, saved sessions, and the
golden decision table are untouched.
- Mode.AUTO_APPROVE ("auto-approve"): gate-identical to INTERACTIVE except
session grants ("always allow this ...") no longer auto-allow - they
route to the reviewer instead (1.5: out-of-band standing policy may skip
the judge; an in-flow click may not). Config allowlists still skip.
- _domain_allowed(include_session=False) checks the user-settings list only.
coworker/config.py: auto_approve flag, off by default, _GLOBAL_ONLY (a
cloned repo cannot hand itself a looser reviewer). agent.py attaches the
Reviewer only when the flag is on; without it AUTO_APPROVE behaves exactly
like INTERACTIVE.
server/manager.py: autonomy audit ranks auto-approve above interactive
(turning the reviewer on IS raising autonomy) and below bypass.
GUI: mode picker label "Full access" -> "Bypass approvals" (wire value
"auto" kept). Verified live against the real sidecar; e2e spec updated;
tsc and all 111 GUI unit tests pass.
Tests: tests/test_auto_approve.py (33) - gate behaviour per mode, fail-
closed parsing, prompt shape, deny asymmetry, retry guard, attended
gating, hard-deny isolation, per-action verdict landing, and that the
reviewer never sees agent prose. Permission suites + golden table: 146
passing unchanged.
coworker/session_facts.py: KnownWorld (roots + git remotes + hosts, captured
once before the agent acts) and SessionFacts (per-turn ingestion records).
Nothing consumes either in v1 - the known world will be rendered into the
reviewer prefix in step 2, ingestion exists so the v2 question ("would this
fact have changed a verdict?") is answerable from a shadow run replay
instead of re-argued.
- render() shows folders and remotes ONLY. Hostnames are held for DST-1 in
v2 but never rendered: a host list is only useful to a reviewer that can
do a suffix match, which models get wrong and Python does not.
- Freezing is the point: a remote added after capture stays unknown, so an
agent cannot make its own destination look familiar.
- Ingestion records the fact and the hostname, never the content and never
a full URL (a query string is exactly where a payload rides). Failed
calls record nothing. Local reads are deliberately excluded; the cost of
that exclusion is recorded in the spec.
- Engine hook is inert when session_facts is None; classification keys on
tool metadata category (web/connector/mcp) so new connectors are covered
the day they ship.
22 new tests. Permission suites (146) and the golden decision table pass
unchanged - this step alters no decision.
Spec: ocw-context/docs/reviewed-auto-mode.md Part 0, 2.4, Part 6 step 1.
1. Grant validation. POST /v1/inbox/{id}/resolve takes a raw resolution string
and approval_outcome() previously honoured whatever it named. The GUI
deliberately withholds the tool-wide "always allow" for run_shell (the
command-scoped grant is the narrower option), for save_skill (every skill
proposal gets its own review) and for connectors; Slack mirrors render only
approve/deny. So any local API caller could mint a session-wide,
any-argument shell grant -- a vocabulary the design says must not exist.
_grant_offered() now mirrors the card's own rules on the server and
downgrades an unoffered grant to a one-time approval, writing a
`grant_refused` audit row. Applied to the "always" channel vocabulary too,
so a Slack reply cannot mint what the in-app card would refuse. MCP tools
are covered alongside connectors: they are not category=connector but are
external, and the grant would be unbounded over every future argument.
A failed always_task mint is now audited rather than silently downgraded.
2. Autonomy transitions. Mode changes (WS set_mode) and the unattended toggle
were unrecorded, so "who turned on auto mode, and when" was unanswerable
from the audit store -- at odds with the per-call trail the engine keeps
everywhere else. Both now write an audit row tagged raised/lowered, so
autonomy increases can be filtered. set_unattended moves onto the manager
so REST and any future surface record it the same way; no-op flips are not
recorded.
15 new tests. test_server's two failures are pre-existing on the unmodified
tree (Windows file-permission errors in pathlib), unrelated to this change.
Design of record: ocw-context/docs/reviewed-auto-mode.md Part 3.
The old rule -- any shell operator disqualifies the whole command -- was wrong
in both directions, verified by running it:
find . -delete -> ALLOW (destructive, no prompt)
find . -exec rm {} + -> ALLOW (destructive, no prompt)
git status && git diff -> ask (two allowed reads, refused)
It judged punctuation rather than danger. `-delete` and `-exec` need no
separator, so a bare `find` prefix auto-ran them; meanwhile two independently
allowed reads were refused for containing `&&`.
Now:
- Constructs whose contents we cannot evaluate -- substitution, redirection,
variable expansion, grouping -- still disqualify the whole command, because
the unexamined tail after a prefix match must only ever be arguments.
- Compound commands are split on &&, ||, ;, |, |&, & and newlines, and EVERY
part must be independently covered by an allowlist entry.
- Parts that run code named in their arguments are never prefix-eligible:
argument executors (xargs, sudo, timeout, env, docker, npx, ssh...),
interpreters carrying inline code (python -c, bash -c, node -e), and
execution/deletion flags (-exec, -execdir, -delete, -ok).
- Matching stays on parsed words, so `git status` covers `git status -s` but
never `git statusfoo` or a bare `git`.
Splitting is textual and does not respect quoted separators. That is
deliberate: over-splitting yields MORE parts to justify, never fewer, so it
cannot loosen a verdict.
37 new tests including metamorphic cases (spacing, quoting, absolute program
path must not loosen `find . -delete`). Golden matrix: three rows flip as
intended, two added. 164 permission tests green.
Design of record: ocw-context/docs/reviewed-auto-mode.md Part 2 (CMD-1/3/4).
in-project files that execute later
Two floors, both mode-independent.
1. The settings files (config.toml, risk_overrides.json, workspace_trust.json,
unattended.json, coworker.db which holds the saved grants, secrets.json,
inbox_routing.json) cannot be modified by any tool, in any mode. The check
runs BEFORE mode, allowlists and every auto-approve path, and returns a hard
refusal rather than an approvable prompt -- loosening requires editing the
files out-of-band.
The escalation this blocks: approve one ordinary-looking `python setup.py`,
it quietly appends to risk_overrides.json, and every future session is more
permissive. That happens in the DEFAULT interactive mode, so the protection
cannot be a property of a sandbox or of any single mode.
Covered for write tools (resolved path), patch/diff blobs (path extracted),
and run_shell (command text). Shell matching is deliberately full-path only:
an earlier draft also matched bare filenames, which hard-denied any command
merely mentioning `secrets.json` -- caught by
test_shell_commands_not_auto_allowed_by_default, which reads that path with
`cat`. Shell is parser depth: it stops accidents and casual attempts, not a
determined adversary. That needs the OS sandbox (OPE-14).
2. Files inside the workspace that run on a later, innocuous action --
.git/hooks/**, .github/workflows/**, .vscode/tasks.json, .coworker/** --
stay writable but never WITHOUT a human. Auto mode, custom-mode auto_allow
and session "always allow" all lose to this. Otherwise accept-edits is a
clean bypass: write .git/hooks/pre-commit, then run an innocuous git commit.
19 new tests, incl. every mode parametrized and a lookalike case
(docs/pre-commit.md stays ordinary). 167 permission-related tests green.
test_standing_approvals::test_blocked_run_does_not_stall_other_tasks is an
intermittent pre-existing scheduler timing flake (fails ~1 in 3 on the
unmodified tree).
Design of record: ocw-context/docs/reviewed-auto-mode.md Part 3.
Three gate defects, each verified by direct execution before and after.
1. web_fetch was RiskClass.READ, so is_consequential() was False and evaluate()
returned allow on its third rung -- before any rule, mode or PDP, in EVERY
mode including plan/discuss. A URL's query string carries data outbound, so
this was an ungated egress path. New RiskClass.EGRESS covers model-chosen
network reads; web_search stays READ (fixed configured provider, not a
model-chosen host). Adds an allowed_domains allowlist (exact host or
subdomain; 'evil-python.org' never matches 'python.org'), a session-scoped
"always allow this domain" grant, and ApprovalOutcome.ALWAYS_DOMAIN.
2. A risk override could DOWNGRADE a built-in: marking write_file as read made
is_write False (skipping path scoping) and consequential False (skipping the
read-only gate) at once -- one settings line disabling two protections, in
every future session. Overrides may now only tighten a built-in write/exec/
egress tool; relaxing a metadata/MCP tool (the intended use) still works.
3. Path scoping read a literal "path" argument, so apply_patch and
apply_unified_diff -- whose paths live inside the patch/diff blob -- were
never scoped at all. write_paths() extracts them from the blob and scopes
every one; a write whose path cannot be located now fails closed to approval
rather than slipping through auto/custom unscoped.
allowed_domains is user-global only, alongside auto_allow: a cloned repo must
not be able to widen the agent's network reach.
Golden matrix: web_fetch interactive allow->ask, plan allow->deny, plus new
egress/patch rows (31 rows green). test_permissions_risk's override test
asserted the old downgrade behavior and is updated to the tightening rule.
Full suite: 22 failures, all pre-existing on the unmodified tree (boto3 absent,
Windows symlink privilege, Slack socket timeouts) -- none introduced here.
Design of record: ocw-context/docs/reviewed-auto-mode.md Part 3.
Freezes today's evaluate() verdict across 26 (mode, tool, args, grants)
situations, so any later permission change shows up as a row-diff. Four rows
are marked BASELINE-WRONG / BASELINE-ANNOYING on purpose: they record known
gaps (shell auto with no sandbox, find -delete and find -exec auto-allowed via
a find prefix, git status && git diff rejected for the operator, web_fetch
never gating in any mode). The PRs that fix these flip their rows here as the
visible proof.
Design of record: ocw-context/docs/reviewed-auto-mode.md Parts 3 and 7.
The live smoke exposed a harness trap: driving each turn through its own
asyncio.run() binds the engine asyncio primitives to the first loop, and
every later stream silently takes the interrupted path - full provider
replies persisted as empty assistant messages. The scripted smoke had
the same latent artifact and did not assert reply content, so it stayed
green. Now the whole scenario runs on ONE loop (like the real server)
and every turn asserts a real reply.
A long multi-turn session driven through the real SessionManager with a
forced 3k-token cap: repeated compactions advance the boundary, later
summaries fold the previous one in, the provider verifiably receives the
compacted view (summary block + verbatim tail, bounded) while the
canonical transcript keeps every turn, state survives a mid-conversation
rebuild, and the persisted record round-trips the final boundary.
Scripted stand-in for the live-model smoke: intent survival across a
real summarizer (prompt tuning) still needs a configured provider key.
Settings -> Models grows a Context compaction card next to Token savings:
the trigger % of the context window (10-95), the absolute token cap
(clamped 10k-2M), and the summarizer-model pin (default: the session's
own model). POST /v1/settings/compaction persists them; engines read the
knobs live per check, so changes apply to running sessions immediately.
The "context compacted" divider rides the existing notice machinery: the
persisted `compacted` notice replays on reload (itemsFromMessages) and
the live COMPACTED event appends the same info notice mid-turn. The
transcript itself stays intact - outbound-only by construction.
Covered by vitest (marker replay), a settings-card e2e (defaults +
clamped POSTs + model pin), and a mid-session divider e2e driven by the
fixtures' scripted `compacted` event.
Minimal engine footprint: a checkpoint at each iteration top (between tool
turns and before a new turn), the usage signal captured per round-trip
(context_tokens; chars/4 estimate when never reported), and
_outbound_messages consulting the boundary. The summarizer runs off-loop
through the normal provider router, so the Settings model pin is just an
id.
Failure policy per spec: retry once in both modes; attended sessions get
the Retry / Trim-oldest-10% prompt (via the ask_user plumbing, gated by an
is_attended callback the WS surface wires); unattended runs auto-trim and
continue — never parked on internal bookkeeping. Raw context-overflow 400s
from the main model route into the same policy, progress-guarded so a
still-overflowing model terminates in the error path.
CompactionState persists on the session record (new sqlite column, same
defensive parse as grants), so reloads keep the compacted view. A
persisted compacted notice + a new COMPACTED event mark the spot for
the GUI divider (rendered in commit 3).
Trigger math (usage signal, chars/4 estimate fallback, min(80% x window,
250k cap) with overridable knobs), boundary picking that never splits a
turn (user-message starts preferred, iteration starts inside a giant tool
loop), the 8-section summarizer prompt with the continuation contract,
mechanical working-state extraction from tool records, deterministic
user-message preservation, the trim-oldest fallback, outbound-view
application, and context-overflow detection. Injectable provider seam;
no engine changes yet.
Options accept {label, description, recommended, preview} objects (plain
strings unchanged — old sessions render as today's pills), and `questions`
groups up to 4 questions into one call, rendered as a stepper via the
header chips. Any option preview switches the card to a two-pane layout:
options left, monospace pane right, following hover/focus.
Grouped calls resolve with a JSON map keyed by header-or-question and
return {answers: {...}} to the agent (single stays {answer: ...});
a grouped item's first question doubles as its title/options so channel
mirrors and legacy surfaces degrade sensibly. Channel buttons use option
labels; grouped items mirror as text with the open-the-app hint.
openai_provider is now the compat workhorse (vendors, resellers, Ollama,
custom endpoints); the effort pin stays for GPT-5.6 reached through a
custom endpoint. base.py lists the current provider set.
_build_openai: no custom base_url -> OpenAIResponsesProvider; a custom
endpoint (Azure /openai/v1, vLLM, compat gateways) keeps the Chat
Completions OpenAIProvider, as do Ollama and every compat vendor. Verify
path (raw GET /models) and matrix ids are untouched.
/v1/chat/completions rejects function tools with any reasoning_effort
other than none on GPT-5.6, so native OpenAI has run with reasoning OFF.
OpenAIResponsesProvider speaks /v1/responses instead: reasoning + tools
at real effort, streamed reasoning summaries into the existing
reasoning_delta plumbing, and CoT continuity across tool round-trips via
store:false + encrypted reasoning replayed through a new _openai message
sidecar (same extras contract as _anthropic/_gemini). Not routed yet.
Coworkers remember durable things you tell them and use them in future sessions.
One Settings screen lists everything remembered - edit, delete, or stop new saves; standing instructions ride along.
Knowledge is session-stable, the save switch is per-message; sqlite gains a summary column via in-place migration.