Commit Graph
86 Commits
Author SHA1 Message Date
Devika Verma 330010cc66 compaction: harden the smoke against per-turn event loops (OPE-27)
The live smoke exposed a harness trap: driving each turn through its own
asyncio.run() binds the engine asyncio primitives to the first loop, and
every later stream silently takes the interrupted path - full provider
replies persisted as empty assistant messages. The scripted smoke had
the same latent artifact and did not assert reply content, so it stayed
green. Now the whole scenario runs on ONE loop (like the real server)
and every turn asserts a real reply.
2026-07-29 18:00:38 +05:30
Devika Verma 0bf9b87800 compaction: repeated-compaction smoke through the manager (OPE-27 4/4)
A long multi-turn session driven through the real SessionManager with a
forced 3k-token cap: repeated compactions advance the boundary, later
summaries fold the previous one in, the provider verifiably receives the
compacted view (summary block + verbatim tail, bounded) while the
canonical transcript keeps every turn, state survives a mid-conversation
rebuild, and the persisted record round-trips the final boundary.

Scripted stand-in for the live-model smoke: intent survival across a
real summarizer (prompt tuning) still needs a configured provider key.
2026-07-29 16:23:26 +05:30
Devika Verma 4fa8acffed compaction: Settings overrides + GUI divider (OPE-27 3/4)
Settings -> Models grows a Context compaction card next to Token savings:
the trigger % of the context window (10-95), the absolute token cap
(clamped 10k-2M), and the summarizer-model pin (default: the session's
own model). POST /v1/settings/compaction persists them; engines read the
knobs live per check, so changes apply to running sessions immediately.

The "context compacted" divider rides the existing notice machinery: the
persisted `compacted` notice replays on reload (itemsFromMessages) and
the live COMPACTED event appends the same info notice mid-turn. The
transcript itself stays intact - outbound-only by construction.

Covered by vitest (marker replay), a settings-card e2e (defaults +
clamped POSTs + model pin), and a mid-session divider e2e driven by the
fixtures' scripted `compacted` event.
2026-07-29 16:20:24 +05:30
Devika Verma f08a3c425b compaction: engine hook, failure policy, persistence (OPE-27 2/4)
Minimal engine footprint: a checkpoint at each iteration top (between tool
turns and before a new turn), the usage signal captured per round-trip
(context_tokens; chars/4 estimate when never reported), and
_outbound_messages consulting the boundary. The summarizer runs off-loop
through the normal provider router, so the Settings model pin is just an
id.

Failure policy per spec: retry once in both modes; attended sessions get
the Retry / Trim-oldest-10% prompt (via the ask_user plumbing, gated by an
is_attended callback the WS surface wires); unattended runs auto-trim and
continue — never parked on internal bookkeeping. Raw context-overflow 400s
from the main model route into the same policy, progress-guarded so a
still-overflowing model terminates in the error path.

CompactionState persists on the session record (new sqlite column, same
defensive parse as grants), so reloads keep the compacted view. A
persisted compacted notice + a new COMPACTED event mark the spot for
the GUI divider (rendered in commit 3).
2026-07-29 16:13:14 +05:30
Devika Verma 028d42eb3b compaction: pure module + tests (OPE-27 1/4)
Trigger math (usage signal, chars/4 estimate fallback, min(80% x window,
250k cap) with overridable knobs), boundary picking that never splits a
turn (user-message starts preferred, iteration starts inside a giant tool
loop), the 8-section summarizer prompt with the continuation contract,
mechanical working-state extraction from tool records, deterministic
user-message preservation, the trim-oldest fallback, outbound-view
application, and context-overflow detection. Injectable provider seam;
no engine changes yet.
2026-07-29 16:05:30 +05:30
Rohit Prasad f96ad4c8e6 Merge pull request #304 from andrewyng/rpTokenMetering
Token usage metering: per-session usage in the app + Anthropic prompt caching
2026-07-28 12:34:12 -07:00
Rohit C Prasad 8674e301a8 Pin mcp<2 — 2.0.0 removed streamablehttp_client
CI installs latest mcp; today's 2.0.0 release breaks every MCP-client import.
2026-07-28 11:58:31 -07:00
Rohit C Prasad 27311cd97f Usage popover: 'Uncached input' when a cache split exists
Input rows then read as components: uncached + cache reads + cache
writes = Total input; plain 'Input' stays for cacheless backends.
2026-07-28 11:48:21 -07:00
Rohit C Prasad d1524b3376 Usage popover: label rows as session totals
Section header 'Session totals' + pluralized cache rows make the
cumulative semantics explicit.
2026-07-28 11:42:44 -07:00
Rohit C Prasad a35b505659 Usage popover: add cumulative Total input row
Fresh + cache read + cache write — the session's billed input volume;
shown only when a cache split exists.
2026-07-28 11:39:15 -07:00
Rohit C Prasad 92c1833223 Usage popover: one field per line
Stacked label/value rows instead of the wrapped inline stats; values are
session sums per model (fresh input split from the cache rows).
2026-07-28 11:30:57 -07:00
Rohit C Prasad 8991d303e0 Enable prompt caching on the Anthropic provider
Two ephemeral breakpoints per request (last system block, final message's
last block) so append-only history re-reads the prior turns' cache;
outbound-only, persisted history stays clean.
2026-07-27 22:47:30 -07:00
Rohit C Prasad 7a108b25f9 Show per-session token usage in the composer
Quiet chip (context-fill meter + session total) opening a per-model
input/output/cache breakdown popover; accumulation from live events,
rebuilt from persisted sidecars on load; unit + e2e coverage.
2026-07-27 21:15:30 -07:00
Rohit C Prasad 979badbd3c Meter token usage across all model providers
Normalized TokenUsage (input/output/cache split) captured in every provider's
stream and complete paths, persisted as an assistant-message sidecar and sent
on the assistant_message event; matrix gains verified context-window sizes.
2026-07-27 21:01:31 -07:00
Rohit Prasad 3766805d10 Merge pull request #259 from andrewyng/rpModelProviders
Add AWS Bedrock, Google Vertex AI, and OpenRouter model providers
2026-07-27 20:41:54 -07:00
Rohit Prasad d3863966c9 Update README with badge from trendshift 2026-07-27 15:15:54 -07:00
Rohit C Prasad 33d3efd3b2 Vertex: countTokens verify, global-location host, honest region help
Model listing 403/404s under plain ADC; countTokens is free and proves
project+location+API in one call. Verified live: Gemini (global), Qwen MaaS (us-south1).
2026-07-27 12:58:16 -07:00
Rohit C Prasad f281b29ff1 Provider auth redesign: joined segments, method panels, Vertex methods
Segmented track + inset per-method panel with its own Test & save footer.
Vertex gains the same treatment: Google Cloud login (default), service account,
and API key (express mode, Gemini-only with a clear error elsewhere).
2026-07-26 21:10:59 -07:00
Rohit C Prasad b3a2b130d2 Vendor Bedrock, Vertex, and OpenRouter brand marks
Same MIT lobe-icons set as the existing gallery logos.
2026-07-26 21:04:49 -07:00
Rohit C Prasad b719227a9a Bedrock settings: one auth method at a time
'Connect with' segmented choice (API key / profile / IAM keys) shows only that
method's fields; non-selected fields are dropped at build so stale values can't leak.
2026-07-25 22:55:28 -07:00
Rohit C Prasad 2a882c09d3 Add Nemotron Super 3 120B to the Bedrock model matrix
Live-verified on Converse; tool calls are sequential, so parallel stays off.
2026-07-25 22:11:08 -07:00
Rohit C Prasad 333f589c80 Support Bedrock API keys (bearer auth)
New optional field: paste the console-generated key, no CLI/IAM setup needed.
Takes precedence over SigV4 credentials, matching boto3; live-tested on Converse.
2026-07-25 21:58:24 -07:00
Rohit C Prasad 50463bba00 Package boto3 for Bedrock and pin google-auth
New [bedrock] extra; desktop builds and CI install it; PyInstaller collects
boto3/botocore so the lazy import works in the bundled sidecar.
2026-07-25 16:23:56 -07:00
Rohit C Prasad 241af5e15f GUI: multi-field provider credentials and add-model family dropdown
Test button and saved pill follow the required-secret field (or the first field for
cloud providers); Bedrock/Vertex add-model rows get a family selector.
2026-07-25 16:22:27 -07:00
Rohit C Prasad 050cc894e7 Add Google Vertex AI provider with per-family dispatch
gemini/ and claude/ ids reuse the native providers; openweight/ goes through the
MaaS OpenAI-compat endpoint with an auto-refreshed google-auth bearer.
Credentials: service-account JSON or Application Default Credentials.
2026-07-25 16:19:00 -07:00
Rohit C Prasad 8cb9524f1f Add AWS Bedrock provider with per-family dispatch
claude/ ids use Anthropic's native Bedrock client; everything else goes via Converse.
Credentials: explicit keys, named profile (incl. SSO), or the ambient AWS chain.
2026-07-25 16:15:18 -07:00
Rohit C Prasad ee495b9006 Add OpenRouter as an OpenAI-compatible reseller provider
Descriptor + curated matrix rows + sk-or- key auto-detect (server and GUI).
2026-07-25 16:07:33 -07:00
Rohit Prasad db93d75bf6 Merge pull request #101 from andrewyng/meta-muse-spark
Add Meta Model API provider with Muse Spark 1.1
2026-07-25 01:27:39 -07:00
Rohit C Prasad fc6ce501dd Add Meta Model API provider with Muse Spark 1.1 2026-07-25 01:21:07 -07:00
Rohit Prasad f467c4ca73 Merge pull request #115 from andrewyng/rpSlackApprovalOwners
Harden Slack approval handling
2026-07-25 00:38:50 -07:00
Rohit P 7656952692 security: harden Slack approval handling 2026-07-24 22:33:13 -07:00
Rohit Prasad 56b3c62864 Merge pull request #107 from andrewyng/rpSecurityHardeningFollowup
Security Hardening
2026-07-24 22:05:43 -07:00
Rohit P ac83bc0490 security: complete local access protections 2026-07-24 18:44:39 -07:00
Rohit P 3f5ac872ca security: refine workspace trust controls 2026-07-24 18:44:39 -07:00
Rohit P 8ee0a0d082 security: strengthen session handling 2026-07-24 18:44:39 -07:00
Rohit P 9cc2761998 security: tighten permission handling 2026-07-24 18:44:39 -07:00
Rohit Prasad 54b4bfd82d Merge pull request #103 from andrewyng/fix-ci-ollama-probe
Pin the Ollama probe in the curated-models test
2026-07-24 17:37:29 -07:00
Rohit Prasad dd33f2cab0 Merge pull request #49 from fahadsiddiqui/harden/boundary-oauth-shell-ws
fix: harden local trust boundaries (shell allowlist, MCP OAuth loopback, WS ingestion)
2026-07-24 15:43:59 -07:00
Rohit C Prasad 2325728f3f Pin the Ollama probe in the curated-models test
The picker gates ollama:* on a live local probe, so this passed only where Ollama runs and failed in CI.
The probe's own behaviour stays covered by test_ollama_models_gated_on_liveness.
2026-07-24 12:48:38 -07:00
Fahad Siddiqui 657cf03460 fix: harden local trust boundaries (shell allowlist, MCP OAuth loopback, WS ingestion)
Boundary-hardening pass addressing three audit findings on the local sidecar.

Shell command allowlist (andrewyng/openworker#28):
- Replace prefix-string matching in PermissionEngine._command_allowed with
  argv-aware matching: reject any command containing shell operators
  (; & | > < ` $( ( and newlines) before consulting the allowlist, then require
  the allowlisted entry's tokens to be an exact argv prefix. This closes the
  auto-run bypass where an allowlisted "git status" also auto-ran
  "git status && rm -rf ~", pipes, redirection, and command substitution.
- Drop language interpreters / package managers (python, python3, node, npm,
  npx) from DEFAULT_ALLOWED_COMMANDS — allowlisting an interpreter allowlists
  arbitrary code (python3 -c "..."), defeating approval gating. Read-only
  inspection commands and pytest remain.

MCP OAuth loopback (andrewyng/openworker#29):
- Verify the OAuth state at the loopback boundary. The MCP SDK already validates
  state (compare_digest), so this is not a CSRF fix but defense-in-depth: capture
  the state from the authorize URL and have deliver_callback ignore a callback
  whose state does not match WITHOUT consuming the pending future, so a stray or
  forged local hit can no longer abort a user's in-progress sign-in. Falls back to
  prior accept-any behavior when no state was captured.

WebSocket ingestion caps (andrewyng/openworker#38):
- Bound a single user_message frame in the session WS loop: max text length,
  max attachment count, and max total attachment bytes. Oversized frames get a
  visible error frame and are dropped instead of being buffered into a turn; the
  socket stays alive. Guards the unauthenticated loopback socket against cheap
  memory spikes.

Tests:
- Allowlist: reject operator chaining (8 variants), argv-boundary matching, and
  interpreters-not-auto-allowed-by-default.
- OAuth: state extraction, and mismatched/missing state ignored without consuming
  the flow while the matching state still resolves it.
- WS: oversized text and too-many-attachments rejected with an error frame, and a
  normal message still works afterwards.

Full suite: 865 passed (1 pre-existing unrelated failure in
test_provider_router::test_manager_curated_models, present on origin/main).
2026-07-24 01:41:05 +05:00
Yashas 4766e59c47 Keep ripgrep searches out of generated directories (#10)
* Skip generated directories in ripgrep searches

* Apply ripgrep exclusions after user globs
2026-07-23 12:29:19 -07:00
Rohit P 4ffc73f1e8 README: Minor updates. 2026-07-23 11:26:19 -07:00
Rohit C Prasad f7c70a2478 README: add how-it-works diagram from the website 2026-07-23 09:29:00 -07:00
Rohit C Prasad 062b1b12b3 Update README 2026-07-23 09:23:56 -07:00
Rohit C Prasad 0e48499075 README: link the website 2026-07-23 09:14:26 -07:00
Rohit C Prasad 9fc4bc43e6 Rewrite README: download links, capabilities, architecture, run-from-source 2026-07-23 09:01:29 -07:00
Rohit C Prasad da1d25373c Prepare app release 0.1.6: version bump v0.1.6 2026-07-23 08:49:30 -07:00
Rohit C Prasad 27d72e5c5e Add Kimi K2.7 Code via Together to the model matrix 2026-07-23 08:41:21 -07:00
Rohit C Prasad eae5fbd8d0 Fix Anthropic extended thinking for current model families
Adaptive thinking for 4.6+/newest models, budgets kept for older ones.
Safety-refusal fallback to Opus for the newest tier; add Inkling via Together.
2026-07-23 08:25:47 -07:00
Rohit C Prasad ba99978e03 Enable Claude extended thinking by default, drop the settings field
Fixed 8192 budget; the provider profile key stays a hidden override (0 = off).
Per-turn composer control is future work.
2026-07-23 07:57:04 -07:00