feat(cli/telemetry): surface unrecognized agents in the agent_runtime=null bucket (#1978)

* feat(cli/telemetry): surface unrecognized agents in the agent_runtime=null bucket

agent_runtime is a closed allowlist: an agent we have no rule for collapses
to null with no trace of what it was, so ~18% of CLI users are unattributable
and new agents stay invisible until reverse-engineered by hand.

Add detectAgentHints(), a self-populating residual signal computed only for
the null bucket (gated off classified events):
- agent_hint: value of AGENT / AI_AGENT (the emerging self-identification
  convention; Crush and Goose set AGENT=<name>) — names agents the allowlist
  misses.
- term_program: raw TERM_PROGRAM (editor name) — catches the IDE-terminal
  class the same way the cursor/windsurf rules do.
- agent_env_hints: sorted, comma-joined "agent-ish" env-var KEY names present
  but matched by no vendor rule — a fingerprint that clusters by agent.

Privacy stays consistent with the existing "never read secret-shaped values"
stance: agent_env_hints emits key names only; the three value-reads are vars
whose sole purpose is non-secret identification, each passed through a strict
short-slug allowlist so anything long/spaced/secret-shaped is dropped.

Breaking down agent_hint / agent_env_hints filtered to agent_runtime IS NULL
AND is_tty=false gives a ranked leaderboard of new agents to promote into
VENDOR_RULES.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(cli/telemetry): guard agent_hint/term_program against short credential-shaped values

Review feedback (Magi, #1978): the short-slug allowlist in sanitizeHint() still
accepted short credential-shaped values (AGENT=sk-ant-api03,
AGENT=AKIAIOSFODNN7EXAMPLE, AGENT=github_pat_abc), so the "never emit a secret"
claim wasn't actually enforced — only overlong values were dropped.

Add a credential-shape guard on top of the slug allowlist:
- known token/credential prefixes (sk-, ghp_, github_pat_, akia, ya29, ...)
- any unbroken alphanumeric run >= 16 chars (key bodies, hex, base64-ish),
  while agent names segment on _/-/. and keep each run short.

Replace the single overlong-value test with the short credential shapes from the
review (parametrized) plus a positive case (gemini_managed_agent) proving real
multi-segment names still pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
James Russo
2026-07-05 22:31:13 -07:00
committed by GitHub
co-authored by Claude Opus 4.8
parent 9de41ce316
commit 232d591479
4 changed files with 229 additions and 1 deletions
+4
View File
@@ -92,6 +92,10 @@ export function trackEvent(
is_tty: sys.is_tty,
sandbox_runtime: sys.sandbox_runtime ?? undefined,
agent_runtime: sys.agent_runtime ?? undefined,
// New-agent discovery signals — populated only when agent_runtime is null.
agent_hint: sys.agent_hint ?? undefined,
term_program: sys.term_program ?? undefined,
agent_env_hints: sys.agent_env_hints ?? undefined,
},
timestamp: new Date().toISOString(),
});