Files
hyperframes/packages/cli/src/tts/manager.test.ts
T
James Russo 4a55bc8673 feat(cli): add --lang and auto-infer phonemizer locale from voice prefix (#351)
* feat(cli): add --lang and auto-infer phonemizer locale from voice prefix

`hyperframes tts` was calling Kokoro's `model.create(text, voice=, speed=)`
with no language argument, so Kokoro's default phonemizer (en-us) was
applied regardless of the voice selected. Picking `ef_dora` or `jf_alpha`
and feeding it Spanish or Japanese text produced English-phonemized
output.

Closes #349.

- `manager.ts`: add `SUPPORTED_LANGS`, `inferLangFromVoiceId`, and
  `isSupportedLang`. Attach a `defaultLang` field to every bundled voice
  and expand the bundled list with `ef_dora`, `ff_siwis`, `jf_alpha`,
  `zf_xiaobei` so `--list` surfaces multilingual options.
- `synthesize.ts`: accept optional `lang: SupportedLang` in
  `SynthesizeOptions`, forward it to the Python worker as `argv[7]`.
  The worker introspects `Kokoro.create`'s signature and only passes
  `lang=` when the installed kokoro-onnx version supports it. Returned
  metadata now includes `lang` and `langApplied` so callers can detect
  silent no-ops. Bump the cached script filename to `synth-v2.py` so
  existing installs pick up the new script automatically.
- `commands/tts.ts`: add `--lang, -l` with validation against
  `SUPPORTED_LANGS`. Resolution order is explicit `--lang` > inferred
  from voice prefix > `en-us`. When explicit lang disagrees with the
  voice-implied lang (legitimate for stylized accents), emit a
  dim-level hint; suppress under `--json`. When kokoro-onnx silently
  ignores the kwarg, log that too. Update `--list` with a new
  "Lang code" column and add multilingual examples.
- Tests: new `manager.test.ts` covering every supported prefix, the
  unknown-prefix fallback, case-insensitivity, `isSupportedLang`
  validation, and a regression guard that every bundled voice has a
  valid `defaultLang` matching its ID.
- Docs: `docs/packages/cli.mdx` and `skills/hyperframes/references/tts.md`
  updated with the flag, examples, the espeak-ng dependency note for
  non-English phonemization, and the voice-prefix → lang table.

Backward compatibility:
- English voices (a*/b* prefixes) continue to phonemize as en-us / en-gb
  — no change.
- Non-English voices now phonemize correctly by default (bug fix, not a
  regression).
- Older kokoro-onnx versions that don't know the `lang` kwarg keep
  working via signature introspection; the CLI logs a dim note if
  `--lang` was requested but ignored.

Verification:
- `bun --cwd packages/cli test` — 128 tests pass (incl. 17 new).
- `bunx oxlint` and `bunx oxfmt --check` clean on changed files.
- `bun run build` succeeds.
- `npx tsx packages/cli/src/cli.ts tts --help` / `--list` render cleanly;
  invalid `--lang` produces a clean error with the valid-codes list.

* refactor(cli): simplify tts --lang implementation

Post-review cleanup on #351. Net -21 lines.

- Drop `defaultLang` field + `makeVoice()` helper from VoiceInfo —
  compute via `inferLangFromVoiceId(v.id)` at read time in listVoices.
  The only reader was the --list table; caching the derived value on
  every voice added a self-consistency invariant we had to test.
- Drop redundant `lang` field from SynthesizeResult — caller already
  knows the requested lang since it passed it in; only `langApplied`
  carries information the caller can't derive.
- Use `errorBox` for --lang validation to match the house style in
  render.ts (other validation errors already use errorBox).
- Reuse existing `langList` module constant in the validation error
  instead of re-joining SUPPORTED_LANGS.
- Inline `DEFAULT_LANG` — used once in inferLangFromVoiceId.
- Trim WHAT-restating comments and the duplicate prefix-enumeration
  JSDoc on inferLangFromVoiceId (VOICE_PREFIX_LANG already carries
  per-row comments).
- Clean up orphaned `synth*.py` files in ~/.cache/hyperframes/tts
  when writing the current versioned script, so repeated upgrades
  don't leak files.
- Drop the `EN-US` case-sensitive-rejection test assertion — the CLI
  lowercases input before validation, so accepting mixed case is a
  feature, not a bug.

Tests: 16/16 in `manager.test.ts`, 127/127 full CLI suite pass.
Lint + format + typecheck clean.
2026-04-20 10:51:00 -07:00

62 lines
1.8 KiB
TypeScript

import { describe, expect, it } from "vitest";
import {
BUNDLED_VOICES,
SUPPORTED_LANGS,
inferLangFromVoiceId,
isSupportedLang,
} from "./manager.js";
describe("inferLangFromVoiceId", () => {
it.each([
["af_heart", "en-us"],
["am_adam", "en-us"],
["bf_emma", "en-gb"],
["bm_george", "en-gb"],
["ef_dora", "es"],
["ff_siwis", "fr-fr"],
["hf_alpha", "hi"],
["if_sara", "it"],
["jf_alpha", "ja"],
["pf_dora", "pt-br"],
["zf_xiaobei", "zh"],
])("maps voice %s to lang %s", (voiceId, expected) => {
expect(inferLangFromVoiceId(voiceId)).toBe(expected);
});
it("falls back to en-us for unknown prefixes", () => {
expect(inferLangFromVoiceId("xf_test")).toBe("en-us");
expect(inferLangFromVoiceId("")).toBe("en-us");
});
it("is case-insensitive on the prefix letter", () => {
expect(inferLangFromVoiceId("EF_dora")).toBe("es");
expect(inferLangFromVoiceId("ZF_xiaobei")).toBe("zh");
});
});
describe("isSupportedLang", () => {
it("accepts every value in SUPPORTED_LANGS", () => {
for (const lang of SUPPORTED_LANGS) {
expect(isSupportedLang(lang)).toBe(true);
}
});
it("rejects invalid or misspelled lang codes", () => {
expect(isSupportedLang("english")).toBe(false);
expect(isSupportedLang("de")).toBe(false);
expect(isSupportedLang("")).toBe(false);
});
});
describe("BUNDLED_VOICES", () => {
// --lang is user-facing, so the voice list must give users a working
// example in at least the most common non-English locales.
it("exposes at least one voice per non-English language", () => {
const langs = new Set(BUNDLED_VOICES.map((v) => inferLangFromVoiceId(v.id)));
expect(langs.has("es")).toBe(true);
expect(langs.has("fr-fr")).toBe(true);
expect(langs.has("ja")).toBe(true);
expect(langs.has("zh")).toBe(true);
});
});