* feat(plugin): add MiniMax-H3 /v2 video generation to the hailuo task plugin
MiniMax-H3 speaks a different contract from the other Hailuo models, so the
hailuo task plugin now branches on the upstream model instead of adding a Go
adaptor:
- submit builds /v2/video_generation with a multimodal `content` array
(text, first/last frame images, reference video/audio, or a full
`metadata.content` passthrough), an explicit `ratio`, and 768P/2K
resolutions; `metadata.callback_url` and `metadata.aigc_watermark` pass
through
- query uses /v2/query/video_generation/{task_id} and parses the
`{"task": {...}}` envelope, falling back to the /v1 shapes for every other
model
- the /v2 result is a public CDN URL, so its artifact is proxied
credentialless instead of through /v1/files/download
- request bounds (duration 4-15, resolution 768P/2K, ratio whitelist, at most
2 frame images and 9/3/3 reference images/videos/audios) are enforced while
the request body is built, which the host runs during validation, so an
out-of-range duration is rejected with a 400 before it can become a billing
multiplier
- duration and resolution are reported as usage facts only. Like the rest of
this plugin, extractUsage returns no billing ratios, so per-call pricing is
flat and 2K/duration pricing is expressed through the model's tiered billing
expression over those facts.
Query hooks are driver hooks and are documented to receive `ctx.model` and
`ctx.upstreamModel`, but polling has no relay info and never populated them.
The polling and realtime-fetch call sites now carry the persisted task model
properties and the plugin adaptor maps them onto the query context, with
`upstreamModel` falling back to the origin name for tasks submitted without a
channel mapping.
* fix(plugin): validate Hailuo H3 requests and errors
* feat(relaykit): preserve hosted tools across conversions
- add protocol-neutral hosted-tool DTOs, conversion metadata, and loss policies
- bridge citations, grounding metadata, and hosted-tool stream lifecycles
- document the public conversion behavior and channel policy controls
* refactor(relaykit): normalize reasoning and thinking intent
- centralize provider-neutral reasoning intent, effort, and budget mappings
- parse model suffixes at the host entry boundary while preserving provider-owned tails
- keep adaptive Claude thinking and explicit zero-token compatibility consistent
* fix(billing): preserve authoritative usage across relay hops
- carry native BillingUsage sidecars through direct and streamed protocol bridges
- merge partial and terminal usage monotonically with safe fallback settlement
- retain cache metadata, penultimate usage, and per-call Gemini tool surcharges
* feat(relay): bridge Responses with Claude and Gemini protocols
- add direct request, response, and stream converters across supported relay formats
- expose Claude count_tokens and Chat-to-Responses compatibility endpoints
- carry conversion diagnostics through the host while retaining the curated public goldens
* fix(relay): wire relaykit conversions into host channels
- connect handlers, adaptors, and channel settings to the standalone conversion layer
- keep model mapping, pricing identity, retries, and provider-specific suffix behavior aligned
- ignore local audit artifacts and retain focused public regression coverage
* fix(relay): bound the wait for upstream response headers (fixes unbounded heap growth)
The relay transport sets a dial timeout, a TLS handshake timeout and an expect-continue
timeout, but nothing bounds how long it waits for the upstream *response headers* after
the request has been written. An upstream that accepts the connection and then never
answers -- without sending FIN/RST, which is what happens when a NAT/firewall silently
drops the flow or the provider hangs -- parks the goroutine in
net/http.(*persistConn).roundTrip forever.
That goroutine keeps the whole request alive, which in practice means three copies of the
request body stay reachable for the lifetime of the process: the raw bytes from
io.ReadAll in CreateBodyStorageFromReader, the decoded messages held as json.RawMessage,
and the re-marshalled upstream body from common.Marshal. BodyStorageCleanup cannot help
here: it runs after c.Next() returns, and for these requests c.Next() never returns.
Measured on v1.0.0-rc.23 in production (see #6947 for the full evidence):
- 23 goroutines stuck in persistConn.roundTrip on a single 40h-old instance,
blocked between 353 and 1894 minutes (5.9h to 31.5h)
- 96.9% of the live heap, sampled after a forced GC, attributable to those three
body copies (HeapAlloc 892 MiB surviving three GC cycles; HeapObjects dropping
30x while bytes dropped only 25%)
- the live floor grows with uptime: 33.7 MiB at 0.1h, 89.2 at 13.8h, 510.0 at 40.1h,
955.2 at 146.8h, OOMKilled at 172.9h -- same image, same config, same load
Doubling the memory limit and adding GOMEMLIMIT only moved the OOM from 132h to 172.9h.
RELAY_TIMEOUT (http.Client.Timeout) cannot be used for this: it covers the whole response
read and would cut legitimate long streaming calls, which is why it defaults to 0.
ResponseHeaderTimeout only bounds the wait for the headers; streaming after they arrive is
unaffected.
The default is deliberately generous. Non-streaming upstreams usually send the response
headers only once generation has finished, so the value has to leave room for a long
completion. 1800s is 12x shorter than the shortest hang observed here while leaving
several times the headroom a normal non-streaming request needs; 0 restores the previous
unbounded behaviour.
The assignment goes next to the other transport.* lines rather than inside the else
branch: newRelayHTTPTransport() normally takes the http.DefaultTransport.Clone() path,
and DefaultTransport does not set ResponseHeaderTimeout either.
This repo already sets ResponseHeaderTimeout on its other outbound transports
(controller/model_sync.go, controller/ratio_sync.go); the relay path appears to have
been missed.
Refs #6947. Likely also the root cause of #6731, which reported the same symptom
(production OOM on /v1/responses after ~64h) but was closed for template reasons.
* review: clamp overflowing timeout values and switch the test to testify
Addresses the two CodeRabbit findings on this PR.
Overflow (common/init.go:113): a RELAY_RESPONSE_HEADER_TIMEOUT beyond ~9.2e9 seconds
overflows time.Duration and can wrap into a *tiny positive* timeout, which would cut
every relay request instead of only the stuck ones. The value is now clamped before the
conversion, with regression tests for both the negative and the overflowing input.
I did not add fail-on-startup validation for negative values, for two reasons: the
existing `if seconds > 0` guard already treats them as "disabled", and the neighbouring
env-driven timeouts in this file are less strict still -- RelayIdleConnTimeout is
converted with no guard at all. Failing startup on a bad value would be a behaviour
change out of step with the rest of the file; happy to add it if you'd prefer that
direction repo-wide.
Test style: switched to testify (require.Equal / require.Zero / require.Positive), which
is what every other test under service/ uses.
go build, go vet and go test ./common/... ./service/... pass.
(`go build ./...` fails on the `web/dist` embed both with and without this change -- the
frontend bundle is not checked in.)
* feat(token): support custom auto group order
* feat(keys): enhance auto group presentation
* fix(keys): rework Auto flow border and compact inherited order
The Auto group highlight previously tinted the whole control surface
with a gradient and animated only a 1px top sweep, which read as a
background color rather than a flowing border. Replace it with a
border-only effect: an aria-hidden, pointer-events-none overlay whose
conic gradient is masked down to a thin ring hugging the rounded
perimeter, so the highlight travels around all four edges and corners
every 3.2s. The interior stays neutral with a restrained static
primary border and glow; prefers-reduced-motion hides the moving
layer while keeping the static emphasis.
The inherited global Auto order also rendered as spacious two-line
rows with circular sequence markers, wasting drawer space. Render it
as a compact wrapping strip of one-line chips (index, name, ratio
badge) with descriptions kept accessible via title and sr-only text,
scrolling only past a much smaller max height.
Custom add/remove/reorder editing, empty-array inheritance semantics,
and the submit payload are unchanged.
* fix(keys): preserve Auto inheritance and unify effects
* refactor(keys): temporarily disable AutoGroupBadge in api-key-group-cell
Follow-up to #6518 (issue #6480) addressing three review findings:
- Document and lock in arrears semantics for the wallet Reserve top-up:
when an auto-group retry lands on a more expensive group, the full
reservation delta is deducted unconditionally (balance may go
negative), mirroring settlement, so the logged pre-consumed quota
always reconciles with the actual balance movement. Genuine DB
errors still fail the attempt with update_data_error. Subscription
funding keeps its insufficient-quota behavior: subscriptions enforce
a hard used<=total cap and do not support arrears.
- PriceData.FreeModel is cleared when a retry switches from a free
group to a paid one, keeping it consistent with the billing session
created at that point.
- getChannel refreshes GroupRatioInfo only after channel selection
succeeds, and the retry loop records the channel in use_channel
before PrepareTieredBillingForSelectedGroup can fail.
* test(relayconvert): add golden snapshot matrix and relaykit boundary guard
Phase 0 of the relaykit extraction plan: pin byte-level output of every
registered (from,to) request/response/stream conversion route, and
forbid kit-bound packages from growing host-only imports.
* wip(relayconvert): drop gin.Context from converter signatures; add convmeta draft
Phase 1 in progress: relayconvert now takes context.Context; host media
resolver adapts gin.Context back at the service boundary.
* refactor(relayconvert): decouple converters from RelayInfo, gin, and settings
Phase 1 of the relaykit extraction plan:
- converters now depend on convmeta.Meta (implemented by RelayInfo) instead
of *relaycommon.RelayInfo; ClaudeConvertInfo and the format guesser move
to convmeta with aliases left behind
- host settings reach converters via a convmeta.Options snapshot built in
RelayInfo.ConvOptions; no more model_setting/reasoning global reads inside
the conversion layer
- effort-suffix helpers move to service/relayconvert/reasoning (old package
forwards); chat-to-responses upgrade policy moves to service (host routing
logic, not conversion)
- golden conversion matrix unchanged
* test(relayconvert): tighten boundary — kit packages now free of gin/setting imports
* refactor(dto): drop gin and logger dependencies
Phase 2 (part 1): dto.Request.IsStream now takes *http.Request instead of
*gin.Context (Gemini's impl reads query/path off the std request); dto's
three logger calls become common.SysError. Boundary test allowlist is now
empty — kit-bound packages import no gin/setting/logger/model.
* refactor(kit): extract dependency-free kitutil; dto/types/relayconvert stop importing common
Phase 2 of the relaykit extraction plan:
- new service/relayconvert/kitutil holds the pure helpers the kit needs
(JSON wrappers, pointer/string/uuid/timestamp utils, MaskSensitiveInfo,
pluggable LogInfo/LogError hooks, Debug flag)
- dto, types, and all relayconvert packages now use kitutil; their only
remaining internal deps are dto/types/constant
- common keeps every original symbol (MaskSensitiveInfo delegates to
kitutil) so host code is untouched; main.go routes kit logging into
common.SysLog/SysError and mirrors DebugEnabled
- golden conversion matrix unchanged
* refactor(kit): move EndpointType/FinishReason to types; OpenRouter dialect via Options
Kit packages (dto/types/relayconvert/reasonmap) no longer import constant:
- EndpointType and finish-reason values live in types; constant re-exports
- the OpenRouter special-case in claude->openai request conversion reads
Options.OpenRouterDialect, set by the host from the channel type;
InitChannelMeta invalidates the cached snapshot on channel switch
* refactor: extract relaykit submodule (dto/types/relayconvert/reasonmap)
Phase 3 of the relaykit extraction plan:
- new go module github.com/QuantumNous/new-api/relaykit containing dto
(minus task family), types, relayconvert (with convmeta/kitutil/reasoning),
and reasonmap; host consumes it via require + replace, go.work for dev
- task-family dto (task/suno/midjourney/video) stays in the host dto
package; dual-consumer host files alias it as taskdto
- relaykit builds and tests standalone (GOWORK=off): no host imports,
no gin, no DB, no settings
- golden conversion matrix unchanged
* build(docker): copy relaykit/go.mod before go mod download
The local-replace submodule's go.mod must exist inside the build context
for the main module graph to resolve.
* fix: address relaykit extraction regressions
* fix: address relaykit review regressions
* docs: document Meta nil receiver contract
* fix(relaykit): fail OpenAI→Claude conversion without max_tokens; reject negative default_max_tokens
The Claude Messages API requires max_tokens (omitting it is a 400
"Field required"), but with a nil Options.Claude.DefaultMaxTokens hook
the converters silently emitted a request the upstream is guaranteed to
reject. Both OpenAI Chat and Responses → Claude conversions now return
sharedclaude.ErrMissingMaxTokens when no path (client value, default
hook, thinking-adapter floor) supplied one. Unreachable in the host,
which always configures the hook.
Host side, claude.default_max_tokens now rejects negative values at the
option API before persisting — they would wrap into huge unsigned values
during conversion. Zero stays allowed: the current API treats
max_tokens: 0 as cache pre-warming.
* fix: make Gemini safety settings read path race-free
When an upstream error response parses as valid JSON but yields no usable
error message (e.g. an aggregator gateway returning {"error":{"message":""}}),
RelayErrorHandler previously produced a bare "bad response status code N"
error with no trace of the original body, making the failure undiagnosable.
Log the body preview in that case, mirroring the existing behavior for
unparseable bodies.
* fix(playground): resolve auto group model listing
- merge and deduplicate available models in configured auto group order.
- reuse special usable group rules and add model filtering regression coverage.
* refactor: extract GetGroupsEnabledModels to dedupe group model expansion
* fix(i18n): clarify Go regex and field passthrough copy
* feat(channel): support Codex upstream model discovery
* Revert "fix(i18n): clarify Go regex and field passthrough copy"
This reverts commit d63d7975db.
When a function call was already registered under its output_index key
via response.output_item.added, the synthetic events built from the
terminal response.completed output carry no output_index and resolve to
a different item-based key. ensureToolForEvent then created a second
tool index and resent the full arguments, so Chat Completions clients
received the same tool call twice.
Reuse the tool registered under itemIDToKey/callIDToKey before creating
a new one, and alias the new key to the existing tool.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Parse OpenAI's native cache_write_tokens (chat prompt_tokens_details /
responses input_tokens_details), bill it at the cache-creation ratio, and
clamp the uncached prompt remainder at zero since cached + cache-write can
exceed prompt_tokens. Propagate the field through chat/responses/claude
format conversions and tiered expression billing (cc variable).
* refactor: consolidate relay protocol converters
* refactor relayconvert text converters
* feat: refine relay converters and advanced custom routing
* refactor: enhance logging and add thought signature handling for Gemini requests
* refactor: enhance channel cache and pricing endpoint handling for advanced custom models
* feat: preserve billing usage semantics
* feat: add protocol-aware billing usage
* Delete useless files
* chore: update action versions in workflow files
* chore: update Docker action versions in workflow files
* fix: harden billing usage settlement and hot-path route matching
- estimate Gemini completion tokens locally when billable usageMetadata is
prompt-only but output content was received (e.g. client aborts the stream
before the final chunk), and rebuild the attached billing_usage as estimated
so settlement does not bill zero output tokens
- guard NewClaudeMessagesBillingUsage against all-zero ClaudeUsage, matching
the OpenAI/Gemini constructors, so a zero billing_usage cannot override a
non-zero top-level usage during settlement
- cache compiled advanced-custom route model regexes; they run on the request
hot path and were recompiled per request
- move the effectiveBillingUsage remap to PostTextConsumeQuota only, and
document that calculateTextQuotaSummary expects remapped usage
- document the updatePricingLock -> channelSyncLock lock ordering that
InitChannelCache/CacheUpdateChannel rely on, and the aux-struct pitfall in
GeminiChatResponse.UnmarshalJSON
controller/task_video.go and service/pre_consume_quota.go were already
removed in ba25ba88f and 116004fd4 (logic lives in service/task_polling.go
and BillingSession now), then brought back as stale copies by a42b39760.
Both have zero callers.
controller/swag_video.go is a leftover swag stub: the swaggo pipeline is
gone and docs/openapi/relay.json already documents these routes.
Thread int32 saturation clamps from tiered settlement and video task
recompute into the consume/task logs under admin_info, so oversized or
malformed billing inputs stay auditable. Clamp negative audio duration
before token conversion and gate the saturation UI markers on admin.
Bound max-tokens fields across all relay format validators, saturate
tiered-expression rounding and audio/tool/task token conversions, and
route legacy remix ratios through the guarded setter.
Bound user-supplied count/duration parameters at request validation,
route ratio multipliers through guarded setters, and use saturating
int conversions in all quota math paths.
* fix(openai): harden Chat-to-Responses compatibility
Add a shared Responses-to-Chat stream state machine and use it from the OpenAI relay path. Preserve assistant text alongside tool calls, bind tool argument deltas by output_index, map incomplete finish reasons, support reasoning/custom tool events, and buffer upstream SSE for non-stream Chat clients.
Add deterministic service tests and relay SSE tests for the conversion path.
Related to #5745.
* refactor: rename openaicompat to relayconvert for improved clarity
* feat(gemini): support responses request conversion
* feat: add responses to chat conversion support
* fix: harden responses chat conversion edge cases
Add a shared Responses-to-Chat stream state machine and use it from the OpenAI relay path. Preserve assistant text alongside tool calls, bind tool argument deltas by output_index, map incomplete finish reasons, support reasoning/custom tool events, and buffer upstream SSE for non-stream Chat clients.
Add deterministic service tests and relay SSE tests for the conversion path.
Related to #5745.
Async task usage logs (LogQuotaData node dimension) were recorded
under whichever node happened to poll the task to completion, not the
node that submitted it. For token/adaptor-billed video tasks the
pre-deduction is often 0, so the entire quota landed on the last
polling node.
Snapshot common.NodeName into TaskPrivateData at submit time and use
it when writing the settlement consume log; fall back to the current
node when empty so existing tasks stay compatible.
* feat: add system instance reporting
* feat: show system instance resources
* fix: update translations for heartbeat messages in Russian and Vietnamese