mirror of
https://github.com/QuantumNous/new-api.git
synced 2026-09-11 14:41:21 +00:00
feat(relay): hosted-tool conversion fidelity, reasoning normalization, and billing usage integrity (#7137)
* feat(relaykit): preserve hosted tools across conversions - add protocol-neutral hosted-tool DTOs, conversion metadata, and loss policies - bridge citations, grounding metadata, and hosted-tool stream lifecycles - document the public conversion behavior and channel policy controls * refactor(relaykit): normalize reasoning and thinking intent - centralize provider-neutral reasoning intent, effort, and budget mappings - parse model suffixes at the host entry boundary while preserving provider-owned tails - keep adaptive Claude thinking and explicit zero-token compatibility consistent * fix(billing): preserve authoritative usage across relay hops - carry native BillingUsage sidecars through direct and streamed protocol bridges - merge partial and terminal usage monotonically with safe fallback settlement - retain cache metadata, penultimate usage, and per-call Gemini tool surcharges * feat(relay): bridge Responses with Claude and Gemini protocols - add direct request, response, and stream converters across supported relay formats - expose Claude count_tokens and Chat-to-Responses compatibility endpoints - carry conversion diagnostics through the host while retaining the curated public goldens * fix(relay): wire relaykit conversions into host channels - connect handlers, adaptors, and channel settings to the standalone conversion layer - keep model mapping, pricing identity, retries, and provider-specific suffix behavior aligned - ignore local audit artifacts and retain focused public regression coverage
This commit is contained in:
@@ -1,9 +1,12 @@
|
||||
// Package reasoning re-exports the pure model-name effort-suffix helpers,
|
||||
// which moved to the conversion kit (service/relayconvert/reasoning) as part
|
||||
// which moved to the conversion kit (relaykit/relayconvert/reasoning) as part
|
||||
// of the relaykit extraction. Host code keeps importing this path unchanged.
|
||||
package reasoning
|
||||
|
||||
import kitreasoning "github.com/QuantumNous/new-api/relaykit/relayconvert/reasoning"
|
||||
import (
|
||||
kitreasoning "github.com/QuantumNous/new-api/relaykit/relayconvert/reasoning"
|
||||
"github.com/QuantumNous/new-api/setting/model_setting"
|
||||
)
|
||||
|
||||
var (
|
||||
EffortSuffixes = kitreasoning.EffortSuffixes
|
||||
@@ -12,8 +15,13 @@ var (
|
||||
)
|
||||
|
||||
var (
|
||||
TrimEffortSuffix = kitreasoning.TrimEffortSuffix
|
||||
TrimEffortSuffixWithSuffixes = kitreasoning.TrimEffortSuffixWithSuffixes
|
||||
ParseOpenAIReasoningEffortFromModelSuffix = kitreasoning.ParseOpenAIReasoningEffortFromModelSuffix
|
||||
ParseDeepSeekV4ThinkingSuffix = kitreasoning.ParseDeepSeekV4ThinkingSuffix
|
||||
TrimEffortSuffixWithSuffixes = kitreasoning.TrimEffortSuffixWithSuffixes
|
||||
ParseDeepSeekV4ThinkingSuffix = kitreasoning.ParseDeepSeekV4ThinkingSuffix
|
||||
TrimGeminiThinkingSuffix = kitreasoning.TrimGeminiThinkingSuffix
|
||||
)
|
||||
|
||||
// ParseOpenAIReasoningEffortFromModelSuffix applies the host effort-tail
|
||||
// whitelist so real model IDs such as qwen-max are not treated as aliases.
|
||||
func ParseOpenAIReasoningEffortFromModelSuffix(modelName string) (string, string) {
|
||||
return kitreasoning.ParseOpenAIReasoningEffortFromModelSuffix(modelName, model_setting.ShouldPreserveEffortTail)
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user