mirror of
https://github.com/QuantumNous/new-api.git
synced 2026-09-11 14:41:21 +00:00
refactor: extract protocol conversion layer into standalone relaykit module (#6369)
* test(relayconvert): add golden snapshot matrix and relaykit boundary guard Phase 0 of the relaykit extraction plan: pin byte-level output of every registered (from,to) request/response/stream conversion route, and forbid kit-bound packages from growing host-only imports. * wip(relayconvert): drop gin.Context from converter signatures; add convmeta draft Phase 1 in progress: relayconvert now takes context.Context; host media resolver adapts gin.Context back at the service boundary. * refactor(relayconvert): decouple converters from RelayInfo, gin, and settings Phase 1 of the relaykit extraction plan: - converters now depend on convmeta.Meta (implemented by RelayInfo) instead of *relaycommon.RelayInfo; ClaudeConvertInfo and the format guesser move to convmeta with aliases left behind - host settings reach converters via a convmeta.Options snapshot built in RelayInfo.ConvOptions; no more model_setting/reasoning global reads inside the conversion layer - effort-suffix helpers move to service/relayconvert/reasoning (old package forwards); chat-to-responses upgrade policy moves to service (host routing logic, not conversion) - golden conversion matrix unchanged * test(relayconvert): tighten boundary — kit packages now free of gin/setting imports * refactor(dto): drop gin and logger dependencies Phase 2 (part 1): dto.Request.IsStream now takes *http.Request instead of *gin.Context (Gemini's impl reads query/path off the std request); dto's three logger calls become common.SysError. Boundary test allowlist is now empty — kit-bound packages import no gin/setting/logger/model. * refactor(kit): extract dependency-free kitutil; dto/types/relayconvert stop importing common Phase 2 of the relaykit extraction plan: - new service/relayconvert/kitutil holds the pure helpers the kit needs (JSON wrappers, pointer/string/uuid/timestamp utils, MaskSensitiveInfo, pluggable LogInfo/LogError hooks, Debug flag) - dto, types, and all relayconvert packages now use kitutil; their only remaining internal deps are dto/types/constant - common keeps every original symbol (MaskSensitiveInfo delegates to kitutil) so host code is untouched; main.go routes kit logging into common.SysLog/SysError and mirrors DebugEnabled - golden conversion matrix unchanged * refactor(kit): move EndpointType/FinishReason to types; OpenRouter dialect via Options Kit packages (dto/types/relayconvert/reasonmap) no longer import constant: - EndpointType and finish-reason values live in types; constant re-exports - the OpenRouter special-case in claude->openai request conversion reads Options.OpenRouterDialect, set by the host from the channel type; InitChannelMeta invalidates the cached snapshot on channel switch * refactor: extract relaykit submodule (dto/types/relayconvert/reasonmap) Phase 3 of the relaykit extraction plan: - new go module github.com/QuantumNous/new-api/relaykit containing dto (minus task family), types, relayconvert (with convmeta/kitutil/reasoning), and reasonmap; host consumes it via require + replace, go.work for dev - task-family dto (task/suno/midjourney/video) stays in the host dto package; dual-consumer host files alias it as taskdto - relaykit builds and tests standalone (GOWORK=off): no host imports, no gin, no DB, no settings - golden conversion matrix unchanged * build(docker): copy relaykit/go.mod before go mod download The local-replace submodule's go.mod must exist inside the build context for the main module graph to resolve. * fix: address relaykit extraction regressions * fix: address relaykit review regressions * docs: document Meta nil receiver contract * fix(relaykit): fail OpenAI→Claude conversion without max_tokens; reject negative default_max_tokens The Claude Messages API requires max_tokens (omitting it is a 400 "Field required"), but with a nil Options.Claude.DefaultMaxTokens hook the converters silently emitted a request the upstream is guaranteed to reject. Both OpenAI Chat and Responses → Claude conversions now return sharedclaude.ErrMissingMaxTokens when no path (client value, default hook, thinking-adapter floor) supplied one. Unreachable in the host, which always configures the hook. Host side, claude.default_max_tokens now rejects negative values at the option API before persisting — they would wrap into huge unsigned values during conversion. Zero stays allowed: the current API treats max_tokens: 0 as cache pre-warming. * fix: make Gemini safety settings read path race-free
This commit is contained in:
@@ -10,7 +10,7 @@ import (
|
||||
"strings"
|
||||
|
||||
"github.com/QuantumNous/new-api/common"
|
||||
"github.com/QuantumNous/new-api/types"
|
||||
"github.com/QuantumNous/new-api/relaykit/types"
|
||||
"github.com/samber/lo"
|
||||
"github.com/tidwall/gjson"
|
||||
"github.com/tidwall/sjson"
|
||||
|
||||
@@ -7,9 +7,9 @@ import (
|
||||
"testing"
|
||||
|
||||
common2 "github.com/QuantumNous/new-api/common"
|
||||
"github.com/QuantumNous/new-api/types"
|
||||
"github.com/QuantumNous/new-api/relaykit/types"
|
||||
|
||||
"github.com/QuantumNous/new-api/dto"
|
||||
"github.com/QuantumNous/new-api/relaykit/dto"
|
||||
"github.com/QuantumNous/new-api/setting/model_setting"
|
||||
"github.com/samber/lo"
|
||||
"github.com/stretchr/testify/require"
|
||||
|
||||
+138
-17
@@ -10,11 +10,12 @@ import (
|
||||
|
||||
"github.com/QuantumNous/new-api/common"
|
||||
"github.com/QuantumNous/new-api/constant"
|
||||
"github.com/QuantumNous/new-api/dto"
|
||||
"github.com/QuantumNous/new-api/pkg/billingexpr"
|
||||
relayconstant "github.com/QuantumNous/new-api/relay/constant"
|
||||
"github.com/QuantumNous/new-api/relaykit/dto"
|
||||
"github.com/QuantumNous/new-api/relaykit/relayconvert/convmeta"
|
||||
"github.com/QuantumNous/new-api/relaykit/types"
|
||||
"github.com/QuantumNous/new-api/setting/model_setting"
|
||||
"github.com/QuantumNous/new-api/types"
|
||||
|
||||
"github.com/gin-gonic/gin"
|
||||
"github.com/gorilla/websocket"
|
||||
@@ -28,22 +29,15 @@ type ThinkingContentInfo struct {
|
||||
}
|
||||
|
||||
const (
|
||||
LastMessageTypeNone = "none"
|
||||
LastMessageTypeText = "text"
|
||||
LastMessageTypeTools = "tools"
|
||||
LastMessageTypeThinking = "thinking"
|
||||
LastMessageTypeNone = convmeta.LastMessageTypeNone
|
||||
LastMessageTypeText = convmeta.LastMessageTypeText
|
||||
LastMessageTypeTools = convmeta.LastMessageTypeTools
|
||||
LastMessageTypeThinking = convmeta.LastMessageTypeThinking
|
||||
)
|
||||
|
||||
type ClaudeConvertInfo struct {
|
||||
LastMessagesType string
|
||||
Index int
|
||||
Usage *dto.Usage
|
||||
FinishReason string
|
||||
Done bool
|
||||
|
||||
ToolCallBaseIndex int
|
||||
ToolCallMaxIndexOffset int
|
||||
}
|
||||
// ClaudeConvertInfo now lives with the converters (convmeta); the alias keeps
|
||||
// host code and adaptors compiling unchanged.
|
||||
type ClaudeConvertInfo = convmeta.ClaudeConvertInfo
|
||||
|
||||
type RerankerInfo struct {
|
||||
Documents []any
|
||||
@@ -184,6 +178,9 @@ type RelayInfo struct {
|
||||
|
||||
StreamStatus *StreamStatus
|
||||
|
||||
// convOptions caches the converter settings snapshot (see ConvOptions).
|
||||
convOptions *convmeta.Options
|
||||
|
||||
ThinkingContentInfo
|
||||
TokenCountMeta
|
||||
*ClaudeConvertInfo
|
||||
@@ -239,6 +236,10 @@ func (info *RelayInfo) InitChannelMeta(c *gin.Context) {
|
||||
|
||||
info.ChannelMeta = channelMeta
|
||||
|
||||
// Channel identity feeds the converter options snapshot (e.g.
|
||||
// OpenRouterDialect); drop the cache so a cross-channel retry rebuilds it.
|
||||
info.convOptions = nil
|
||||
|
||||
// reset some fields based on channel meta
|
||||
// 重置某些字段,例如模型名称等
|
||||
if info.Request != nil {
|
||||
@@ -459,7 +460,7 @@ func genBaseRelayInfo(c *gin.Context, request dto.Request) *RelayInfo {
|
||||
isStream := false
|
||||
|
||||
if request != nil {
|
||||
isStream = request.IsStream(c)
|
||||
isStream = request.IsStream(c.Request)
|
||||
}
|
||||
c.Set(string(constant.ContextKeyIsStream), isStream)
|
||||
|
||||
@@ -679,13 +680,133 @@ func GenRelayInfoAlphaSearch(c *gin.Context, request *dto.AlphaSearchRequest) *R
|
||||
//}
|
||||
|
||||
func (info *RelayInfo) SetEstimatePromptTokens(promptTokens int) {
|
||||
if info == nil {
|
||||
return
|
||||
}
|
||||
info.estimatePromptTokens = promptTokens
|
||||
}
|
||||
|
||||
func (info *RelayInfo) GetEstimatePromptTokens() int {
|
||||
if info == nil {
|
||||
return 0
|
||||
}
|
||||
return info.estimatePromptTokens
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// convmeta.Meta implementation — the view format converters see. Keep these
|
||||
// thin: they only expose protocol state, never billing/user fields.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
var _ convmeta.Meta = (*RelayInfo)(nil)
|
||||
|
||||
func (info *RelayInfo) GetOriginModelName() string {
|
||||
if info == nil {
|
||||
return ""
|
||||
}
|
||||
return info.OriginModelName
|
||||
}
|
||||
|
||||
func (info *RelayInfo) GetUpstreamModelName() string {
|
||||
if info == nil || info.ChannelMeta == nil {
|
||||
return ""
|
||||
}
|
||||
return info.UpstreamModelName
|
||||
}
|
||||
|
||||
func (info *RelayInfo) HasChannelMeta() bool { return info != nil && info.ChannelMeta != nil }
|
||||
|
||||
func (info *RelayInfo) GetChannelID() int {
|
||||
if info == nil || info.ChannelMeta == nil {
|
||||
return 0
|
||||
}
|
||||
return info.ChannelId
|
||||
}
|
||||
|
||||
func (info *RelayInfo) GetChannelType() int {
|
||||
if info == nil || info.ChannelMeta == nil {
|
||||
return 0
|
||||
}
|
||||
return info.ChannelType
|
||||
}
|
||||
|
||||
func (info *RelayInfo) GetIsStream() bool {
|
||||
return info != nil && info.IsStream
|
||||
}
|
||||
|
||||
func (info *RelayInfo) GetReasoningEffort() string {
|
||||
if info == nil {
|
||||
return ""
|
||||
}
|
||||
return info.ReasoningEffort
|
||||
}
|
||||
|
||||
func (info *RelayInfo) SetReasoningEffort(effort string) {
|
||||
if info == nil {
|
||||
return
|
||||
}
|
||||
info.ReasoningEffort = effort
|
||||
}
|
||||
|
||||
func (info *RelayInfo) EnsureClaudeConvertInfo() *convmeta.ClaudeConvertInfo {
|
||||
if info == nil {
|
||||
return &convmeta.ClaudeConvertInfo{
|
||||
LastMessagesType: convmeta.LastMessageTypeNone,
|
||||
}
|
||||
}
|
||||
if info.ClaudeConvertInfo == nil {
|
||||
info.ClaudeConvertInfo = &convmeta.ClaudeConvertInfo{
|
||||
LastMessagesType: convmeta.LastMessageTypeNone,
|
||||
}
|
||||
}
|
||||
return info.ClaudeConvertInfo
|
||||
}
|
||||
|
||||
func (info *RelayInfo) GetSendResponseCount() int {
|
||||
if info == nil {
|
||||
return 0
|
||||
}
|
||||
return info.SendResponseCount
|
||||
}
|
||||
|
||||
func (info *RelayInfo) IncrSendResponseCount() {
|
||||
if info == nil {
|
||||
return
|
||||
}
|
||||
info.SendResponseCount++
|
||||
}
|
||||
|
||||
// ConvOptions snapshots host settings for the converters. Rebuilt on each
|
||||
// call site's first use; cached so one relay session sees one snapshot.
|
||||
func (info *RelayInfo) ConvOptions() *convmeta.Options {
|
||||
if info != nil && info.convOptions != nil {
|
||||
return info.convOptions
|
||||
}
|
||||
|
||||
claudeSettings := model_setting.GetClaudeSettings()
|
||||
geminiSettings := model_setting.GetGeminiSettings()
|
||||
options := &convmeta.Options{
|
||||
Claude: convmeta.ClaudeOptions{
|
||||
ThinkingAdapterEnabled: claudeSettings.ThinkingAdapterEnabled,
|
||||
ThinkingAdapterBudgetTokensPercentage: claudeSettings.ThinkingAdapterBudgetTokensPercentage,
|
||||
DefaultMaxTokens: claudeSettings.GetDefaultMaxTokens,
|
||||
},
|
||||
Gemini: convmeta.GeminiOptions{
|
||||
ThinkingAdapterEnabled: geminiSettings.ThinkingAdapterEnabled,
|
||||
ThinkingAdapterBudgetTokensPercentage: geminiSettings.ThinkingAdapterBudgetTokensPercentage,
|
||||
FunctionCallThoughtSignatureEnabled: geminiSettings.FunctionCallThoughtSignatureEnabled,
|
||||
SupportsImagine: model_setting.IsGeminiModelSupportImagine,
|
||||
SafetySetting: model_setting.GetGeminiSafetySetting,
|
||||
},
|
||||
OpenRouterDialect: info != nil && info.GetChannelType() == constant.ChannelTypeOpenRouter,
|
||||
PreserveThinkingSuffix: model_setting.ShouldPreserveThinkingSuffix,
|
||||
}
|
||||
if info != nil {
|
||||
info.convOptions = options
|
||||
}
|
||||
return options
|
||||
}
|
||||
|
||||
func (info *RelayInfo) SetFirstResponseTime() {
|
||||
if info.isFirstResponse {
|
||||
info.FirstResponseTime = time.Now()
|
||||
|
||||
@@ -0,0 +1,25 @@
|
||||
package common
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"github.com/QuantumNous/new-api/setting/model_setting"
|
||||
"github.com/stretchr/testify/assert"
|
||||
)
|
||||
|
||||
func TestRelayInfoConvOptionsUsesNormalizedGeminiSafetySettings(t *testing.T) {
|
||||
settings := model_setting.GetGeminiSettings()
|
||||
original := settings.SafetySettings
|
||||
t.Cleanup(func() {
|
||||
settings.SafetySettings = original
|
||||
})
|
||||
settings.SafetySettings = map[string]string{
|
||||
"HARM_CATEGORY_HATE_SPEECH": "",
|
||||
"HARM_CATEGORY_DANGEROUS_CONTENT": "BLOCK_ONLY_HIGH",
|
||||
}
|
||||
|
||||
options := (&RelayInfo{}).ConvOptions()
|
||||
|
||||
assert.Equal(t, "OFF", options.Gemini.SafetySetting("HARM_CATEGORY_HATE_SPEECH"))
|
||||
assert.Equal(t, "BLOCK_ONLY_HIGH", options.Gemini.SafetySetting("HARM_CATEGORY_DANGEROUS_CONTENT"))
|
||||
}
|
||||
@@ -3,7 +3,9 @@ package common
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"github.com/QuantumNous/new-api/types"
|
||||
"github.com/QuantumNous/new-api/relaykit/relayconvert/convmeta"
|
||||
"github.com/QuantumNous/new-api/relaykit/types"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
)
|
||||
|
||||
@@ -38,3 +40,41 @@ func TestRelayInfoGetFinalRequestRelayFormatNilReceiver(t *testing.T) {
|
||||
var info *RelayInfo
|
||||
require.Equal(t, types.RelayFormat(""), info.GetFinalRequestRelayFormat())
|
||||
}
|
||||
|
||||
func TestRelayInfoMetaTypedNilReceiver(t *testing.T) {
|
||||
var info *RelayInfo
|
||||
var meta convmeta.Meta = info
|
||||
|
||||
assert.Empty(t, meta.GetOriginModelName())
|
||||
assert.Empty(t, meta.GetUpstreamModelName())
|
||||
assert.False(t, meta.HasChannelMeta())
|
||||
assert.Zero(t, meta.GetChannelID())
|
||||
assert.Zero(t, meta.GetChannelType())
|
||||
assert.False(t, meta.GetIsStream())
|
||||
assert.Empty(t, meta.GetReasoningEffort())
|
||||
assert.Zero(t, meta.GetEstimatePromptTokens())
|
||||
assert.Zero(t, meta.GetSendResponseCount())
|
||||
|
||||
assert.NotPanics(t, func() {
|
||||
meta.SetReasoningEffort("high")
|
||||
meta.IncrSendResponseCount()
|
||||
meta.AppendRequestConversion(types.RelayFormatClaude)
|
||||
})
|
||||
|
||||
firstState := meta.EnsureClaudeConvertInfo()
|
||||
secondState := meta.EnsureClaudeConvertInfo()
|
||||
require.NotNil(t, firstState)
|
||||
require.NotNil(t, secondState)
|
||||
assert.Equal(t, convmeta.LastMessageTypeNone, firstState.LastMessagesType)
|
||||
assert.NotSame(t, firstState, secondState)
|
||||
|
||||
firstOptions := meta.ConvOptions()
|
||||
secondOptions := meta.ConvOptions()
|
||||
require.NotNil(t, firstOptions)
|
||||
require.NotNil(t, secondOptions)
|
||||
assert.NotSame(t, firstOptions, secondOptions)
|
||||
assert.NotNil(t, firstOptions.Claude.DefaultMaxTokens)
|
||||
assert.NotNil(t, firstOptions.Gemini.SupportsImagine)
|
||||
assert.NotNil(t, firstOptions.Gemini.SafetySetting)
|
||||
assert.NotNil(t, firstOptions.PreserveThinkingSuffix)
|
||||
}
|
||||
|
||||
@@ -1,31 +1,14 @@
|
||||
package common
|
||||
|
||||
import (
|
||||
"github.com/QuantumNous/new-api/dto"
|
||||
"github.com/QuantumNous/new-api/types"
|
||||
"github.com/QuantumNous/new-api/relaykit/relayconvert/convmeta"
|
||||
"github.com/QuantumNous/new-api/relaykit/types"
|
||||
)
|
||||
|
||||
// GuessRelayFormatFromRequest moved to convmeta with the converters; the
|
||||
// delegation keeps host callers unchanged.
|
||||
func GuessRelayFormatFromRequest(req any) (types.RelayFormat, bool) {
|
||||
switch req.(type) {
|
||||
case *dto.GeneralOpenAIRequest, dto.GeneralOpenAIRequest:
|
||||
return types.RelayFormatOpenAI, true
|
||||
case *dto.OpenAIResponsesRequest, dto.OpenAIResponsesRequest:
|
||||
return types.RelayFormatOpenAIResponses, true
|
||||
case *dto.ClaudeRequest, dto.ClaudeRequest:
|
||||
return types.RelayFormatClaude, true
|
||||
case *dto.GeminiChatRequest, dto.GeminiChatRequest:
|
||||
return types.RelayFormatGemini, true
|
||||
case *dto.EmbeddingRequest, dto.EmbeddingRequest:
|
||||
return types.RelayFormatEmbedding, true
|
||||
case *dto.RerankRequest, dto.RerankRequest:
|
||||
return types.RelayFormatRerank, true
|
||||
case *dto.ImageRequest, dto.ImageRequest:
|
||||
return types.RelayFormatOpenAIImage, true
|
||||
case *dto.AudioRequest, dto.AudioRequest:
|
||||
return types.RelayFormatOpenAIAudio, true
|
||||
default:
|
||||
return "", false
|
||||
}
|
||||
return convmeta.GuessRelayFormatFromRequest(req)
|
||||
}
|
||||
|
||||
func AppendRequestConversionFromRequest(info *RelayInfo, req any) {
|
||||
|
||||
@@ -29,9 +29,9 @@ type StreamErrorEntry struct {
|
||||
}
|
||||
|
||||
type StreamStatus struct {
|
||||
EndReason StreamEndReason
|
||||
EndError error
|
||||
endOnce sync.Once
|
||||
EndReason StreamEndReason
|
||||
EndError error
|
||||
endOnce sync.Once
|
||||
|
||||
mu sync.Mutex
|
||||
Errors []StreamErrorEntry
|
||||
|
||||
@@ -7,7 +7,7 @@ import (
|
||||
"strings"
|
||||
|
||||
"github.com/QuantumNous/new-api/common"
|
||||
"github.com/QuantumNous/new-api/dto"
|
||||
"github.com/QuantumNous/new-api/relaykit/dto"
|
||||
"github.com/QuantumNous/new-api/setting/operation_setting"
|
||||
)
|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@ import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/QuantumNous/new-api/dto"
|
||||
"github.com/QuantumNous/new-api/relaykit/dto"
|
||||
"github.com/QuantumNous/new-api/setting/operation_setting"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
|
||||
Reference in New Issue
Block a user