refactor: extract protocol conversion layer into standalone relaykit module (#6369)

* test(relayconvert): add golden snapshot matrix and relaykit boundary guard

Phase 0 of the relaykit extraction plan: pin byte-level output of every
registered (from,to) request/response/stream conversion route, and
forbid kit-bound packages from growing host-only imports.

* wip(relayconvert): drop gin.Context from converter signatures; add convmeta draft

Phase 1 in progress: relayconvert now takes context.Context; host media
resolver adapts gin.Context back at the service boundary.

* refactor(relayconvert): decouple converters from RelayInfo, gin, and settings

Phase 1 of the relaykit extraction plan:
- converters now depend on convmeta.Meta (implemented by RelayInfo) instead
  of *relaycommon.RelayInfo; ClaudeConvertInfo and the format guesser move
  to convmeta with aliases left behind
- host settings reach converters via a convmeta.Options snapshot built in
  RelayInfo.ConvOptions; no more model_setting/reasoning global reads inside
  the conversion layer
- effort-suffix helpers move to service/relayconvert/reasoning (old package
  forwards); chat-to-responses upgrade policy moves to service (host routing
  logic, not conversion)
- golden conversion matrix unchanged

* test(relayconvert): tighten boundary — kit packages now free of gin/setting imports

* refactor(dto): drop gin and logger dependencies

Phase 2 (part 1): dto.Request.IsStream now takes *http.Request instead of
*gin.Context (Gemini's impl reads query/path off the std request); dto's
three logger calls become common.SysError. Boundary test allowlist is now
empty — kit-bound packages import no gin/setting/logger/model.

* refactor(kit): extract dependency-free kitutil; dto/types/relayconvert stop importing common

Phase 2 of the relaykit extraction plan:
- new service/relayconvert/kitutil holds the pure helpers the kit needs
  (JSON wrappers, pointer/string/uuid/timestamp utils, MaskSensitiveInfo,
  pluggable LogInfo/LogError hooks, Debug flag)
- dto, types, and all relayconvert packages now use kitutil; their only
  remaining internal deps are dto/types/constant
- common keeps every original symbol (MaskSensitiveInfo delegates to
  kitutil) so host code is untouched; main.go routes kit logging into
  common.SysLog/SysError and mirrors DebugEnabled
- golden conversion matrix unchanged

* refactor(kit): move EndpointType/FinishReason to types; OpenRouter dialect via Options

Kit packages (dto/types/relayconvert/reasonmap) no longer import constant:
- EndpointType and finish-reason values live in types; constant re-exports
- the OpenRouter special-case in claude->openai request conversion reads
  Options.OpenRouterDialect, set by the host from the channel type;
  InitChannelMeta invalidates the cached snapshot on channel switch

* refactor: extract relaykit submodule (dto/types/relayconvert/reasonmap)

Phase 3 of the relaykit extraction plan:
- new go module github.com/QuantumNous/new-api/relaykit containing dto
  (minus task family), types, relayconvert (with convmeta/kitutil/reasoning),
  and reasonmap; host consumes it via require + replace, go.work for dev
- task-family dto (task/suno/midjourney/video) stays in the host dto
  package; dual-consumer host files alias it as taskdto
- relaykit builds and tests standalone (GOWORK=off): no host imports,
  no gin, no DB, no settings
- golden conversion matrix unchanged

* build(docker): copy relaykit/go.mod before go mod download

The local-replace submodule's go.mod must exist inside the build context
for the main module graph to resolve.

* fix: address relaykit extraction regressions

* fix: address relaykit review regressions

* docs: document Meta nil receiver contract

* fix(relaykit): fail OpenAI→Claude conversion without max_tokens; reject negative default_max_tokens

The Claude Messages API requires max_tokens (omitting it is a 400
"Field required"), but with a nil Options.Claude.DefaultMaxTokens hook
the converters silently emitted a request the upstream is guaranteed to
reject. Both OpenAI Chat and Responses → Claude conversions now return
sharedclaude.ErrMissingMaxTokens when no path (client value, default
hook, thinking-adapter floor) supplied one. Unreachable in the host,
which always configures the hook.

Host side, claude.default_max_tokens now rejects negative values at the
option API before persisting — they would wrap into huge unsigned values
during conversion. Zero stays allowed: the current API treats
max_tokens: 0 as cache pre-warming.

* fix: make Gemini safety settings read path race-free
This commit is contained in:
Calcium-Ion
2026-07-27 15:56:21 +08:00
committed by GitHub
parent f51dd4d808
commit 86ac0f7745
368 changed files with 7144 additions and 1594 deletions
+1 -1
View File
@@ -10,7 +10,7 @@ import (
"strings"
"github.com/QuantumNous/new-api/common"
"github.com/QuantumNous/new-api/types"
"github.com/QuantumNous/new-api/relaykit/types"
"github.com/samber/lo"
"github.com/tidwall/gjson"
"github.com/tidwall/sjson"
+2 -2
View File
@@ -7,9 +7,9 @@ import (
"testing"
common2 "github.com/QuantumNous/new-api/common"
"github.com/QuantumNous/new-api/types"
"github.com/QuantumNous/new-api/relaykit/types"
"github.com/QuantumNous/new-api/dto"
"github.com/QuantumNous/new-api/relaykit/dto"
"github.com/QuantumNous/new-api/setting/model_setting"
"github.com/samber/lo"
"github.com/stretchr/testify/require"
+138 -17
View File
@@ -10,11 +10,12 @@ import (
"github.com/QuantumNous/new-api/common"
"github.com/QuantumNous/new-api/constant"
"github.com/QuantumNous/new-api/dto"
"github.com/QuantumNous/new-api/pkg/billingexpr"
relayconstant "github.com/QuantumNous/new-api/relay/constant"
"github.com/QuantumNous/new-api/relaykit/dto"
"github.com/QuantumNous/new-api/relaykit/relayconvert/convmeta"
"github.com/QuantumNous/new-api/relaykit/types"
"github.com/QuantumNous/new-api/setting/model_setting"
"github.com/QuantumNous/new-api/types"
"github.com/gin-gonic/gin"
"github.com/gorilla/websocket"
@@ -28,22 +29,15 @@ type ThinkingContentInfo struct {
}
const (
LastMessageTypeNone = "none"
LastMessageTypeText = "text"
LastMessageTypeTools = "tools"
LastMessageTypeThinking = "thinking"
LastMessageTypeNone = convmeta.LastMessageTypeNone
LastMessageTypeText = convmeta.LastMessageTypeText
LastMessageTypeTools = convmeta.LastMessageTypeTools
LastMessageTypeThinking = convmeta.LastMessageTypeThinking
)
type ClaudeConvertInfo struct {
LastMessagesType string
Index int
Usage *dto.Usage
FinishReason string
Done bool
ToolCallBaseIndex int
ToolCallMaxIndexOffset int
}
// ClaudeConvertInfo now lives with the converters (convmeta); the alias keeps
// host code and adaptors compiling unchanged.
type ClaudeConvertInfo = convmeta.ClaudeConvertInfo
type RerankerInfo struct {
Documents []any
@@ -184,6 +178,9 @@ type RelayInfo struct {
StreamStatus *StreamStatus
// convOptions caches the converter settings snapshot (see ConvOptions).
convOptions *convmeta.Options
ThinkingContentInfo
TokenCountMeta
*ClaudeConvertInfo
@@ -239,6 +236,10 @@ func (info *RelayInfo) InitChannelMeta(c *gin.Context) {
info.ChannelMeta = channelMeta
// Channel identity feeds the converter options snapshot (e.g.
// OpenRouterDialect); drop the cache so a cross-channel retry rebuilds it.
info.convOptions = nil
// reset some fields based on channel meta
// 重置某些字段,例如模型名称等
if info.Request != nil {
@@ -459,7 +460,7 @@ func genBaseRelayInfo(c *gin.Context, request dto.Request) *RelayInfo {
isStream := false
if request != nil {
isStream = request.IsStream(c)
isStream = request.IsStream(c.Request)
}
c.Set(string(constant.ContextKeyIsStream), isStream)
@@ -679,13 +680,133 @@ func GenRelayInfoAlphaSearch(c *gin.Context, request *dto.AlphaSearchRequest) *R
//}
func (info *RelayInfo) SetEstimatePromptTokens(promptTokens int) {
if info == nil {
return
}
info.estimatePromptTokens = promptTokens
}
func (info *RelayInfo) GetEstimatePromptTokens() int {
if info == nil {
return 0
}
return info.estimatePromptTokens
}
// ---------------------------------------------------------------------------
// convmeta.Meta implementation — the view format converters see. Keep these
// thin: they only expose protocol state, never billing/user fields.
// ---------------------------------------------------------------------------
var _ convmeta.Meta = (*RelayInfo)(nil)
func (info *RelayInfo) GetOriginModelName() string {
if info == nil {
return ""
}
return info.OriginModelName
}
func (info *RelayInfo) GetUpstreamModelName() string {
if info == nil || info.ChannelMeta == nil {
return ""
}
return info.UpstreamModelName
}
func (info *RelayInfo) HasChannelMeta() bool { return info != nil && info.ChannelMeta != nil }
func (info *RelayInfo) GetChannelID() int {
if info == nil || info.ChannelMeta == nil {
return 0
}
return info.ChannelId
}
func (info *RelayInfo) GetChannelType() int {
if info == nil || info.ChannelMeta == nil {
return 0
}
return info.ChannelType
}
func (info *RelayInfo) GetIsStream() bool {
return info != nil && info.IsStream
}
func (info *RelayInfo) GetReasoningEffort() string {
if info == nil {
return ""
}
return info.ReasoningEffort
}
func (info *RelayInfo) SetReasoningEffort(effort string) {
if info == nil {
return
}
info.ReasoningEffort = effort
}
func (info *RelayInfo) EnsureClaudeConvertInfo() *convmeta.ClaudeConvertInfo {
if info == nil {
return &convmeta.ClaudeConvertInfo{
LastMessagesType: convmeta.LastMessageTypeNone,
}
}
if info.ClaudeConvertInfo == nil {
info.ClaudeConvertInfo = &convmeta.ClaudeConvertInfo{
LastMessagesType: convmeta.LastMessageTypeNone,
}
}
return info.ClaudeConvertInfo
}
func (info *RelayInfo) GetSendResponseCount() int {
if info == nil {
return 0
}
return info.SendResponseCount
}
func (info *RelayInfo) IncrSendResponseCount() {
if info == nil {
return
}
info.SendResponseCount++
}
// ConvOptions snapshots host settings for the converters. Rebuilt on each
// call site's first use; cached so one relay session sees one snapshot.
func (info *RelayInfo) ConvOptions() *convmeta.Options {
if info != nil && info.convOptions != nil {
return info.convOptions
}
claudeSettings := model_setting.GetClaudeSettings()
geminiSettings := model_setting.GetGeminiSettings()
options := &convmeta.Options{
Claude: convmeta.ClaudeOptions{
ThinkingAdapterEnabled: claudeSettings.ThinkingAdapterEnabled,
ThinkingAdapterBudgetTokensPercentage: claudeSettings.ThinkingAdapterBudgetTokensPercentage,
DefaultMaxTokens: claudeSettings.GetDefaultMaxTokens,
},
Gemini: convmeta.GeminiOptions{
ThinkingAdapterEnabled: geminiSettings.ThinkingAdapterEnabled,
ThinkingAdapterBudgetTokensPercentage: geminiSettings.ThinkingAdapterBudgetTokensPercentage,
FunctionCallThoughtSignatureEnabled: geminiSettings.FunctionCallThoughtSignatureEnabled,
SupportsImagine: model_setting.IsGeminiModelSupportImagine,
SafetySetting: model_setting.GetGeminiSafetySetting,
},
OpenRouterDialect: info != nil && info.GetChannelType() == constant.ChannelTypeOpenRouter,
PreserveThinkingSuffix: model_setting.ShouldPreserveThinkingSuffix,
}
if info != nil {
info.convOptions = options
}
return options
}
func (info *RelayInfo) SetFirstResponseTime() {
if info.isFirstResponse {
info.FirstResponseTime = time.Now()
@@ -0,0 +1,25 @@
package common
import (
"testing"
"github.com/QuantumNous/new-api/setting/model_setting"
"github.com/stretchr/testify/assert"
)
func TestRelayInfoConvOptionsUsesNormalizedGeminiSafetySettings(t *testing.T) {
settings := model_setting.GetGeminiSettings()
original := settings.SafetySettings
t.Cleanup(func() {
settings.SafetySettings = original
})
settings.SafetySettings = map[string]string{
"HARM_CATEGORY_HATE_SPEECH": "",
"HARM_CATEGORY_DANGEROUS_CONTENT": "BLOCK_ONLY_HIGH",
}
options := (&RelayInfo{}).ConvOptions()
assert.Equal(t, "OFF", options.Gemini.SafetySetting("HARM_CATEGORY_HATE_SPEECH"))
assert.Equal(t, "BLOCK_ONLY_HIGH", options.Gemini.SafetySetting("HARM_CATEGORY_DANGEROUS_CONTENT"))
}
+41 -1
View File
@@ -3,7 +3,9 @@ package common
import (
"testing"
"github.com/QuantumNous/new-api/types"
"github.com/QuantumNous/new-api/relaykit/relayconvert/convmeta"
"github.com/QuantumNous/new-api/relaykit/types"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
@@ -38,3 +40,41 @@ func TestRelayInfoGetFinalRequestRelayFormatNilReceiver(t *testing.T) {
var info *RelayInfo
require.Equal(t, types.RelayFormat(""), info.GetFinalRequestRelayFormat())
}
func TestRelayInfoMetaTypedNilReceiver(t *testing.T) {
var info *RelayInfo
var meta convmeta.Meta = info
assert.Empty(t, meta.GetOriginModelName())
assert.Empty(t, meta.GetUpstreamModelName())
assert.False(t, meta.HasChannelMeta())
assert.Zero(t, meta.GetChannelID())
assert.Zero(t, meta.GetChannelType())
assert.False(t, meta.GetIsStream())
assert.Empty(t, meta.GetReasoningEffort())
assert.Zero(t, meta.GetEstimatePromptTokens())
assert.Zero(t, meta.GetSendResponseCount())
assert.NotPanics(t, func() {
meta.SetReasoningEffort("high")
meta.IncrSendResponseCount()
meta.AppendRequestConversion(types.RelayFormatClaude)
})
firstState := meta.EnsureClaudeConvertInfo()
secondState := meta.EnsureClaudeConvertInfo()
require.NotNil(t, firstState)
require.NotNil(t, secondState)
assert.Equal(t, convmeta.LastMessageTypeNone, firstState.LastMessagesType)
assert.NotSame(t, firstState, secondState)
firstOptions := meta.ConvOptions()
secondOptions := meta.ConvOptions()
require.NotNil(t, firstOptions)
require.NotNil(t, secondOptions)
assert.NotSame(t, firstOptions, secondOptions)
assert.NotNil(t, firstOptions.Claude.DefaultMaxTokens)
assert.NotNil(t, firstOptions.Gemini.SupportsImagine)
assert.NotNil(t, firstOptions.Gemini.SafetySetting)
assert.NotNil(t, firstOptions.PreserveThinkingSuffix)
}
+5 -22
View File
@@ -1,31 +1,14 @@
package common
import (
"github.com/QuantumNous/new-api/dto"
"github.com/QuantumNous/new-api/types"
"github.com/QuantumNous/new-api/relaykit/relayconvert/convmeta"
"github.com/QuantumNous/new-api/relaykit/types"
)
// GuessRelayFormatFromRequest moved to convmeta with the converters; the
// delegation keeps host callers unchanged.
func GuessRelayFormatFromRequest(req any) (types.RelayFormat, bool) {
switch req.(type) {
case *dto.GeneralOpenAIRequest, dto.GeneralOpenAIRequest:
return types.RelayFormatOpenAI, true
case *dto.OpenAIResponsesRequest, dto.OpenAIResponsesRequest:
return types.RelayFormatOpenAIResponses, true
case *dto.ClaudeRequest, dto.ClaudeRequest:
return types.RelayFormatClaude, true
case *dto.GeminiChatRequest, dto.GeminiChatRequest:
return types.RelayFormatGemini, true
case *dto.EmbeddingRequest, dto.EmbeddingRequest:
return types.RelayFormatEmbedding, true
case *dto.RerankRequest, dto.RerankRequest:
return types.RelayFormatRerank, true
case *dto.ImageRequest, dto.ImageRequest:
return types.RelayFormatOpenAIImage, true
case *dto.AudioRequest, dto.AudioRequest:
return types.RelayFormatOpenAIAudio, true
default:
return "", false
}
return convmeta.GuessRelayFormatFromRequest(req)
}
func AppendRequestConversionFromRequest(info *RelayInfo, req any) {
+3 -3
View File
@@ -29,9 +29,9 @@ type StreamErrorEntry struct {
}
type StreamStatus struct {
EndReason StreamEndReason
EndError error
endOnce sync.Once
EndReason StreamEndReason
EndError error
endOnce sync.Once
mu sync.Mutex
Errors []StreamErrorEntry
+1 -1
View File
@@ -7,7 +7,7 @@ import (
"strings"
"github.com/QuantumNous/new-api/common"
"github.com/QuantumNous/new-api/dto"
"github.com/QuantumNous/new-api/relaykit/dto"
"github.com/QuantumNous/new-api/setting/operation_setting"
)
+1 -1
View File
@@ -4,7 +4,7 @@ import (
"strings"
"testing"
"github.com/QuantumNous/new-api/dto"
"github.com/QuantumNous/new-api/relaykit/dto"
"github.com/QuantumNous/new-api/setting/operation_setting"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"