mirror of
https://github.com/QuantumNous/new-api.git
synced 2026-09-13 15:54:34 +00:00
refactor: extract protocol conversion layer into standalone relaykit module (#6369)
* test(relayconvert): add golden snapshot matrix and relaykit boundary guard Phase 0 of the relaykit extraction plan: pin byte-level output of every registered (from,to) request/response/stream conversion route, and forbid kit-bound packages from growing host-only imports. * wip(relayconvert): drop gin.Context from converter signatures; add convmeta draft Phase 1 in progress: relayconvert now takes context.Context; host media resolver adapts gin.Context back at the service boundary. * refactor(relayconvert): decouple converters from RelayInfo, gin, and settings Phase 1 of the relaykit extraction plan: - converters now depend on convmeta.Meta (implemented by RelayInfo) instead of *relaycommon.RelayInfo; ClaudeConvertInfo and the format guesser move to convmeta with aliases left behind - host settings reach converters via a convmeta.Options snapshot built in RelayInfo.ConvOptions; no more model_setting/reasoning global reads inside the conversion layer - effort-suffix helpers move to service/relayconvert/reasoning (old package forwards); chat-to-responses upgrade policy moves to service (host routing logic, not conversion) - golden conversion matrix unchanged * test(relayconvert): tighten boundary — kit packages now free of gin/setting imports * refactor(dto): drop gin and logger dependencies Phase 2 (part 1): dto.Request.IsStream now takes *http.Request instead of *gin.Context (Gemini's impl reads query/path off the std request); dto's three logger calls become common.SysError. Boundary test allowlist is now empty — kit-bound packages import no gin/setting/logger/model. * refactor(kit): extract dependency-free kitutil; dto/types/relayconvert stop importing common Phase 2 of the relaykit extraction plan: - new service/relayconvert/kitutil holds the pure helpers the kit needs (JSON wrappers, pointer/string/uuid/timestamp utils, MaskSensitiveInfo, pluggable LogInfo/LogError hooks, Debug flag) - dto, types, and all relayconvert packages now use kitutil; their only remaining internal deps are dto/types/constant - common keeps every original symbol (MaskSensitiveInfo delegates to kitutil) so host code is untouched; main.go routes kit logging into common.SysLog/SysError and mirrors DebugEnabled - golden conversion matrix unchanged * refactor(kit): move EndpointType/FinishReason to types; OpenRouter dialect via Options Kit packages (dto/types/relayconvert/reasonmap) no longer import constant: - EndpointType and finish-reason values live in types; constant re-exports - the OpenRouter special-case in claude->openai request conversion reads Options.OpenRouterDialect, set by the host from the channel type; InitChannelMeta invalidates the cached snapshot on channel switch * refactor: extract relaykit submodule (dto/types/relayconvert/reasonmap) Phase 3 of the relaykit extraction plan: - new go module github.com/QuantumNous/new-api/relaykit containing dto (minus task family), types, relayconvert (with convmeta/kitutil/reasoning), and reasonmap; host consumes it via require + replace, go.work for dev - task-family dto (task/suno/midjourney/video) stays in the host dto package; dual-consumer host files alias it as taskdto - relaykit builds and tests standalone (GOWORK=off): no host imports, no gin, no DB, no settings - golden conversion matrix unchanged * build(docker): copy relaykit/go.mod before go mod download The local-replace submodule's go.mod must exist inside the build context for the main module graph to resolve. * fix: address relaykit extraction regressions * fix: address relaykit review regressions * docs: document Meta nil receiver contract * fix(relaykit): fail OpenAI→Claude conversion without max_tokens; reject negative default_max_tokens The Claude Messages API requires max_tokens (omitting it is a 400 "Field required"), but with a nil Options.Claude.DefaultMaxTokens hook the converters silently emitted a request the upstream is guaranteed to reject. Both OpenAI Chat and Responses → Claude conversions now return sharedclaude.ErrMissingMaxTokens when no path (client value, default hook, thinking-adapter floor) supplied one. Unreachable in the host, which always configures the hook. Host side, claude.default_max_tokens now rejects negative values at the option API before persisting — they would wrap into huge unsigned values during conversion. Zero stays allowed: the current API treats max_tokens: 0 as cache pre-warming. * fix: make Gemini safety settings read path race-free
This commit is contained in:
@@ -0,0 +1,31 @@
|
||||
package convmeta
|
||||
|
||||
import (
|
||||
"github.com/QuantumNous/new-api/relaykit/dto"
|
||||
"github.com/QuantumNous/new-api/relaykit/types"
|
||||
)
|
||||
|
||||
// GuessRelayFormatFromRequest infers the relay format from a request DTO's
|
||||
// concrete type. Moved from relay/common (which keeps a delegating alias).
|
||||
func GuessRelayFormatFromRequest(req any) (types.RelayFormat, bool) {
|
||||
switch req.(type) {
|
||||
case *dto.GeneralOpenAIRequest, dto.GeneralOpenAIRequest:
|
||||
return types.RelayFormatOpenAI, true
|
||||
case *dto.OpenAIResponsesRequest, dto.OpenAIResponsesRequest:
|
||||
return types.RelayFormatOpenAIResponses, true
|
||||
case *dto.ClaudeRequest, dto.ClaudeRequest:
|
||||
return types.RelayFormatClaude, true
|
||||
case *dto.GeminiChatRequest, dto.GeminiChatRequest:
|
||||
return types.RelayFormatGemini, true
|
||||
case *dto.EmbeddingRequest, dto.EmbeddingRequest:
|
||||
return types.RelayFormatEmbedding, true
|
||||
case *dto.RerankRequest, dto.RerankRequest:
|
||||
return types.RelayFormatRerank, true
|
||||
case *dto.ImageRequest, dto.ImageRequest:
|
||||
return types.RelayFormatOpenAIImage, true
|
||||
case *dto.AudioRequest, dto.AudioRequest:
|
||||
return types.RelayFormatOpenAIAudio, true
|
||||
default:
|
||||
return "", false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,215 @@
|
||||
// Package convmeta defines the conversion-context contract between format
|
||||
// converters (future relaykit) and the hosting application. Converters read
|
||||
// protocol state and per-request options exclusively through the Meta
|
||||
// interface; the host's RelayInfo implements it.
|
||||
package convmeta
|
||||
|
||||
import (
|
||||
"github.com/QuantumNous/new-api/relaykit/dto"
|
||||
"github.com/QuantumNous/new-api/relaykit/types"
|
||||
)
|
||||
|
||||
// Meta is the only view of the relay session that format converters may use.
|
||||
// It is satisfied by *relaycommon.RelayInfo on the host side; other embedders
|
||||
// (tests, external relaykit users) can use *Values.
|
||||
// Implementations backed by pointer types must make every method safe on a nil
|
||||
// receiver: a typed-nil pointer stored in Meta is still a non-nil interface,
|
||||
// and relaykit deliberately does not use reflection to detect that case.
|
||||
type Meta interface {
|
||||
GetOriginModelName() string
|
||||
GetUpstreamModelName() string
|
||||
// HasChannelMeta reports whether upstream channel information is attached;
|
||||
// converters use it to decide if GetUpstreamModelName is meaningful.
|
||||
HasChannelMeta() bool
|
||||
GetChannelID() int
|
||||
GetChannelType() int
|
||||
GetIsStream() bool
|
||||
GetReasoningEffort() string
|
||||
// SetReasoningEffort records the effort level a converter derived from a
|
||||
// model-name suffix so downstream billing/logging can see it.
|
||||
SetReasoningEffort(effort string)
|
||||
GetEstimatePromptTokens() int
|
||||
|
||||
// EnsureClaudeConvertInfo lazily creates and returns the mutable
|
||||
// OpenAI→Claude stream conversion state. For non-nil receivers, the same
|
||||
// instance must be returned for the lifetime of one streaming session; a
|
||||
// nil receiver may return a temporary initialized state.
|
||||
EnsureClaudeConvertInfo() *ClaudeConvertInfo
|
||||
|
||||
// GetSendResponseCount / IncrSendResponseCount expose the shared
|
||||
// downstream-chunk counter (the host may also increment it).
|
||||
GetSendResponseCount() int
|
||||
IncrSendResponseCount()
|
||||
|
||||
// AppendRequestConversion records a hop in the request format chain.
|
||||
AppendRequestConversion(format types.RelayFormat)
|
||||
|
||||
// ConvOptions returns the request-scoped conversion options snapshot.
|
||||
// Must never return nil.
|
||||
ConvOptions() *Options
|
||||
}
|
||||
|
||||
// ClaudeConvertInfo carries mutable state for OpenAI chat → Claude Messages
|
||||
// stream conversion. Moved here from relay/common (which keeps an alias).
|
||||
type ClaudeConvertInfo struct {
|
||||
LastMessagesType string
|
||||
Index int
|
||||
Usage *dto.Usage
|
||||
FinishReason string
|
||||
Done bool
|
||||
|
||||
ToolCallBaseIndex int
|
||||
ToolCallMaxIndexOffset int
|
||||
}
|
||||
|
||||
const (
|
||||
LastMessageTypeNone = "none"
|
||||
LastMessageTypeText = "text"
|
||||
LastMessageTypeTools = "tools"
|
||||
LastMessageTypeThinking = "thinking"
|
||||
)
|
||||
|
||||
// Values is a plain-struct Meta implementation for tests and non-RelayInfo
|
||||
// hosts (the relaykit-native entry point).
|
||||
type Values struct {
|
||||
OriginModelName string
|
||||
UpstreamModelName string
|
||||
ChannelMetaAttached bool
|
||||
ChannelID int
|
||||
ChannelType int
|
||||
IsStream bool
|
||||
ReasoningEffort string
|
||||
EstimatePromptTokens int
|
||||
|
||||
ClaudeConvertInfo *ClaudeConvertInfo
|
||||
SendResponseCount int
|
||||
ConversionChain []types.RelayFormat
|
||||
|
||||
Options *Options
|
||||
}
|
||||
|
||||
var _ Meta = (*Values)(nil)
|
||||
|
||||
func (v *Values) GetOriginModelName() string {
|
||||
if v == nil {
|
||||
return ""
|
||||
}
|
||||
return v.OriginModelName
|
||||
}
|
||||
|
||||
func (v *Values) GetUpstreamModelName() string {
|
||||
if v == nil {
|
||||
return ""
|
||||
}
|
||||
return v.UpstreamModelName
|
||||
}
|
||||
|
||||
func (v *Values) HasChannelMeta() bool {
|
||||
return v != nil && v.ChannelMetaAttached
|
||||
}
|
||||
|
||||
func (v *Values) GetChannelID() int {
|
||||
if v == nil {
|
||||
return 0
|
||||
}
|
||||
return v.ChannelID
|
||||
}
|
||||
|
||||
func (v *Values) GetChannelType() int {
|
||||
if v == nil {
|
||||
return 0
|
||||
}
|
||||
return v.ChannelType
|
||||
}
|
||||
|
||||
func (v *Values) GetIsStream() bool {
|
||||
return v != nil && v.IsStream
|
||||
}
|
||||
|
||||
func (v *Values) GetReasoningEffort() string {
|
||||
if v == nil {
|
||||
return ""
|
||||
}
|
||||
return v.ReasoningEffort
|
||||
}
|
||||
|
||||
func (v *Values) SetReasoningEffort(effort string) {
|
||||
if v != nil {
|
||||
v.ReasoningEffort = effort
|
||||
}
|
||||
}
|
||||
|
||||
func (v *Values) GetEstimatePromptTokens() int {
|
||||
if v == nil {
|
||||
return 0
|
||||
}
|
||||
return v.EstimatePromptTokens
|
||||
}
|
||||
|
||||
func (v *Values) EnsureClaudeConvertInfo() *ClaudeConvertInfo {
|
||||
if v == nil {
|
||||
return &ClaudeConvertInfo{LastMessagesType: LastMessageTypeNone}
|
||||
}
|
||||
if v.ClaudeConvertInfo == nil {
|
||||
v.ClaudeConvertInfo = &ClaudeConvertInfo{LastMessagesType: LastMessageTypeNone}
|
||||
}
|
||||
return v.ClaudeConvertInfo
|
||||
}
|
||||
|
||||
func (v *Values) GetSendResponseCount() int {
|
||||
if v == nil {
|
||||
return 0
|
||||
}
|
||||
return v.SendResponseCount
|
||||
}
|
||||
|
||||
func (v *Values) IncrSendResponseCount() {
|
||||
if v != nil {
|
||||
v.SendResponseCount++
|
||||
}
|
||||
}
|
||||
|
||||
func (v *Values) AppendRequestConversion(format types.RelayFormat) {
|
||||
if v == nil || format == "" {
|
||||
return
|
||||
}
|
||||
if n := len(v.ConversionChain); n > 0 && v.ConversionChain[n-1] == format {
|
||||
return
|
||||
}
|
||||
v.ConversionChain = append(v.ConversionChain, format)
|
||||
}
|
||||
|
||||
func (v *Values) ConvOptions() *Options {
|
||||
if v == nil {
|
||||
return &Options{}
|
||||
}
|
||||
if v.Options == nil {
|
||||
v.Options = &Options{}
|
||||
}
|
||||
return v.Options
|
||||
}
|
||||
|
||||
// UpstreamModelName / ChannelTypeOf are nil-safe accessors for optional Meta
|
||||
// values (converters are often called with a nil Meta in tests and compat
|
||||
// shims).
|
||||
func UpstreamModelName(m Meta) string {
|
||||
if m == nil || !m.HasChannelMeta() {
|
||||
return ""
|
||||
}
|
||||
return m.GetUpstreamModelName()
|
||||
}
|
||||
|
||||
func ChannelTypeOf(m Meta) int {
|
||||
if m == nil || !m.HasChannelMeta() {
|
||||
return 0
|
||||
}
|
||||
return m.GetChannelType()
|
||||
}
|
||||
|
||||
// OptionsOf returns m's conversion options, or empty defaults when m is nil.
|
||||
func OptionsOf(m Meta) *Options {
|
||||
if m == nil {
|
||||
return &Options{}
|
||||
}
|
||||
return m.ConvOptions()
|
||||
}
|
||||
@@ -0,0 +1,38 @@
|
||||
package convmeta
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"github.com/QuantumNous/new-api/relaykit/types"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
)
|
||||
|
||||
func TestValuesTypedNilMetaIsSafe(t *testing.T) {
|
||||
var values *Values
|
||||
var meta Meta = values
|
||||
|
||||
assert.Empty(t, meta.GetOriginModelName())
|
||||
assert.Empty(t, meta.GetUpstreamModelName())
|
||||
assert.False(t, meta.HasChannelMeta())
|
||||
assert.Zero(t, meta.GetChannelID())
|
||||
assert.Zero(t, meta.GetChannelType())
|
||||
assert.False(t, meta.GetIsStream())
|
||||
assert.Empty(t, meta.GetReasoningEffort())
|
||||
assert.Zero(t, meta.GetEstimatePromptTokens())
|
||||
assert.Zero(t, meta.GetSendResponseCount())
|
||||
|
||||
require.NotPanics(t, func() {
|
||||
meta.SetReasoningEffort("high")
|
||||
meta.IncrSendResponseCount()
|
||||
meta.AppendRequestConversion(types.RelayFormatClaude)
|
||||
})
|
||||
|
||||
convertInfo := meta.EnsureClaudeConvertInfo()
|
||||
require.NotNil(t, convertInfo)
|
||||
assert.Equal(t, LastMessageTypeNone, convertInfo.LastMessagesType)
|
||||
require.NotNil(t, meta.ConvOptions())
|
||||
require.NotNil(t, OptionsOf(meta))
|
||||
assert.Empty(t, UpstreamModelName(meta))
|
||||
assert.Zero(t, ChannelTypeOf(meta))
|
||||
}
|
||||
@@ -0,0 +1,79 @@
|
||||
package convmeta
|
||||
|
||||
// Options is the per-request snapshot of host configuration that converters
|
||||
// consult. The host fills it from its settings system when constructing the
|
||||
// Meta (see relaycommon.RelayInfo.ConvOptions); relaykit users fill it
|
||||
// directly. Zero value = every adaptation disabled, no defaults applied.
|
||||
type Options struct {
|
||||
Claude ClaudeOptions
|
||||
Gemini GeminiOptions
|
||||
|
||||
// OpenRouterDialect marks the upstream as OpenRouter's OpenAI-compatible
|
||||
// surface, which accepts extra fields (reasoning config, cache_control on
|
||||
// system parts) that converters emit only for that dialect. The host sets
|
||||
// it from the channel type.
|
||||
OpenRouterDialect bool
|
||||
|
||||
// PreserveThinkingSuffix reports models whose -thinking/-nothinking/effort
|
||||
// suffix must be kept on the outgoing model name (host blacklist lookup).
|
||||
// Nil means "never preserve".
|
||||
PreserveThinkingSuffix func(modelName string) bool
|
||||
}
|
||||
|
||||
type ClaudeOptions struct {
|
||||
// ThinkingAdapterEnabled turns "-thinking"-suffixed OpenAI model names
|
||||
// into Claude extended-thinking requests.
|
||||
ThinkingAdapterEnabled bool
|
||||
// ThinkingAdapterBudgetTokensPercentage sizes thinking budget_tokens as a
|
||||
// fraction of max_tokens when the adapter fires.
|
||||
ThinkingAdapterBudgetTokensPercentage float64
|
||||
// DefaultMaxTokens returns the max_tokens to inject when the source
|
||||
// request carries none. The Claude Messages API requires max_tokens
|
||||
// (omitting it is a 400), so when this hook is nil and no other path
|
||||
// supplies a value, OpenAI→Claude request conversion fails with an
|
||||
// explicit error instead of emitting a request the upstream is
|
||||
// guaranteed to reject. The new-api host always provides this hook;
|
||||
// standalone relaykit users must supply one or guarantee max_tokens on
|
||||
// every request.
|
||||
DefaultMaxTokens func(modelName string) int
|
||||
}
|
||||
|
||||
type GeminiOptions struct {
|
||||
// ThinkingAdapterEnabled maps -thinking/-nothinking/effort suffixes to
|
||||
// Gemini thinkingConfig.
|
||||
ThinkingAdapterEnabled bool
|
||||
// ThinkingAdapterBudgetTokensPercentage sizes thinkingBudget as a fraction
|
||||
// of maxOutputTokens when the adapter fires.
|
||||
ThinkingAdapterBudgetTokensPercentage float64
|
||||
// FunctionCallThoughtSignatureEnabled attaches thoughtSignature bypass
|
||||
// values to function-call parts.
|
||||
FunctionCallThoughtSignatureEnabled bool
|
||||
// SupportsImagine reports whether the model supports image generation
|
||||
// (switches response modalities). Nil means "never".
|
||||
SupportsImagine func(modelName string) bool
|
||||
// SafetySetting returns the harm threshold for a category. Nil or empty
|
||||
// return means no safetySettings are attached.
|
||||
SafetySetting func(category string) string
|
||||
}
|
||||
|
||||
func (o *ClaudeOptions) DefaultMaxTokensFor(modelName string) (int, bool) {
|
||||
if o == nil || o.DefaultMaxTokens == nil {
|
||||
return 0, false
|
||||
}
|
||||
return o.DefaultMaxTokens(modelName), true
|
||||
}
|
||||
|
||||
func (o *GeminiOptions) SupportsImagineModel(modelName string) bool {
|
||||
return o != nil && o.SupportsImagine != nil && o.SupportsImagine(modelName)
|
||||
}
|
||||
|
||||
func (o *GeminiOptions) SafetySettingFor(category string) string {
|
||||
if o == nil || o.SafetySetting == nil {
|
||||
return ""
|
||||
}
|
||||
return o.SafetySetting(category)
|
||||
}
|
||||
|
||||
func (o *Options) ShouldPreserveThinkingSuffix(modelName string) bool {
|
||||
return o != nil && o.PreserveThinkingSuffix != nil && o.PreserveThinkingSuffix(modelName)
|
||||
}
|
||||
Reference in New Issue
Block a user