Skip to content

About

A lite AI Gateway for any LLM API, any Models and any Agents.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Model Proxy v3

One proxy, five API schemas. Talk to Claude, Gemini, OpenAI Responses, and OpenAI Chat Completions models through a single endpoint, no matter which API format your client speaks.

The proxy accepts requests in the Claude Messages, Gemini (GenerateContent / Interactions), OpenAI Responses and Chat Completions format, converts them for whatever upstream provider you've configured, and converts the response back. On top of translation it handles exact / wildcard / catch-all model routing, composite aliases (weighted, primary+fallback, fusion fan-out, planner→executor coordinator), schedule-based timetable routing, per-model request/response transform hooks, global and per-alias token limits, privacy / compression sidecars, and a web + terminal dashboard for per-model, per-tool, and per-agent usage stats and some configs modification.

 Claude / Gemini genContent & interactions / OpenAI Responses & Chat Completions
                               β”‚
                               β–Ό            privacy-filter
                        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     compression
     sidecar plugins <- β”‚ Model Proxy β”‚ ->  image-fetch & encoding
                        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     tool judge
                                            auth & usage stats
                               β”‚ 
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β–Ό              β–Ό               β–Ό              β–Ό
   Anthropic        Gemini         OpenAI Chat      OpenAI
   Messages       GenContent &     Completions      Responses
                  Interactions     upstream         upstream
                  upstream

Proxy ↔ remote auth & stats service

The two optional remote sidecars ([remote] auth_server and [remote] record_server) can be the same service or two separate ones. auth_server gates admission; record_server collects per-request usage after the response. When they are the same service, the proxy can authenticate and report stats against one backend.

sequenceDiagram
    participant C as Client
    participant P as Model Proxy
    participant A as Auth Service<br/>([remote] auth_server)
    participant S as Stats Service<br/>([remote] record_server)
    participant U as Upstream Provider

    C->>P: POST /v1/messages<br/>(Authorization / x-api-key)
    Note over P: auth_with_model/auth_with_body = false β†’ auth now (GET)<br/>either = true β†’ defer until body parsed
    P->>A: auth_server<br/>GET (default) or POST (auth_with_body: whole request body)<br/>forward: Authorization, x-api-key, x-goog-api-key,<br/>user-agent, request_id, endpoint,<br/>[x-resource-for], [x-forwarded-for, x-real-ip]
    A-->>P: 200 OK<br/>header: one_time_auth_code / OTAC (optional)<br/>body: {version (required), targets[] (optional)}
    Note over P: if body carries targets[]<br/>β†’ walk rungs in order, fail over on retryable error,<br/>skip config-file model resolution
    P->>U: forwarded request (native or converted)
    U-->>P: response (streaming or JSON)
    P-->>C: response (converted back to client schema)
    P-)S: POST record_server<br/>{request_id, endpoint, user_key, model, response_status, token counters,<br/>[response_body if record_response_body=true]}<br/>header: one_time_auth_code, x-forwarded-for, [x-real-ip]
Loading

Auth targets[] failover ladder (response body). The auth service's 200 body MUST carry version (a non-empty string β€” "v1" is the current era). A 200 whose body is missing version, or is not an object, is rejected with 401 before any routing β€” a service that does not speak the versioned contract is never silently trusted. It MAY additionally carry targets β€” an ordered list of self-contained target descriptors. When present, the proxy uses them directly for this single request, starting on targets[0] and advancing to the next rung on a retryable upstream failure, and skips resolving the model from [models.*] / [composite] / [schedule] in the config file:

{ "version": "v1",
  "targets": [
    { "target": "claude-opus-4-6", "mode": "anthropic-messages",
      "base": "https://api.anthropic.com", "key": "sk-…", "timeout": 30000 },
    { "target": "gpt-5", "mode": "openai-completions",
      "base": "https://api.openai.com/v1", "timeout": 10000, "retry_on": [503], "retry": 2 }
] }
Target field Type Meaning
target string required Real upstream model id to send (like an alias target).
base / base_url string required Upstream base URL.
mode / upstream_mode string (optional) Upstream protocol: anthropic-messages, openai-completions, openai-responses, gemini-generatecontent, gemini-interactions. Defaults to [default_upstream].upstream_mode, else openai-completions.
key / api_key string (optional) Upstream API key for this rung only. When present it replaces the caller's credential for that rung (see the notice below); when omitted, the caller's credential is forwarded (subject to auth_passthrough_with).
otac string (optional) Per-rung value replacing the one_time_auth_code header for the upstream call and the stats record. Overrides the auth response's OTAC header for this rung.
transforms string (optional) Comma-separated [transforms.*] set names to apply. When omitted, no transforms are attached.
timeout number (optional) Whole-request upstream abort deadline for this rung, in milliseconds. Overrides UPSTREAM_BODY_TIMEOUT_MS; on expiry the proxy aborts the attempt and fails over to the next entry.
retry_on number[] (optional) Upstream statuses that re-hit this same rung before the ladder advances (axis 2). Bounded by [remote] max_target_retries (default 1; 0 disables) unless this rung sets retry.
retry number (optional) Max same-rung retries for this rung only, overriding [remote] max_target_retries (axis 2). retry_on still gates which statuses trigger it; 0 disables same-rung retry for the rung.

Descriptors are self-contained β€” target and base are required. A descriptor is not merged onto [default_upstream] / section / entry; the only inherited field is mode. A missing target/base would fall back to http://localhost with no key, so entries that omit them are rejected (dropped with an error) rather than silently mis-routed.

Notice β€” a rung's key replaces the caller's credential. When a descriptor carries a non-empty key, the proxy sends that key upstream and does not forward the caller's. The rung's key overwrites the mode's auth header (Authorization for openai-completions, x-api-key for anthropic-messages, x-goog-api-key for the Gemini modes) rather than being added alongside it. This applies without auth_passthrough_with = "config_key" β€” that setting governs config-resolved routes, not ladder rungs, so an auth service can pin credentials per rung regardless of the client's passthrough setting. It is per-rung, not per-ladder: an entry with no key still forwards the caller's credential, so one ladder may mix server-pinned and caller-supplied keys. A rung's otac likewise overrides the one_time_auth_code header for that rung only.

Failover (axis 1). The ladder advances on HTTP 429, any 5xx, a transport failure (β†’ 502), or an abort/timeout (β†’ 504). A deterministic 4xx (400/401/422/…) is terminal β€” the ladder stops and the client sees that rung's status. Total attempts are bounded by [remote] max_targets (default 16).

Bounds and validation. Entries are validated, deduplicated (target@base@key), then capped at max_targets; an invalid entry is dropped with an error and the ladder continues β€” even when it is targets[0]. If the version-carrying body has no targets (or every entry is invalid), the proxy falls back to normal config resolution; a missing version, a non-200 auth call, or an unparseable body is a contract failure (see the response table above). A rung's base host is not checked against the config host allowlist β€” the auth server is a trusted routing authority, so a descriptor may target any well-formed host (only base URL syntax is validated).

The ladder is per-request and ephemeral β€” never cached, never written to config, and does not persist across requests. It requires a parsed JSON request body (the proxy re-serializes it for each rung), so it applies to body-carrying endpoints. The auth call may run early or deferred β€” auth_with_model / auth_with_body are not required; the targets[] override is captured on either path.

See Auth & Stats Service Protocol below for the full wire-level contract.

Features

  • Five API formats in, any provider out β€” accept requests in any of these schemas:
    • /v1/messages β€” Claude Messages API
    • /v1beta/models/{model}:generateContent (+ :streamGenerateContent, :countTokens) β€” Gemini GenerateContent API
    • /v1/interactions β€” Gemini Interactions API
    • /v1/responses β€” OpenAI Responses API
    • /v1/chat/completions β€” OpenAI Chat Completions API (always enabled; per-model routed passthrough)
  • Embeddings β€” /v1/embeddings proxied to an OpenAI-compatible upstream.
  • Model-based routing β€” route each model name to its own upstream URL, API key, and protocol via a simple TOML config. Exact model keys are supported in every [models.*] category; provider wildcards (claude-*) and the final catch-all (*) are scoped as described in Model Routing & Aliases.
  • Composite aliases β€” group several models under one name with weighted-random, primary/fallback (with runtime share decay on failure), fusion fan-out, or a plannerβ†’executor coordinator that hands off once a trigger tool appears.
  • Schedule aliases β€” timetable-based routing: pick which model (or composite) serves a request based on server-local hour-of-day and day-of-week, with a fallback target for any time outside the configured windows.
  • Transform hooks β€” per-model / per-upstream request and response rewriting via five lifecycle hooks (request_ingress, before_conversion, before_upstream, after_upstream, response_egress), with Tier-1 field ops (rename / set / default / remove / map_value) and Tier-2 builtins for cross-message fixes. See Configuration Reference.
  • Extended thinking / reasoning β€” Claude-style thinking blocks, with conversion to OpenAI reasoning_effort for upstreams that need it. Handles tag-based (<think>...</think>) and reasoning_content extraction, including streaming and cross-chunk tool-call fragment stitching.
  • Usage accounting β€” per-model token and request stats, plus per-tool and per-agent stats (request/response counts, payload size, block counts), viewable in a web dashboard or a live terminal UI.
  • Token limits β€” global and per-alias token caps over a configurable window (sliding Nh/Nd or calendar 1w/1m). Returns HTTP 413 when exceeded.
  • Sidecars β€” optional privacy-filter (sidecar or local hash-only mode), compression, image-encode, and tool-judge sidecars: redact or shrink request payloads, fetch images, or prune irrelevant tools before the request reaches the upstream.
  • Runs anywhere β€” Node.js server or Docker.

Quick Start

1. Install

git clone <repo-url>
cd model_proxy_v3
npm install

Do not use npm install --dry-run --omit=dev (or --production) to preview a production install. Despite --dry-run, npm actually prunes devDependencies from the local node_modules β€” typescript, tsx, and every agent-SDK dev dep disappear, and the next npm run build fails with sh: tsc: command not found. Fix it by running npm install again. Such runs can also strip the optionalDependencies block from package.json β€” then the @github/keytar link is never restored by reinstalling alone; see "OS keychain key storage" under Configuration Reference for the recovery steps. To verify what a consumer would install, use npm pack --dry-run instead.

Node β‰₯ 19 recommended. The proxy uses the Web Crypto global (crypto.randomUUID()) available natively in Node β‰₯ 19. On Node 18.17–18.x, either:

  • export NODE_OPTIONS=--experimental-global-webcrypto before running any command (npm run server, npm run test:unit, etc.); or
  • in each source file that calls crypto.randomUUID(), add at the top:
    import { webcrypto } from 'node:crypto';
    const crypto = webcrypto;

Node < 18.17 is not supported.

2. Configure

Copy the example config and edit it:

cp docs/getting-started/proxy_config.example.toml proxy_config.toml

A minimal-config walkthrough (model categories, upstream_mode, per-model overrides, wildcards) and the behavioral notes on extended thinking (inline # comment support, third-party anthropic-messages endpoints, synthetic thinking signatures, budget_tokens vs max_tokens, the kimi-k2.7-code thinking_budget collision, tag-based and reasoning_content extraction) live in docs/getting-started/configuration-guide.md.

See proxy_config.example.toml for a fully commented config covering every section and option.

3. Run

npm run server        # starts on http://localhost:8788

Send a request:

curl http://localhost:8788/v1/messages \
  -H "Content-Type: application/json" \
  -H "x-api-key: your-key" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

The proxy reads proxy_config.toml from the working directory by default; when that is absent it falls back to ~/.config/model-proxy-v3/proxy_config.toml. Point it elsewhere with PROXY_CONFIG_PATH=./other.toml npm run server. Change the port with PORT=7777.

4. Watch live stats (optional)

Start with the terminal dashboard:

npm run server -- --tui
# equivalent:
TUI=true npm run server

--dashboard (or DASHBOARD=true) turns on token-stats persistence on its own β€” the JSONL dump and the startup restore below β€” without starting a UI; --tui and --agent imply it. It composes with every other mode (--rpc --dashboard is fine).

You get a live view of configured models, token usage, response times, and tool stats; a web dashboard is also available at GET /dashboard. The Q key (documented in the h help panel) opens a model picker with a per-row usage suffix β€” (58%), (7472/20000, 63%) (remaining/limit + used%), (Β₯43.97) β€” and shows the highlighted model's full quota on a status line at the bottom of the panel when moving through the list (minimax, deepseek, kimi, openrouter, zhipu, qnaigc coding plans β€” see GET /dashboard/api/quota); anthropic-routed models instead show the 5h-window left percent recorded from the anthropic-ratelimit-unified-5h-utilization response header on proxied traffic. In the composite-aliases panel each target model shows its usage left after the timing suffix, e.g. [0.11/2.12/63.93s] (58% left) or (6930/12000, 42% left) for count-based providers. The TUI key bindings, the model_proxy_tokens.jsonl usage-dump format, and the startup stats-restoration rules are documented in docs/guides/live-stats.md.

Mode-flag gating: what gets tracked

The proxy only records detailed stats and enforces token limits when at least one of --dashboard, --tui, --agent, or --rpc is enabled. In "plain proxy" mode (no flags), memory usage is minimal:

Store Flags set No flags
modelStats (per-model totals) βœ… βœ…
dailyTokenStats (daily rollup) βœ… βœ…
tokenHeatmapEvents (sliding-window events) βœ… ❌
compositeAliasStates.events (alias limit windows) βœ… ❌
agentStats / toolRequestChars / upstreamResponseToolStats βœ… ❌
requestEndpointStats / timing / upstream / status codes βœ… ❌

Token limits enforced: βœ… Global + Composite alias | ❌ Both disabled

This means without any UI/runtime flag the proxy runs with minimal overhead β€” only cumulative per-model and daily totals are kept, no heatmap, no tool tracking, and no token-limit enforcement.

Performance benchmarks (10,000 rpm β‰ˆ 167 req/s)

Mode Heap (stats) Concurrency (30 s streams) Total RSS tokenHeatmapEvents shift CPU
Plain (no flags) ~50 KB 5,000 Γ— ~50 KB = 250 MB ~300–400 MB 0% (events disabled)
Flags + 1-day retention 14.4M events = 922 MB 250 MB ~1.2 GB ~283% (2.8Γ— realtime)

CPU figures measured on Node 22.23.2: Array.prototype.shift() on object arrays is O(n) memmove (no left-trim optimization). Benchmark: 100k→443 ¡s/op, 1M→2.7 ms/op, 10M→12 ms/op. At 167 ops/s, 1-day retention (14.4M array) saturates ~3 CPU cores on shift() alone.

Takeaway: plain mode scales to 10k+ rpm with flat ~350 MB. Enabling any dashboard flag (--dashboard, --tui, --agent, --rpc) at 10k rpm requires per-second event aggregation or a ring buffer β€” the current per-request array + shift() design caps at ~1–2k rpm with 1-day retention.

CLI Commands

The server binary doubles as a config inspection tool. With no arguments it starts the HTTP server as before; with one of the commands below it runs, prints, and exits.

npm run server -- --list-models          # human-readable table
npm run server -- --list-models --json   # machine-readable (dashboard config shape)
npm run server -- --validate-config      # check the config file
npm run server -- --export-pi-models     # pi models file + default provider/model
npm run server -- --export-pi-models --default-model smart-coder   # pick the default model
npm run server -- --export-openclaw-providers   # OpenClaw models.providers block
npm run server -- --help                 # usage

(npx model-proxy-v3 … or node dist/server.js … once built.)

Command Purpose
--list-models [--json] List configured target models plus composite and schedule aliases. The default output is a grouped table β€” target models (CATEGORY/ID/TARGET/BASE URL/MODE, with category-level base_url/upstream_mode shown when an entry inherits them), composite aliases with their targets (one per line) and token limit, and schedule aliases with their timed windows (one per line). Headers are listed directly above the rows β€” no dash rule. --json emits the sanitized payload the dashboard config endpoint serves, with api_key values stripped.
--validate-config Parse and validate the config, printing the [ERROR]/[WARN] lines the parser finds plus a counts summary; exits 1 when there are errors.
--export-pi-models [--default-model <id>] Print what pi needs to route through this proxy: defaultProvider/defaultModel (for ~/.pi/agent/settings.json) plus a providers block (for ~/.pi/agent/models.json) β€” one provider (model-proxy-v3) holding a pi-ai Model object per configured target model and alias, each pointed at this proxy's own loopback origin (http://127.0.0.1:$PORT, default 8788). defaultModel is the --default-model alias, or the first configured model when the flag is omitted; an unknown id is a usage error. apiKey is a dummy (sk-hi): the proxy's client auth is a presence check, and configured target api_key values are never emitted.
--export-openclaw-providers Print the models.providers block for ~/.openclaw/openclaw.json β€” a single provider (model-proxy-v3) holding one entry per configured target model and alias, each pointed at this proxy's own loopback origin (http://127.0.0.1:$PORT, default 8788). The provider carries api: 'anthropic-messages', auth: 'api-key', and the dummy apiKey (sk-hi); the surrounding models.mode is merge, so the block can be merged into an existing OpenClaw config without dropping its other providers. Configured target api_key values are never emitted.
--help, -h Print usage.

Commands read the local TOML file only ($PROXY_CONFIG_PATH); Consul/Apollo remote sources are not consulted. When PROXY_CONFIG_PATH is unset, the config path defaults to ./proxy_config.toml if it exists, else ~/.config/model-proxy-v3/proxy_config.toml. When neither exists the home path is returned anyway and its directory (~/.config/model-proxy-v3) is created, so a GUI-launched binary whose working directory is not the repo still finds its config in the home directory.

Exit codes: 0 success (for --validate-config: no errors), 1 command failed (unreadable config, or a config with errors), 2 usage error (unknown argument, more than one command, --json without --list-models, or --default-model without --export-pi-models / without a model id / naming an unknown one).

API Endpoints

Endpoint Purpose
POST /v1/messages Claude Messages API
POST /v1/messages/count_tokens Count tokens (Claude/OpenAI format)
POST /v1/responses OpenAI Responses API
POST /v1/responses/input_tokens Count input tokens for a Responses request (forwarded to /v1/responses/input_tokens under openai-responses, or bridged through Chat Completions token counting under other modes)
POST /v1/responses/compact Compact a Responses conversation (forwarded to /v1/responses/compact under openai-responses, or bridged through Chat Completions under other modes)
POST /v1beta/models/{model}:generateContent Gemini content (also :streamGenerateContent, :countTokens)
POST /v1/interactions Gemini Interactions API
POST /v1/embeddings Embeddings (proxied to an OpenAI-compatible upstream)
GET /v1/models List available models (no auth required)
GET /dashboard Web dashboard for config + stats
GET /dashboard/api/quota?model=<id> Remaining usage/credits for a model's route (minimax, deepseek, kimi, openrouter, zhipu, qnaigc coding plans; provider detected from the route host). ?base_url=<origin> variant serves the web dashboard's per-URL "Usage Left" column, falling back to the recorded anthropic 5h percent. Dashboard api_key auth.
GET /config-reload Reload config from PROXY_CONFIG_CONSUL or PROXY_CONFIG_APOLLO. Only meaningful when a remote config source is set; returns 400/500 otherwise. Clears the config cache and re-fetches.
GET /health (also GET /) Health check. Probes the resolved default-category / [default_upstream] upstream /v1/models; returns {status:"ok", models, cached, version} on success or 404 when no models are reachable. No auth required.
GET /favicon.ico Returns 204 No Content (browser plumbing).
POST /decision The Clef API β€” typed decision questions (noul/choice/score) on arbitrary state. Proxies to a configured upstream (local Laya sidecar or Cloudflare Clef). See Decision endpoint.
/{protocol}/{host}/... dynamic route Per-request upstream override. See Dynamic routing.
POST /passthrough/v1/... Passthrough mode: the path after /passthrough is forwarded verbatim to a [passthrough] target whose mode matches it (see Passthrough mode). Client auth, auth_server, logging, usage recording, and timeouts still apply; model routing, composite/schedule, transforms, privacy filtering, and kompress do not.

A Gemini /v1/models/{model}:... variant exists for each /v1beta/models/{model}:... endpoint. :countTokens is supported too: native Gemini routes forward to Gemini countTokens, while non-Gemini upstream modes are bridged through the proxy's token-counting path. For streaming conversions, the proxy forwards final upstream usage where the upstream emits it: OpenAI Chat Completions uses stream_options.include_usage, Gemini streams use final usageMetadata / interaction usage, and cache-read tokens are preserved through Claude/Responses usage fields when available. Full request/response examples live in the API reference docs.

Supported upstream_mode

Each endpoint can be routed to one or more upstream API families (upstream_mode). The mode is selected by the route's defaultMode / model config:

Client endpoint anthropic-messages openai-completions openai-responses gemini-generatecontent gemini-interactions
POST /v1/messages Native passthrough to /v1/messages; request stays Claude Messages format end-to-end. Direct transform: Claude Messages β†’ Chat Completions β†’ Claude Messages. If input is already OpenAI-shaped, it can pass through. Indirect transform via openai-completions: Claude Messages β†’ Chat Completions β†’ Responses input β†’ Claude Messages. Basic tools and streaming are supported; max_tokens is rewritten to max_output_tokens. Direct transform: Claude Messages β†’ Gemini generateContent β†’ Claude Messages. Direct transform: Claude Messages β†’ Gemini Interactions/generateContent-compatible upstream β†’ Claude Messages.
POST /v1/responses Direct transform: Responses input/instructions β†’ Claude Messages β†’ Responses. Text and tool-use are supported for non-streaming and streaming. Direct transform: Responses β†’ Chat Completions β†’ Responses. For api.qnaigc.com, keeps legacy max_tokens; otherwise uses max_completion_tokens. Native passthrough to /v1/responses. Direct transform via Claude Messages: Responses β†’ Claude Messages β†’ Gemini generateContent β†’ Claude Messages β†’ Responses. Direct transform via Claude Messages: Responses β†’ Claude Messages β†’ Gemini Interactions/generateContent β†’ Claude Messages β†’ Responses.
POST /v1/chat/completions Convert passthrough: Chat Completions body is converted to Claude Messages format and forwarded to /v1/messages; response (streaming and non-streaming) is converted back to OpenAI completions format. Tool schema types are lowercased; content: "" on assistant messages with tool_calls is normalized to null; consecutive tool messages are grouped into one user turn. Native passthrough. Uses the resolved per-model route; composite aliases and target-mapped model ids are resolved and the model field in the forwarded body is rewritten to the target model id. Transform passthrough: Chat Completions body is converted to Responses input and forwarded to /v1/responses using the resolved per-model route. Transform passthrough: Chat Completions body (including image_url blocks) is converted to Gemini generateContent body (inline_data for data-URI images; http(s) image URLs are fetched server-side with an SSRF guard) and forwarded to :generateContent / :streamGenerateContent?alt=sse. Text deltas and finishReason round-trip; tool-call/thinking response parts and any model-generated image output are dropped (response schemas for Claude Messages and OpenAI Completions do not carry image output β€” see image I/O notes). Same as gemini-generatecontent; not separately wired today.
POST /v1beta/models/{model}:generateContent / :streamGenerateContent Indirect transform via openai-completions: generateContent β†’ Chat Completions β†’ Claude Messages β†’ generateContent. Forwards upstream to /v1/messages; text, tool calls, and streaming text deltas return as Gemini candidates[].content.parts; tool calls become functionCall parts. Direct transform: generateContent β†’ Chat Completions β†’ generateContent. Forwards upstream to /v1/chat/completions. Indirect transform via openai-completions: generateContent β†’ Chat Completions β†’ Responses input β†’ generateContent. Forwards upstream to /v1/responses; system/developer messages become Responses instructions; content-part arrays are normalized to text. Native passthrough to :generateContent / :streamGenerateContent using the configured Gemini API version. Native Gemini-family route; forwards to Gemini generateContent/stream endpoint using Interactions-compatible mode.
POST /v1/interactions Indirect transform via openai-completions: Interactions β†’ Chat Completions β†’ Claude Messages β†’ Interactions. Forwards upstream to /v1/messages; text, tool calls, and streaming text deltas return in Interactions shape. Direct transform: Interactions β†’ Chat Completions β†’ Interactions. Forwards upstream to /v1/chat/completions. Indirect transform via openai-completions: Interactions β†’ Chat Completions β†’ Responses input β†’ Interactions. Forwards upstream to /v1/responses; system/developer messages become Responses instructions; content-part arrays are normalized to text. Native Gemini-family route; forwards to Gemini generateContent/stream endpoint. Native Gemini-family route; forwards to Gemini generateContent/stream endpoint using Interactions-compatible mode.
GET /v1/models Passthrough model listing; no upstreamMode conversion is applied. Passthrough model listing; no upstreamMode conversion is applied. Passthrough model listing; no upstreamMode conversion is applied. Passthrough model listing; no upstreamMode conversion is applied. Passthrough model listing; no upstreamMode conversion is applied.
POST /v1/embeddings Not supported. Only supported mode; forwards to OpenAI-compatible embeddings upstream. Not supported. Not supported. Not supported.

Notes:

  • Native passthrough means the client endpoint and upstream API family already match, so the request body is not converted to another provider's format.
  • Direct transform means the proxy converts directly between the client endpoint format and the selected upstream family, then converts the response directly back to the client endpoint shape.
  • Direct transform via Claude Messages means Responses uses Claude Messages as its internal bridge before calling Gemini; it does not go through openai-completions.
  • Indirect transform via openai-completions means the request body is routed through OpenAI Chat Completions as an intermediate shape before reaching the target upstream family. This covers two cases: (a) Gemini endpoint input becomes Chat Completions, then becomes Claude Messages or OpenAI Responses; (b) /v1/messages routed to an openai-responses upstream becomes Chat Completions, then Responses input. This reuses the Chat Completions middle mode for code reuse while preserving the original client endpoint response shape.
  • Direct transforms are preferred long-term for endpoint fidelity. The current /v1/interactions β†’ anthropic-messages / openai-responses routes use the indirect openai-completions bridge for code reuse; see Routing transform review for tradeoffs and recommendations.

Token usage statistics columns

The TUI "Top Models" panel and the stats sidecar record token usage per request from the upstream response (JSON body or final SSE usage event), via extractUsageFromResponsePayload / createUsageTrackingTransformStream (src/utils/dashboard-stats.ts). Column semantics per endpoint:

Endpoint in cached wrote out
/v1/messages input_tokens (uncached input only) cache_read_input_tokens cache_creation_input_tokens output_tokens
/v1/chat/completions prompt_tokens (includes cached) prompt_cache_hit_tokens or prompt_tokens_details.cached_tokens prompt_cache_miss_tokens (DeepSeek-style) completion_tokens
/v1/responses input_tokens input_tokens_details.cached_tokens β€” (always 0) output_tokens
/v1beta/models/{model}:generateContent promptTokenCount (includes cached) cachedContentTokenCount β€” (always 0) candidatesTokenCount
/v1/interactions input_tokens ?? total_input_tokens only if upstream sends cache_read_input_tokens only if upstream sends cache_creation_input_tokens output_tokens ?? total_output_tokens

Notes:

  • total is the upstream total_tokens when present, else computed as in + cached + wrote + out.
  • For streaming chat/completions, the proxy forces stream_options.include_usage: true (native passthrough and converted routes alike) so the upstream emits the final usage chunk; the extra chunk is forwarded to the client unchanged.
  • Anthropic's input_tokens excludes cached tokens, but OpenAI Chat Completions prompt_tokens and Gemini promptTokenCount include them β€” so for those endpoints the computed total counts cached tokens in both in and cached.

Endpoint details

Additional endpoint behavior is documented in docs/api/api-endpoints.md:

  • Dynamic routing β€” per-request upstream override routes /{protocol}/{host}/..., disabled by default (opt in with ENABLE_DYNAMIC_ROUTING=true) and guarded by an SSRF allowlist (ALLOWED_HOSTS).
  • Image input/output across format boundaries β€” wire shapes, source-shape handling, who fetches HTTP image URLs, and the model-generated-image limits.
  • OpenAI prompt caching fields β€” which of prompt_cache_key / prompt_cache_options / prompt_cache_breakpoint survive each cross-mode conversion.
  • Dashboard API β€” the /dashboard/api/* JSON routes, optional bearer token, and stats keying by resolved model id.

Decision endpoint (POST /decision)

An endpoint for typed decisions rather than generation: the caller supplies some state and a map of questions, and gets back a probability per question. There is no prompt and no completion β€” questions are answered in a single forward pass.

The wire contract is the Clef API, defined by docs/api/decision/clef-schema-input.json and docs/api/decision/clef-schema-output.json. Those two files are authoritative; the summary below is a convenience.

// request β€” required: model, state, questions
{
  "model": "clef",                 // "clef" or "clef-flash"
  "state": "…",                    // string, or structured data (object/array)
  "questions": {                   // 1–64 questions, answers return under the same ids
    "keep": { "type": "noul",   "instructions": "Is this file safe to edit?" },
    "plan": { "type": "choice", "instructions": "Pick a plan", "criteria": ["free", "pro"] },
    "risk": { "type": "score",  "instructions": "Rate the risk", "legend": ["none", "low", "high"] }
  },
  "images": [ … ]                  // optional
}

// response β€” required: model, answers, usage
{
  "model": "clef",
  "answers": {
    "keep": { "type": "noul",   "noul": 0.87 },
    "plan": { "type": "choice", "choice": "pro", "probabilities": { "free": 0.2, "pro": 0.8 }, "confidence": 0.8 },
    "risk": { "type": "score",  "score": 1.4, "legend": { "0": "none", "1": "low", "2": "high" },
              "probabilities": { "0": 0.1, "1": 0.4, "2": 0.5 }, "confidence": 0.5 }
  },
  "usage": { "input_tokens": 412, "output_tokens": 0 }
}

Configuration β€” url is a full endpoint, not a base URL (the proxy POSTs to it verbatim, so include the path). Setting both backend and url is what enables the route; without them /decision answers 503.

[decision]
backend = "laya"                            # "laya" (text-only) or "cloudflare"/"clef" (images allowed)
url = "http://localhost:8765/decision"      # full endpoint URL, POSTed verbatim
# api_key = "optional-bearer-token"         # sent as Authorization: Bearer
# timeout_ms = 5000                         # default: 5000 for "laya", 30000 for "cloudflare"/"clef"

Two backends, differing in exactly one thing β€” images:

  • laya β€” the local sidecar submodules/laya-mlx/serve_judge.py, serving this contract from the Laya typed-decisions model. Its MLX encoder is text-only, so a request carrying a non-empty images array is rejected with 400 before any upstream call. It fails loud rather than dropping the images, so a caller cannot mistake a text-only answer for one that saw the image. The same token-budget caveat as the tool judge applies β€” see the Laya context-limit note under Documentation.
  • cloudflare β€” any upstream serving the Clef schemas with image support (e.g. Cloudflare's @cf/cloudflare/clef). The images array is forwarded untouched, and the model id in the body selects the remote model. clef is an accepted equivalent spelling for this backend.

Behaviour notes:

  • Verbatim forwarding β€” the request body is forwarded byte-for-byte, and the upstream response body is returned byte-for-byte. The proxy implements no transport of its own beyond the POST, and does no model routing or alias resolution: the model field in the body is the upstream's business.
  • Envelope-only validation β€” the proxy checks only the top-level required fields (model/state/questions in, model/answers/usage out). Per-question shapes, id charset, the 1–64 question bound, and the model pattern are the upstream's validation, and its error is forwarded.
  • Errors pass through β€” a non-2xx upstream response is returned as-is (status, body, and x-request-id). A 2xx whose body is not JSON, or is missing the response envelope, is reported as 502 Invalid upstream response.
  • Auth & stats apply β€” the caller's credential and [remote] auth_server gate run as for other endpoints; requests are recorded to record_server and appear in the dashboard under the model from the request body.

Model Routing & Aliases

Incoming model names resolve through three stacked logic levels (see the Routing Hierarchy table below):

  • Level 1 β€” [models.*] β€” exact key β†’ prefix-* wildcard β†’ * catch-all lookup, with base_url / api_key / upstream_mode inherited per-entry β†’ section β†’ [default_upstream]. An optional per-entry max_tokens (e.g. "m" = {target = "...", max_tokens = 8192}) fills the value when the request omits max_tokens (anthropic-messages) and caps larger client values at the before_upstream hook (all modes); when unset, the field is strict passthrough β€” the proxy never sets, modifies, or caps it. Some upstreams require max_tokens (e.g. DeepSeek's Anthropic-compatible API rejects requests without it) β€” configure max_tokens on those target entries, since the proxy no longer injects a default. See docs/reference/configuration-reference.md. [models.FREE] and [models.EMBEDDING] are exact-only; in [models.FREE] the configured key always wins, elsewhere the caller's key wins by default (auth_passthrough_with).
  • Level 2 β€” [composite] aliases:
    • share / primary+fallback β€” weighted random or ordered fallback, with runtime share decay when a target keeps failing, plus optional per-alias token_limit.
    • fusion β€” fan-out to parallel panel models with an optional judge and synth.
    • coordinator β€” planner β†’ executor hand-off: routes to the planner until a trigger tool call (ExitPlanMode, Edit, Write, …) appears in the conversation history.
  • Level 3 β€” [schedule] aliases β€” pick one target by server-local hour-of-day / day-of-week windows, with an empty-window fallback target.
  • Token limits β€” global (general.global_token_limit) and per-alias caps over sliding (1h–6d) or calendar (1w/1m) windows; HTTP 413 when exceeded.

Note β€” response model field reflects the real upstream model, not the alias. The proxy rewrites the request body's model field from the client's alias to the real target model id before forwarding, but the response body's model field (JSON and streaming chat.completion.chunk events alike) is passthrough by default β€” clients see whatever the real upstream model returns, not the alias they requested. To echo the requested alias back instead, attach the restore_client_model_alias transform built-in to the route β€” see docs/reference/transforms-reference.md.

The full reference β€” category lookup priority tables, base_url/api_key override and "who wins" rules, every composite/fusion/coordinator/schedule option, the token-limit windowing engine, and worked examples β€” lives in docs/reference/routing-and-aliases.md.

Routing Hierarchy (Logic Levels)

The proxy has three logic levels, stacked bottom-up. Each level chooses which level below gets to serve this request:

                        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
   Level 3 (top)        β”‚  [schedule]                   β”‚  ← timetable (hour-of-day, day-of-week)
                        β”‚  "what should serve *now*?"   β”‚
                        β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
   Level 2 (middle)     β”‚  [composite]                  β”‚  ← share / primary+fallback / fusion fan-out / coordinator
                        β”‚  "split or sequence across N?" β”‚
                        β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
   Level 1 (base)       β”‚  [models.*]                   β”‚  ← exact name / prefix-* wildcard / * catch-all
                        β”‚  "which upstream?"             β”‚
                        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
Level Section Selects by Cardinality Re-routes to
3 [schedule] Timetable windows 1 β†’ 1 (one target picked per request) Level 2 or 1
2 [composite] (share / primary+fallback) Weighted random or fallback order 1 β†’ 1 Level 1
2 [composite] (fusion) Role + fusion_options 1 β†’ N β†’ 1 (panelΓ—N + judge + synth) Level 1
2 [composite] (coordinator) Stage detection via toolset in messages history 1 β†’ 1 (planner β†’ executor, one-way) Level 1
1 [models.*] Exact / prefix-* / * catch-all 1 β†’ 1 (one upstream) β€” (sends)

Level-by-level details and worked request-resolution examples are in docs/reference/routing-and-aliases.md.

Deployment

Docker

cp docs/getting-started/proxy_config.example.toml proxy_config.toml
#COMMIT=$(git rev-parse --short HEAD)
#docker build --network=host --build-arg VERSION=$COMMIT -t model-proxy-v3:$COMMIT -t model-proxy-v3:latest .
docker build -t model-proxy-v3 .
docker run --network host -p 8788:8788 -v $(pwd)/proxy_config.toml:/app/proxy_config.toml -e LOG_LEVEL=info model-proxy-v3

For higher throughput, run several containers behind an nginx reverse proxy that load-balances across them. Refer to docs/reference/nginx_conf/ for nginx configuration examples.

npm

The published package ships only the compiled output β€” dist/ plus package.json, README.md, and LICENSE (whitelisted via the files field in package.json; .npmignore is a backup blocklist). No tests, docs, configs, or source are included.

# Preview the exact tarball contents (no publish, no side effects):
npm pack --dry-run

# Publish. `prepublishOnly` runs `npm run build` + `npm run test:unit` first,
# so dist/ is always fresh and green in the tarball:
npm publish

What the published package gives consumers:

Field Value Meaning
main dist/index.js ESM entry β€” the compiled src/index.ts fetch handler.
bin model-proxy-v3 β†’ dist/server.js Runnable via npx model-proxy-v3 or a global install.
files ["dist"] Only compiled JS is packed.
dependencies runtime-only typescript and all agent/test SDKs live in devDependencies and are never installed by consumers.
engines node >= 19 Web Crypto (crypto.randomUUID()) requirement.

Running the published package:

npx model-proxy-v3                      # or: npm i -g model-proxy-v3 && model-proxy-v3
PORT=8788 model-proxy-v3                # default port is 8788

The server reads proxy_config.toml from the current working directory; when that is absent it falls back to ~/.config/model-proxy-v3/proxy_config.toml. Point it elsewhere with PROXY_CONFIG_PATH=/path/to/proxy_config.toml npx model-proxy-v3.

Native single-file binary

npm run build:native     # -> dist/model-proxy-v3-<host-triple>

The output is named <name>-<host target triple> (e.g. model-proxy-v3-x86_64-apple-darwin), the same name the Tauri tray's externalBin stages the sidecar under, so it can be copied across unchanged. A ready-to-use system tray for this proxy lives in git@github.com:qidu/proxy_tray.git: a Tauri v2 app that stages the binary as its sidecar and supervises it over the --rpc JSON-RPC control channel (design: docs/architecture/design_tauri_tray.md).

scripts/build-sea.js bundles the server into one Node SEA executable. The output is a copy of the Node that built it, so that Node must be an official, self-contained build (nodejs.org, nvm, actions/setup-node). Homebrew's is a thin launcher over a shared libnode with SEA compiled out and cannot host one; the script detects that up front and refuses rather than failing mid-build.

If npm run build:native fails on that check, build with an official Node without touching the system install:

npx --yes --package=node@22 node scripts/build-sea.js

Node response compression headers

The Node server adapter normalizes response headers before writing them to the client. It removes content-encoding and content-length because Node fetch() can decode upstream-compressed bodies while leaving the original upstream headers visible on the Response. The previous direct header copy:

Object.fromEntries(response.headers.entries())

could therefore send plain text with a stale content-encoding: br / gzip header. Clients such as opencode may then try to decode Brotli by default and fail to read the response body.

Auth & Stats Service Protocol

The proxy talks to two optional remote services over plain HTTP: an auth service ([remote] auth_server) that gates admission before routing, and a stats service ([remote] record_server) that collects per-request usage records after the response. The exact wire-level contract β€” request/response shapes, forwarded headers, the one_time_auth_code (OTAC) linkage, auth_with_model / auth_with_body timing, the auth targets[] failover ladder, and how to combine both services in one backend β€” is documented in docs/architecture/auth-stats-protocol.md.

Configuration Reference

Most users only need proxy_config.toml; optional environment variables tune behavior. The full field-by-field reference lives in docs/reference/configuration-reference.md:

  • TOML sections β€” [general], [default_upstream], [remote], [transforms.*] / [transform_defaults], [privacy_filter], [decision], [dashboard], [passthrough].

  • OS keychain key storage β€” [general] store_key_in_system = true moves every config api_key into the OS keychain (accounts <target_model_id>/<base_url> under the model_proxy_v3 service) and rewrites the config file to STORE_KEY_IN_SYSTEM sentinels; sentinels are resolved from the keychain on every later load, silently, as long as the same node binary/path keeps reading them β€” a Node upgrade or a different binary path can trigger a one-time OS keychain prompt or a fatal KeyStoreError if keychain access is denied. A single sentinel with no matching keychain entry is not fatal: that slot is cleared (treated as unconfigured, so the normal api_key fallback applies) and reported as a config error in the TUI/dashboard/console, letting the rest of the models load and the proxy start. Scope: configured api_keys of [models.*] targets, default_upstream.default_api_key, and [passthrough] target keys in the local proxy_config.toml only β€” ignored for Consul/Apollo sources, N/A for composite aliases, and caller/user keys from request headers are never stored. Local/dev-host feature only β€” fails loud when no OS keychain is available (Docker, Cloudflare Workers). Backed by the @github/keytar native addon, pinned in optionalDependencies as github:github/node-keytar#v7.10.6: npm install fetches it from GitHub and its install script downloads a prebuilt NAPI binary β€” no compiler or Python needed (see the toolchain note below). A source copy is kept at submodules/node-keytar for local development; it is not used by npm install. If startup fails with Cannot find package '@github/keytar', first check git diff package.json: npm prune-style runs (--dry-run --omit=dev, --production β€” see the warning in Quick Start) can strip the optionalDependencies block, after which no npm install ever re-fetches the addon; restore it with git checkout -- package.json && npm install. When a sentinel's exact <target>/<base_url> account isn't found, resolution falls back to a best-effort search across every stored account for the service (keytar.findCredentials) β€” on macOS this broader enumeration can trigger its own Keychain Access permission prompt in a separate GUI window (not the terminal); if startup appears to hang here, check for that dialog and click "Always Allow". This fallback call is bounded to 60s β€” if it doesn't respond in time (e.g. a stuck/dismissed dialog), it fails loud with a clear timeout error instead of hanging the process indefinitely.

  • store_key_in_system in the win32 single-executable build β€” @github/keytar needs too many dependencies on Windows (Visual Studio Build Tools with the "Desktop development with C++" workload + Python 3, via node-gyp β€” see the toolchain table below), so on Windows only the executable substitutes an in-binary store and keeps the keys in its own file body (process.execPath, using the record format from store-in-body): setPassword appends a JSON record at EOF and lookups walk the record chain backwards. No OS keychain is involved, and the proxy_config.toml behaviour is unchanged (keys are still rewritten to sentinels and resolved on later loads). Three consequences: the keys are plaintext inside the .exe β€” that is portability, not secrecy; appending rewrites the executable, so its directory must be writable (an in-place append at EOF is tried first; otherwise the whole file is copied and swapped, and either way the directory must permit writes) and a full append is O(file size); and rebuilding wipes the stored keys, since a fresh blob is injected into a fresh copy of node. The fallback engages only for a real SEA binary whose process.execPath basename is not node/node.exe, so it can never target a Node install. macOS and Linux keep the OS keychain via @github/keytar; the Docker image and macOS/Linux SEA binaries keep the fatal KeyStoreError.

  • System & toolchain requirements for store_key_in_system = true β€” the feature needs an OS keychain backend at run time. On macOS/Linux the @github/keytar native addon installs as a prebuilt NAPI binary (ABI 3 β€” works on any modern Node), so no build toolchain is required; a native toolchain (Python 3 + a C++ compiler, via node-gyp) is needed only when the prebuilt download fails (e.g. no network access to GitHub releases) and npm falls back to compiling the addon from source. On Windows the addon needs that toolchain in the normal case β€” Visual Studio Build Tools with the "Desktop development with C++" workload + Python 3 β€” which is why the win32 single-executable build skips keytar entirely and stores the keys in the executable's own file body instead (see the bullet above):

    Platform Runtime requirement (OS keychain) Node build toolchain (node-gyp)
    macOS Keychain Services (built in) Xcode Command Line Tools (xcode-select --install) + Python 3 β€” only for the source-build fallback
    Linux a Secret Service provider β€” gnome-keyring (or KWallet) daemon running, with libsecret-1 build-essential (gcc/g++/make) + Python 3 + libsecret-1-dev headers β€” only for the source-build fallback
    Windows β€” node dist/server.js Credential Vault (built in) Visual Studio Build Tools with the "Desktop development with C++" workload + Python 3
    Windows β€” single-executable build not used β€” no OS keychain, keys go in the executable's own file body (see the bullet above) not needed β€” keytar is not installed

    Headless Linux servers need a keyring daemon unlocked in the session (e.g. gnome-keyring-daemon --start --components=secrets with DBus session) or keychain access fails. Unsupported environments β€” Docker/distroless containers, Cloudflare Workers β€” have no OS keychain; with the flag enabled the proxy refuses to start (fatal KeyStoreError, no silent fallback). The Docker image installs with --omit=optional --ignore-scripts, so the addon is absent there by design.

  • Environment variables β€” core/server (PORT, LOG_LEVEL, …), config source (PROXY_CONFIG_PATH / PROXY_CONFIG_CONSUL / PROXY_CONFIG_APOLLO), token counting & upstream, and the privacy-filter / compression / image-encode sidecars.

Also see proxy_config.example.toml and docs/getting-started/README_DETAILS.md.

Config field aliases β€” The proxy accepts both canonical and short field names in different config contexts (e.g., base_url/base, upstream_mode/mode, api_key/key). The rules differ between [models.*] inline tables, section-level defaults, and [passthrough] entries. See Config field aliases in the configuration reference for the full mapping tables.

Testing

# Coverage testcases: run from the project root; the runner builds an isolated test config.
node tests/run-integration-tests.js --all

# Agent SDK / provider tests live under ./tests and need the proxy running first.
node tests/multi-agents-test.ts
node tests/multi-agents-composite.ts
npm run test:unit

# Point testcases at a specific proxy / key
PROXY_URL=http://localhost:8788 API_KEY=sk-test node tests/run-integration-tests.js --all
  • Coverage test cases live in tests/integration/; use node tests/run-integration-tests.js --all or selected suite indices.
  • Agent-SDK and provider tests live in tests/.

Documentation

The docs/ folder has deep-dives on specific topics:

  • Sidecar status & config β€” docs/architecture/status_of_sidecars_of_proxy.md (summary of all sidecars: remote auth, privacy filter, image fetch, tool judge, kompress, coordinator, fusion, composite, schedule, transforms)
  • Tool-judge sidecar (Laya) context limit β€” Laya runs with a per-checkpoint token budget (max_len: 512 for convaiinnovations/laya, 1,024 for multilingual/typed-decisions) shared by question prefix + state. The proxy uses character caps (2000/600/50) that don't map to tokens; exceeding the budget silently truncates the tool list β€” later tools vanish from the judge's view. See docs/architecture/design_tool_judge_sidecar_protocol.md and submodules/laya-mlx/README.md.
  • Routing & aliases β€” docs/reference/routing-and-aliases.md (full [models.*] / [composite] / [schedule] / token-limit reference), plus docs/getting-started/proxy_config.example.toml, docs/architecture/routing_refactor.md, docs/architecture/routing_config_revision.md
  • API endpoint details β€” docs/api/api-endpoints.md (dynamic routing, image I/O across formats, prompt-caching fields, Dashboard JSON API)
  • Configuration guide β€” docs/getting-started/configuration-guide.md (minimal proxy_config.toml walkthrough + thinking/reasoning notes)
  • Configuration reference β€” docs/reference/configuration-reference.md (all TOML sections + environment variables)
  • Auth & stats protocol β€” docs/architecture/auth-stats-protocol.md (wire-level contract for the remote auth/stats sidecars)
  • Config loading β€” docs/reference/config_loader.md
  • Live stats β€” docs/guides/live-stats.md (TUI/web dashboard, JSONL usage-dump format, startup stats restoration)
  • Thinking / reasoning β€” docs/guides/claude-extended-thinking.md, docs/guides/claude-adaptive-thinking.md
  • API formats β€” docs/api/claude-api-reference.md, docs/api/gemini-api-reference.md, docs/api/openai-api-reference.md
  • Fusion & composite design β€” docs/architecture/design_fusion_composite_alias.md
  • Request/response transform hooks β€” docs/reference/transforms-reference.md (current reference: hooks, Tier-1 ops, Tier-2 built-ins incl. restore_client_model_alias, [transforms.*] / [transform_defaults] config) β€” plus docs/architecture/design_request_transform_hooks.md (original design) and docs/contributing/implementation_of_request_transform_hooks.md (implementation log)
  • Agent harness integrations β€” docs/guides/agents/ (per-agent guides; e.g. using this proxy as an LLM provider for deepseek-harness)

Interactive agent session (optional)

Start an interactive pi-agent-core agent session that uses the proxy's own /v1/messages endpoint (loopback) as its LLM provider β€” useful for exercising the proxy's routing/quotas/transforms through a real agent loop without a separate client:

npm run server -- --agent
# equivalent:
AGENT=true npm run server

--agent/AGENT and --tui/TUI are mutually exclusive; if both are set, AGENT wins with a warning. Either flag also enables --dashboard (token-stats persistence).

Flow: pick a working directory (tools are confined to it) β†’ system prompt loads AGENTS.md/CLAUDE.md from that directory, with an optional multi-select for already-installed pi skills (global ~/.pi/agent/skills or project .pi/skills), plus skills installable on demand from other agents via the skills CLI β†’ pick a model alias from proxy_config.toml β†’ a verification message confirms it replies β†’ set a budget (tokens and/or turns, default 50,000,000 tokens / 100 turns) β†’ enter a free-text task. The agent runs with read_file/write_file/bash tools (plus find_skill/add_skill when the skills CLI is available, capped at 5 runtime installs/session) until the task completes or the budget is hit, then prompts for a follow-up task in the same conversation. Set TRAJ=true to also log the full session transcript to a private (0600), per-session file under os.tmpdir().

Logging: proxy diagnostics (every console level, including WARN/ERROR) are shown on a single dedicated row above the ── rule rather than written to the terminal directly β€” a raw stderr write lands on the row the TUI parked the cursor on, which is the > input row, and would smear across your prompt. Only the newest line is kept: each one replaces the previous, truncated to the terminal width so it never wraps, and the row takes no space at all until the first log arrives. Console output goes back to the real stderr when the session ends, so background/non-agent logging is unaffected.

Dependencies: pi-agent-core (bundled), and optionally the skills CLI (npm install skills) for the dynamic skill-install tools β€” omitted with a notice if unavailable.

Risks/limits: ⚠️ no OS-level sandbox β€” bash runs via a plain /bin/sh -c child process with the full privileges of the user running the server. Tool safety is limited to non-bypassable but simple raw string/regex checks (not a real shell parser): path confinement to the working directory + /tmp, and a denylist of destructive patterns (rm -rf, kill -9, git push --force, chmod -R 777, piping curl/wget into a shell). These stop obviously destructive self-inflicted commands, not an adversarial or sufficiently obfuscated one. Only run agent sessions against working directories and models you trust. See src/agent-tools.ts for the exact checks.

🀝 Contributing

  1. Fork the repository
  2. Create a feature branch
  3. Make your changes
  4. Submit a pull request

Models and Tools

Models Involved

  1. DeepSeek-R1, V3.2, V4-Flash, V4-Pro, V4.1-Flash
  2. Minimax-M2.6, M2.7-highspeed, M3
  3. Kimi-K2.6, K2.7-Code, K3
  4. GPT-5.4-Mini, GPT-5.4, GPT-5.5, GPT-5.6
  5. Gemini-2.5-Flash, 3.0-Preview, 3.1-Flash
  6. Claude-Sonnet-4.5, Sonnet-4.6, Opus 4.6, Opus 4.8, Fable 5
  7. Nemotron-3-Super-120b, gpt-oss-120b, Nemotron-3-ultra-550b-a55b
  8. GLM-5.2, GLM-5.3
  9. Qwen3.8-Max
  10. Grok-4.6

Tools Involved

  1. Claude-Code
  2. Kiro
  3. Gemini-Cli
  4. Pi-Coding-Agent
  5. Codex
  6. opencode
  7. model_proxy_v3 (For over 90% of its lifecycle, it functions as a local LLM gateway and continuously improves itself with CC and models.)

License

This project is licensed under the MIT License.

About

A lite AI Gateway for any LLM API, any Models and any Agents.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages