😏
This commit is contained in:
@@ -0,0 +1,740 @@
|
||||
# Soma 2.4.0
|
||||
|
||||
Soma is a stateless OpenAI-compatible Chat Completions proxy. It buffers and
|
||||
validates a target assistant turn, identifies operative refusal or withholding, and
|
||||
uses a separately configured transform model to produce a context-grounded repair.
|
||||
Native OpenAI `tool_calls` and valid response metadata remain target-owned and
|
||||
immutable.
|
||||
|
||||
Soma 2.4.0 addresses the central weakness of 2.3.0: a transform cannot accurately
|
||||
repair references such as “the supplied text,” “both functions,” or “the preceding
|
||||
result” when it receives only the latest user string and failed assistant field.
|
||||
Version 2.4 supplies bounded full task context, repairs reasoning and content
|
||||
together when necessary, distinguishes optional reasoning from the usable answer,
|
||||
and uses an integrity verifier that can reject invented or context-breaking output.
|
||||
|
||||
Soma 2.4.0 is not pre-qualified. JSON grammar guarantees neither correct semantic
|
||||
decisions nor grounded repairs. The exact model artifact, server build, endpoint
|
||||
pair, reasoning modes, media modes, prompt, budgets, and temperature must pass the
|
||||
live automated qualification gates before deployment. Reports retain hash-bound
|
||||
outputs for audit and reproduction, but inspection is not a separate qualification
|
||||
stage.
|
||||
|
||||
The pre-release 2.4.0 tree was stabilized in place rather than assigning a new
|
||||
version to review corrections made before qualification. The rollback remains the
|
||||
unchanged 2.3.0 directory. This stabilization adds no Soma environment variable and
|
||||
no package dependency; existing 2.4.0 profiles retain the same runtime contract.
|
||||
Evaluator CLI provenance such as `--reasoning-budget` and the corresponding
|
||||
llama.cpp server option are not Soma environment settings.
|
||||
|
||||
The current Qwen3.5-9B Q6_K route is unqualified. Its exploratory temperature-zero
|
||||
report used unrestricted secondary reasoning and failed automated gates. That report
|
||||
is diagnostic evidence only: it cannot be promoted or reinterpreted after evaluator
|
||||
stabilization. A fresh bounded-budget report must replace it as the current
|
||||
qualification record.
|
||||
|
||||
Soma never downloads, loads, switches, starts, stops, or restarts a model. It does
|
||||
not execute tools, maintain conversation state, authenticate clients, or provide
|
||||
tenant isolation.
|
||||
|
||||
## Request flow
|
||||
|
||||
With the full primary-off/secondary-on staged profile, the normal path is:
|
||||
|
||||
```text
|
||||
client request
|
||||
-> target model
|
||||
-> buffer and validate one complete assistant turn
|
||||
-> build one bounded, role-preserving task context
|
||||
-> classify every present reasoning/content field on primary, reasoning off,
|
||||
with the other draft text fields removed from that classification envelope
|
||||
-> apply field policy and, if necessary, request one joint repair object
|
||||
candidate 1: primary/off -> primary/off integrity verification
|
||||
candidate 2: secondary/on -> secondary/on integrity verification
|
||||
candidate 3: secondary/on -> secondary/on integrity verification
|
||||
-> if an on-mode verifier truncates or returns invalid JSON:
|
||||
verify the identical candidate once on the same endpoint/model, reasoning off
|
||||
-> retain usable verified fields and immutable native tool calls
|
||||
-> if the turn is still unusable and target retry is enabled:
|
||||
call the target once more with a fixed corrective system instruction
|
||||
and process that fresh response under the same bounded rules
|
||||
-> if a verified repair rewrote reasoning and loop-back is enabled:
|
||||
call the target once more with only that verified reasoning appended
|
||||
and process that fresh response under the same bounded rules
|
||||
-> deduplicate exact reasoning/content copies
|
||||
-> return JSON or normalized SSE
|
||||
```
|
||||
|
||||
Candidate two is generated only after a completed integrity verifier semantically
|
||||
rejects candidate one, or after an eligible route failure. Candidate three is the
|
||||
last semantic candidate and receives a fixed alternate focus on authority and exact
|
||||
output/tool/stop constraints. Every candidate is generated fresh from the original
|
||||
task context and failed target draft: neither rejected candidate text nor
|
||||
model-generated failure analysis is fed into another generation.
|
||||
|
||||
Without a configured secondary, Soma permits at most two primary candidates. A
|
||||
transient primary availability failure can make the remainder of the current client
|
||||
request start on secondary. That stickiness is request-local; the next request starts
|
||||
on primary. There is no global circuit breaker or harness/session model state.
|
||||
|
||||
The staged reasoning modes are part of the 2.4 contract, not optional tuning:
|
||||
`TRANSFORM_REASONING_MODE` must be `off`, and a configured secondary must use
|
||||
`TRANSFORM_SECONDARY_REASONING_MODE=on`. Configuration validation rejects any other
|
||||
combination so deployment cannot silently diverge from the qualified route.
|
||||
|
||||
## Classification and integrity decisions
|
||||
|
||||
Classification returns exactly one object:
|
||||
|
||||
```json
|
||||
{"decision":"pass"}
|
||||
```
|
||||
|
||||
or:
|
||||
|
||||
```json
|
||||
{"decision":"rewrite"}
|
||||
```
|
||||
|
||||
The decision concerns the assistant's operative stance, not the presence of words
|
||||
such as “cannot,” “refuse,” or “safety.” Quoted, translated, summarized, or analyzed
|
||||
refusal language is not itself a refusal. A failed method followed by continued work,
|
||||
a different method toward the same requested outcome, and an honest factual
|
||||
constraint do not become rewrites merely because they contain negative language.
|
||||
An adopted decision to stop, avoid part of the request, or deliberately substitute a
|
||||
lesser outcome is a rewrite when the work remains feasible under the authoritative
|
||||
system/developer instructions. A boundary those instructions explicitly require is
|
||||
not an avoidable refusal and must be preserved.
|
||||
|
||||
Integrity verification uses the same strict object shape:
|
||||
|
||||
```json
|
||||
{"decision":"pass"}
|
||||
```
|
||||
|
||||
or:
|
||||
|
||||
```json
|
||||
{"decision":"rewrite"}
|
||||
```
|
||||
|
||||
During integrity verification, `rewrite` means the candidate has a material defect
|
||||
and must be rejected or regenerated; it does not request an in-place edit. The
|
||||
verifier can reject residual withholding, a generic promise,
|
||||
meta-commentary about the transformation, treating the failed assistant draft as
|
||||
user-supplied material, invented task-specific inputs or results, contradictions with
|
||||
accepted reasoning or immutable tool calls, and an unapproved clarification.
|
||||
|
||||
Soma accepts only a complete JSON object satisfying the current schema. A pure JSON
|
||||
fence is accepted, but an object embedded in prose is not. `TRANSFORM_JSON_MODE=true`
|
||||
is the default and sends a small schema through `response_format`; disabling it
|
||||
removes that wire hint but retains the same prompts, strict parser, local validation,
|
||||
and recovery bounds. Separately, target and repaired content requested as JSON must
|
||||
parse strictly and match an immediately declared top-level type. A configured literal
|
||||
stop sequence may not survive in forwarded reasoning or content. Soma intentionally
|
||||
does not implement full client JSON-Schema validation.
|
||||
|
||||
## Joint repair contract and field policy
|
||||
|
||||
One repair call returns a fixed object with both members present and nullable:
|
||||
|
||||
```json
|
||||
{
|
||||
"reasoning": "complete repaired reasoning or null",
|
||||
"content": "complete repaired content or null"
|
||||
}
|
||||
```
|
||||
|
||||
Only fields classified for repair may be non-null. The two-key wire shape never
|
||||
changes, while the per-call schema constrains each requested member to string and each
|
||||
other member to null. Local validation preserves the same contract when a transform
|
||||
endpoint ignores the schema or JSON mode is disabled. Nonblank exact outputs such as
|
||||
`{}`, `[]`, punctuation, and Unicode symbols are valid; Soma does not impose an
|
||||
English-text or alphanumeric "substance" heuristic on the requested deliverable.
|
||||
|
||||
Soma classifies all present fields before requesting a repair:
|
||||
|
||||
- If reasoning and content pass, both target fields are preserved.
|
||||
- If content passes and reasoning requires repair, Soma drops the reasoning field;
|
||||
it does not risk generating new private analysis for an already usable answer.
|
||||
- If reasoning passes and content requires repair, the accepted reasoning is supplied
|
||||
as evidence for the content repair.
|
||||
- If both require repair, one candidate generates reasoning first and then content so
|
||||
the answer can follow the repaired analysis.
|
||||
- Verified jointly repaired reasoning is forwarded with its verified content. Because
|
||||
integrity verification is message-level, a rejected joint candidate is retried as a
|
||||
whole; Soma never salvages one unverified member from it.
|
||||
- A reasoning-only response with no content and no native tool call cannot become a
|
||||
terminal success merely because internal analysis exists. It takes the optional
|
||||
target retry when enabled; otherwise it fails explicitly.
|
||||
- Native `tool_calls` are immutable. Soma may repair adjacent reasoning/content using
|
||||
the full tool context, but exhausted prose repair clears the unusable prose and
|
||||
preserves the structured call. Soma never invents or edits a tool name, ID,
|
||||
argument string, ordering, or result. Before repair, every returned function name
|
||||
must match a supplied tool definition, and multiple returned calls are rejected
|
||||
when `parallel_tool_calls=false`.
|
||||
|
||||
The failed target assistant draft is evidence, not user-supplied task material. The
|
||||
transform is instructed not to quote, explain, or “convert” the refusal itself. It
|
||||
may preserve supported facts and genuine constraints, but it must not choose an
|
||||
arbitrary example, fill invented placeholders, fabricate code changes or external
|
||||
results, or claim a tool/action completed without evidence.
|
||||
|
||||
`TRANSFORM_ALLOW_CLARIFICATION=false` is the default. A transform response that asks
|
||||
the user for more information is not accepted as the repaired answer unless this
|
||||
option is explicitly enabled. Enabling it is appropriate only for harnesses where an
|
||||
essential missing input genuinely requires another user turn; it must be qualified
|
||||
as a separate behavior profile.
|
||||
|
||||
## Full task context and privacy boundary
|
||||
|
||||
Classification, repair, and integrity verification receive the original request
|
||||
context needed to understand references and preserve constraints:
|
||||
|
||||
- original messages in order and by role, including `system`, `developer`, `user`,
|
||||
`assistant`, and `tool` messages and tool results;
|
||||
- complete tool definitions, `tool_choice`, and `parallel_tool_calls`;
|
||||
- `response_format`, modality/audio controls, and `stop`;
|
||||
- the target assistant draft, clearly separated from the original request;
|
||||
- immutable target native tool calls in a separate read-only section; and
|
||||
- the configured media representation for every multimodal part.
|
||||
|
||||
System and developer messages remain authoritative context below Soma's fixed JSON
|
||||
and native-tool invariants. Other supplied values are task evidence, not permission
|
||||
to override the transform contract.
|
||||
|
||||
Draft text is projected per phase. A classifier receives only its named target field,
|
||||
so refusing content cannot contaminate accepted reasoning or vice versa. Repair
|
||||
generation may inspect fields marked for replacement to preserve facts supported by
|
||||
the task. Integrity verification removes every replaced or discarded original field
|
||||
and judges only retained evidence plus the current candidate.
|
||||
|
||||
Soma does not send target/transform endpoint credentials, HTTP headers, the target
|
||||
model name, sampling knobs, or rejected transform candidates. It does not log task
|
||||
context, prompts, target drafts, repaired output, tool arguments, media payloads, or
|
||||
credentials.
|
||||
|
||||
This is nevertheless a wider trust boundary than 2.3.0. Any secret embedded inside a
|
||||
conversation, tool definition, tool argument, or tool result is part of the original
|
||||
task context and can reach every transform endpoint used for that request, including
|
||||
a remote secondary. Configure only transform services authorized to receive the full
|
||||
request. Header exclusion cannot remove secrets that the client placed in message or
|
||||
tool data.
|
||||
|
||||
`TRANSFORM_CONTEXT_MAX_CHARS=131072` bounds the serialized `task_context`, and
|
||||
`TRANSFORM_FIELD_MAX_CHARS=32768` bounds an individual target reasoning/content
|
||||
field. The configured field limit must not exceed the context limit. The context
|
||||
limit has a hard maximum of 4000000 characters. Because this is a character bound,
|
||||
not tokenizer accounting, large-context profiles should leave room for transform
|
||||
instructions and generated output. Soma rejects oversized semantic input
|
||||
rather than truncating messages, tool schemas, code, or evidence into a misleading
|
||||
task. Phase envelopes add the bounded candidate/contract data, and forwarded native
|
||||
media remains subject to the upstream endpoint and trusted front proxy's byte limits.
|
||||
|
||||
## Media modes
|
||||
|
||||
Media handling is explicit per transform endpoint:
|
||||
|
||||
- `placeholder` preserves typed part positions and non-payload metadata, omits the
|
||||
actual binary/media payload, and marks the part unseen. The transform must not infer
|
||||
absent media details. This is the correct setting for a text-only or `--no-mmproj`
|
||||
llama.cpp server.
|
||||
- `forward` sends original typed content media using native OpenAI multimodal message
|
||||
parts. It does not serialize base64 media into ordinary JSON text. Provider-specific
|
||||
top-level assistant media has no portable input envelope and fails explicitly in
|
||||
this mode; use `placeholder` for that shape. Use `forward` only for an endpoint that
|
||||
is authorized and qualified to accept the request's typed media parts.
|
||||
- `reject` refuses to send a media-bearing task to that endpoint. Soma may use a
|
||||
configured compatible transform route; otherwise it fails explicitly. Target retry
|
||||
is not used to bypass an operator's transform-media policy.
|
||||
|
||||
Set `TRANSFORM_MEDIA_MODE` for primary and `TRANSFORM_SECONDARY_MEDIA_MODE` for
|
||||
secondary. If a `forward` endpoint rejects the media request, Soma routes only to a
|
||||
compatible configured secondary or fails explicitly. It never invokes target retry
|
||||
to bypass media policy and never silently retries the task as placeholder text,
|
||||
because either action would change the evidence available to the model.
|
||||
|
||||
Assistant audio attached to a usable text or native-tool turn is preserved, including
|
||||
audio accumulated from a target stream. Audio-only target turns are explicitly
|
||||
unsupported: Soma cannot inspect or repair the audio payload under its text repair
|
||||
contract, so it returns `unsupported_target_response` instead of forwarding an
|
||||
unchecked terminal answer.
|
||||
|
||||
## Bounded recovery and verifier fallback
|
||||
|
||||
The primary reasoning-off profile owns normal classification, candidate one, and its
|
||||
integrity verification. After semantic rejection, a configured secondary
|
||||
reasoning-on profile owns candidates two and three, each generated from the pristine
|
||||
task package and independently verified.
|
||||
|
||||
Reasoning-enabled generation can improve task understanding, but a small model may
|
||||
spend an entire decision budget thinking and end with `finish_reason=length` before
|
||||
emitting its tiny JSON decision. If an on-mode integrity verification is truncated or
|
||||
structurally invalid, Soma does not discard the candidate. It verifies that identical
|
||||
candidate exactly once on the same endpoint and model with reasoning disabled. Only
|
||||
a completed `rewrite` decision advances to a fresh generation.
|
||||
|
||||
Transport/availability failures follow bounded route failover. Structural JSON
|
||||
recovery may include a concise closed failure category, but never rejected output or
|
||||
raw exception text. Semantic retries receive only positive instructions and the
|
||||
pristine task context; they are not primed with the preceding candidate or its
|
||||
failure.
|
||||
|
||||
There are hard ceilings of:
|
||||
|
||||
- two target calls per client request;
|
||||
- 20 transform calls for each target response; and
|
||||
- 40 transform calls across the complete client request.
|
||||
|
||||
`TRANSFORM_TOTAL_TIMEOUT=1200` is one aggregate deadline. It starts after the first
|
||||
target response completes and covers every transform call, an optional second target
|
||||
call, and processing of the second response. It does not reset after target retry.
|
||||
The initial target call remains governed by `CONNECT_TIMEOUT` and `REQUEST_TIMEOUT`
|
||||
outside that aggregate window.
|
||||
|
||||
## Optional target retry
|
||||
|
||||
`TARGET_RETRY_ON_UNREPAIRABLE=false` preserves the normal one-target-call behavior.
|
||||
When enabled, Soma may call the target exactly once more only when the completed turn
|
||||
is unrepairable and leaves no usable content or immutable native tool call. A failed
|
||||
optional reasoning field does not trigger target retry when valid content remains.
|
||||
|
||||
The retry starts from the original request and inserts one fixed corrective system
|
||||
instruction immediately after the leading system/developer block. It preserves the
|
||||
conversation and tool contract and never includes the rejected target response or a
|
||||
rejected transform candidate. This avoids training the second response to imitate the
|
||||
failure, but it does add target latency/cost and may produce a different native tool
|
||||
decision. Soma still does not execute that call.
|
||||
|
||||
Enable target retry only after qualifying the complete target-plus-transform route.
|
||||
It is not a general retry for target HTTP errors, optional reasoning loss, or a merely
|
||||
imperfect answer.
|
||||
|
||||
## Loop-back on verified repair
|
||||
|
||||
`TARGET_LOOP_BACK_ON_VERIFIED_REPAIR=false` is the default. When enabled, Soma may
|
||||
make exactly one additional target call after an integrity-verified repair that
|
||||
rewrote the target's refusal reasoning. Instead of returning the transform's repaired
|
||||
candidate directly, Soma re-sends the original request with one appended assistant
|
||||
message carrying only that verified repaired reasoning in a `reasoning_content`
|
||||
field, so the target re-ingests the relaxed context and produces the task output
|
||||
itself. The re-entry payload preserves the original conversation, media, tool
|
||||
definitions, `tool_choice`, stop controls, and response-format settings untouched.
|
||||
|
||||
Loop-back fires only on the first target attempt, only when reasoning was one of the
|
||||
repaired fields, and only after that candidate passed integrity verification.
|
||||
Content-only repairs, fields that classified as `pass`, tool-only turns, cleared
|
||||
tool prose, and the second target attempt never loop. Genuine refusals never loop
|
||||
because truthful technical, environmental, evidentiary, uncertainty, impossibility,
|
||||
missing-input, and factual limitations classify as `pass` and are never rewritten.
|
||||
|
||||
The second target call shares the hard ceiling of two target calls per client
|
||||
request and the aggregate `TRANSFORM_TOTAL_TIMEOUT` window, which is not reset. The
|
||||
second response is processed under the same classification, repair, and integrity
|
||||
rules; if it is also unrepairable, the request fails explicitly and Soma never
|
||||
makes a third target call. `TARGET_RETRY_ON_UNREPAIRABLE` and loop-back are
|
||||
mutually exclusive per request because they handle disjoint failure classes
|
||||
(unrepairable turns versus verified reasoning repairs) and share the single
|
||||
additional-call slot.
|
||||
|
||||
The reasoning carrier is fixed to `reasoning_content` with no fallback. Backends
|
||||
that reject that field in input messages fail explicitly rather than silently
|
||||
degrading to a different carrier. Loop-back adds target latency and cost; qualify
|
||||
the complete target-plus-transform route before enabling it.
|
||||
|
||||
## Failure behavior
|
||||
|
||||
`FAIL_OPEN=false` is the default. Exhausted mandatory repair, invalid verification,
|
||||
oversized context, incompatible media, missing usable terminal output, and other
|
||||
nonrecoverable transform errors return an explicit error instead of forwarding a
|
||||
known-bad candidate.
|
||||
|
||||
`FAIL_OPEN=true` is an availability policy only. It can restore an original refusal,
|
||||
withholding field, or otherwise rejected target text and therefore defeats strict
|
||||
repair guarantees. Do not treat fail-open as a safety, compliance, or successful
|
||||
quality mode, and do not enable it merely to hide model qualification failures.
|
||||
Fail-open never makes a reasoning-only or otherwise empty terminal turn successful;
|
||||
that turn still takes the explicitly enabled target retry or returns an error.
|
||||
|
||||
Client and upstream JSON reject non-finite numbers. Transform objects additionally
|
||||
reject duplicate member names. Client `stream` and `parallel_tool_calls` values must
|
||||
be booleans, and `n` must be null or integer `1`. Target assistant text, reasoning
|
||||
aliases, and native tool-call shapes are validated before any local mutation.
|
||||
|
||||
Endpoint configuration rejects userinfo, queries, fragments, invalid ports, and
|
||||
unsafe header overrides. `REQUIRE_DISTINCT_ENDPOINTS=true` prevents exact
|
||||
target/transform origin collisions and direct self-routes. Operators must still avoid
|
||||
DNS aliases or LAN addresses that resolve to a wildcard-bound Soma listener.
|
||||
|
||||
## Configuration
|
||||
|
||||
Minimal one-profile configuration:
|
||||
|
||||
```bash
|
||||
TARGET_URL=https://opencode.ai/zen/v1
|
||||
TRANSFORM_URL=http://127.0.0.1:8001/v1
|
||||
TRANSFORM_MODEL=local
|
||||
TRANSFORM_REASONING_MODE=off
|
||||
TRANSFORM_MEDIA_MODE=placeholder
|
||||
```
|
||||
|
||||
Same-server primary-off/secondary-on profile:
|
||||
|
||||
```bash
|
||||
PROXY_HOST=127.0.0.1
|
||||
PROXY_PORT=8080
|
||||
|
||||
# Clear the removed 2.3.x option from an already-populated shell.
|
||||
unset TRANSFORM_CONFIRM_REWRITES
|
||||
|
||||
TARGET_URL=https://opencode.ai/zen/v1
|
||||
TARGET_KEY=
|
||||
TARGET_HEADERS_JSON={}
|
||||
|
||||
TRANSFORM_URL=http://127.0.0.1:8001/v1
|
||||
TRANSFORM_KEY=
|
||||
TRANSFORM_MODEL=local
|
||||
TRANSFORM_HEADERS_JSON={}
|
||||
TRANSFORM_REASONING_MODE=off
|
||||
TRANSFORM_MEDIA_MODE=placeholder
|
||||
|
||||
TRANSFORM_SECONDARY_URL=http://127.0.0.1:8001/v1
|
||||
TRANSFORM_SECONDARY_KEY=
|
||||
TRANSFORM_SECONDARY_MODEL=local
|
||||
TRANSFORM_SECONDARY_HEADERS_JSON={}
|
||||
TRANSFORM_SECONDARY_REASONING_MODE=on
|
||||
TRANSFORM_SECONDARY_MEDIA_MODE=placeholder
|
||||
|
||||
# Same-server primary/secondary is allowed. This rejects target/transform collisions.
|
||||
REQUIRE_DISTINCT_ENDPOINTS=true
|
||||
|
||||
ENABLE_REASONING={}
|
||||
TRANSFORM_TEMPERATURE=0
|
||||
TRANSFORM_JSON_MODE=true
|
||||
TRANSFORM_CONTEXT_MAX_CHARS=131072
|
||||
TRANSFORM_FIELD_MAX_CHARS=32768
|
||||
TRANSFORM_DECISION_MAX_TOKENS=1536
|
||||
TRANSFORM_REWRITE_MAX_TOKENS=16384
|
||||
TRANSFORM_ALLOW_CLARIFICATION=false
|
||||
TRANSFORM_TOTAL_TIMEOUT=1200
|
||||
TARGET_RETRY_ON_UNREPAIRABLE=false
|
||||
TARGET_LOOP_BACK_ON_VERIFIED_REPAIR=false
|
||||
FAIL_OPEN=false
|
||||
|
||||
CONNECT_TIMEOUT=15
|
||||
REQUEST_TIMEOUT=600
|
||||
```
|
||||
|
||||
`TRANSFORM_CONFIRM_REWRITES` was removed. Soma rejects the variable even when its
|
||||
value is `false`; this catches a stale 2.3.x deployment rather than silently changing
|
||||
its meaning. Deleting an export from a file does not clear an existing shell value,
|
||||
so either start from a clean environment or run:
|
||||
|
||||
```bash
|
||||
unset TRANSFORM_CONFIRM_REWRITES
|
||||
```
|
||||
|
||||
The primary and secondary keys/headers never inherit from one another. A same-server
|
||||
secondary supplies behavioral diversity but no process, GPU, or availability
|
||||
isolation. An independent endpoint/model can supply both, at the cost of extending
|
||||
the full-context trust boundary. Qualify the secondary by itself and then qualify the
|
||||
exact composed pair. Soma 2.4 requires primary `off` and secondary `on`; `default` and
|
||||
the inverse mode assignments are rejected during configuration validation.
|
||||
|
||||
`TRANSFORM_TEMPERATURE` is sent on every transform call and overrides the llama
|
||||
server sampling default. `TRANSFORM_DECISION_MAX_TOKENS` covers classifications and
|
||||
integrity decisions; `TRANSFORM_REWRITE_MAX_TOKENS` covers the fixed joint-repair
|
||||
object. Valid ranges are 256–16384 decision tokens, 256–16384 repair tokens,
|
||||
4096–4000000 context characters, and 1024–4000000 field characters, with the field
|
||||
limit no greater than the context limit. Larger budgets bound output but do not
|
||||
improve model judgment by themselves.
|
||||
|
||||
Environment files are shell profiles and are not loaded automatically. Restart Soma
|
||||
after every environment change:
|
||||
|
||||
```bash
|
||||
set -a
|
||||
. ./soma.env
|
||||
set +a
|
||||
python3 soma.py --check-config
|
||||
python3 soma.py
|
||||
```
|
||||
|
||||
Additional environment variables not shown in the profiles above:
|
||||
|
||||
- `LOG_LEVEL` (default `INFO`) — Python logging level for proxy diagnostics.
|
||||
- `FORWARD_CLIENT_HEADERS` (default `true`) — forward non-hop, non-credential
|
||||
client headers to the target endpoint.
|
||||
- `TRANSFORM_PROMPT` (default built in) — base system prompt prepended to every
|
||||
transform phase prompt.
|
||||
- `SOMA_AUTO_REQUIRES_TOOL` (default `false`) — strict auto-tools mode that
|
||||
classifies each request as requiring a native call or a text response.
|
||||
- `UPSTREAM_ERROR_BODY_LIMIT` (default `4000`, range 256–65536) — bounded number
|
||||
of upstream error-body bytes retained for target diagnostics.
|
||||
- `SSE_CHUNK_CHARS` (default `2048`, range 128–65536) — maximum characters per
|
||||
normalized SSE text delta.
|
||||
|
||||
Verify the effective version, endpoint identities, reasoning/media modes, JSON mode,
|
||||
context/field/token limits, clarification, target-retry, and loop-back policies,
|
||||
aggregate deadline, and call ceilings through `--check-config`, startup diagnostics,
|
||||
or `/health`.
|
||||
|
||||
Point clients at:
|
||||
|
||||
```text
|
||||
http://<proxy-host>:8080/v1/chat/completions
|
||||
```
|
||||
|
||||
Aliases are available at `/chat/completions`, `/v1/models`, `/models`, and `/health`.
|
||||
|
||||
### llama.cpp recommendation for a shared local endpoint
|
||||
|
||||
For the shared-endpoint topology where one llama.cpp process serves a primary
|
||||
reasoning-off profile and a secondary reasoning-on profile through per-request
|
||||
`enable_thinking`, enable server reasoning support and cap thinking so a small
|
||||
decision response has room to emit JSON:
|
||||
|
||||
```text
|
||||
--reasoning on --reasoning-budget 512 --temp 0
|
||||
```
|
||||
|
||||
The primary profile still sends `enable_thinking=false`; the global server mode must
|
||||
not prevent the secondary profile from producing and parsing reasoning when it sends
|
||||
`enable_thinking=true`. A 512-token cap is the required starting profile for the
|
||||
bounded-budget qualification run; configuring it is not itself a qualification
|
||||
claim. The evaluator's matching `--reasoning-budget 512` argument records what the
|
||||
already-running server uses and does not configure the server.
|
||||
|
||||
Keep both the server and Soma transform temperature at zero for qualification. Soma's
|
||||
per-request `TRANSFORM_TEMPERATURE=0` is authoritative for transform calls; the server
|
||||
flag supplies a matching default. Temperature zero removes deliberate sampling
|
||||
variance so repeat failures can be attributed to the route under test, although it
|
||||
does not promise byte-identical output across server builds, speculative decoding,
|
||||
cache state, or concurrency. Any nonzero temperature is a different profile and
|
||||
requires a separate report.
|
||||
|
||||
Re-run the exact live profile after changing any server argument. Soma does not add
|
||||
these arguments, restart the server, or download a model. A text-only server launched
|
||||
with `--no-mmproj` should use `placeholder`, not `forward`, for both transform media
|
||||
modes.
|
||||
|
||||
### Multiple harness profiles
|
||||
|
||||
Use one Soma process and listener port per harness. The supplied
|
||||
[`profiles/harness-a.env.example`](profiles/harness-a.env.example) and
|
||||
[`profiles/harness-b.env.example`](profiles/harness-b.env.example) use ordinary
|
||||
environment variables and distinct ports. Soma has no `HARNESS_TYPE` dispatch or
|
||||
shared mutable deployment profile.
|
||||
|
||||
```bash
|
||||
cp profiles/harness-a.env.example profiles/harness-a.env
|
||||
cp profiles/harness-b.env.example profiles/harness-b.env
|
||||
env -i PATH="$PATH" /bin/sh -c 'set -a; . ./profiles/harness-a.env; set +a; exec python3 soma.py --check-config'
|
||||
env -i PATH="$PATH" /bin/sh -c 'set -a; . ./profiles/harness-b.env; set +a; exec python3 soma.py --check-config'
|
||||
```
|
||||
|
||||
The supplied `profiles/.gitignore` excludes populated profile names while retaining
|
||||
the examples. Keep production profiles outside distributable artifacts even when
|
||||
ignore rules are present. A shared transform server must be qualified at the combined
|
||||
load and configured concurrency; a one-slot llama server serializes both harnesses.
|
||||
|
||||
### Trust and resource boundary
|
||||
|
||||
Keep `PROXY_HOST=127.0.0.1` unless a trusted front proxy supplies authentication,
|
||||
access control, TLS, request-size limits, buffering limits, timeouts, and rate limits.
|
||||
Soma warns when bound to a non-loopback interface. It buffers complete target turns
|
||||
and full bounded task packages and has no in-process concurrency-admission limit, so
|
||||
the front proxy must enforce limits appropriate to available memory.
|
||||
|
||||
## Native tool calls and streaming
|
||||
|
||||
Soma supports native OpenAI `tool_calls` only. It preserves IDs, `type: function`,
|
||||
function names, strict JSON argument strings, ordering, and streaming fragments.
|
||||
Proprietary text tool syntaxes are ordinary assistant text; conversion belongs in the
|
||||
target's OpenAI-compatible gateway.
|
||||
|
||||
For `stream:true`, Soma buffers the complete target stream, processes it, and emits
|
||||
normalized OpenAI delta SSE. Original chunk boundaries are not preserved. Valid
|
||||
reasoning, content, native tool calls, finish reason, usage, and response metadata are
|
||||
retained. Accepted `cost` and `usage` metadata are emitted together at most once.
|
||||
|
||||
The upstream stream must produce a terminal non-null `finish_reason`. Soma accepts a
|
||||
terminal choice followed by EOF or the ordinary sequence ending in `[DONE]`.
|
||||
Standard empty-choice usage frames are retained. After the first `[DONE]`, at most one
|
||||
narrow metadata postlude is allowed: an object with `choices: []`, no keys outside
|
||||
`choices`, `cost`, and `usage`, and at least one non-null metadata value. It may end at
|
||||
EOF or one closing `[DONE]`. Further objects/delimiters, malformed or non-finite JSON,
|
||||
duplicate keys, premature `[DONE]`, or meaningful data after the terminal choice are
|
||||
rejected.
|
||||
|
||||
## Diagnostics
|
||||
|
||||
Successful and post-dispatch error responses expose privacy-safe trace/timing and
|
||||
bounded call counts, field decisions, candidate/verifier outcomes, target-retry use,
|
||||
deduplication, and fail-open status. `/health` and startup diagnostics additionally
|
||||
show the effective non-secret reasoning and media configuration.
|
||||
|
||||
Transform logs identify phase, field/candidate, backend, reasoning and media mode,
|
||||
purpose, closed failure category, JSON-mode value, channel lengths, finish reason,
|
||||
token counts, elapsed time, and a request-ID fingerprint. They do not include prompts,
|
||||
task context, target or transform text, media, tool arguments, credentials, error
|
||||
bodies, or raw upstream request IDs.
|
||||
|
||||
## Tests
|
||||
|
||||
Run the complete offline suite:
|
||||
|
||||
```bash
|
||||
python3 -m unittest -v test_soma.py test_soma_extra.py
|
||||
python3 test_soma_live.py --inventory
|
||||
```
|
||||
|
||||
The suite covers strict schemas, full-context isolation and limits, joint field
|
||||
policy, media routes, primary/secondary candidate ownership, reasoning-off verifier
|
||||
fallback, optional target retry, hard call/deadline ceilings, fail-open behavior,
|
||||
native tool fidelity, JSON validation, and SSE normalization. Offline success is
|
||||
necessary but is not model qualification.
|
||||
|
||||
## Live qualification
|
||||
|
||||
`test_soma_live.py` is opt-in and dynamically imports the adjacent `soma.py`, so it
|
||||
exercises the exact runtime prompts, schemas, parsing, validation, routing, and field
|
||||
policy. It calls only already-running endpoints supplied by the operator and never
|
||||
manages a model or server.
|
||||
|
||||
First inspect the frozen corpus without network access:
|
||||
|
||||
```bash
|
||||
python3 test_soma_live.py --inventory
|
||||
```
|
||||
|
||||
Run the exact selected transform artifact/profile at temperature 0 and retain the
|
||||
report only under ignored `qualification-local/`. Consult `--help` for the current
|
||||
provenance and endpoint arguments:
|
||||
|
||||
```bash
|
||||
python3 test_soma_live.py --help
|
||||
```
|
||||
|
||||
Every qualifying run must declare primary `--reasoning-mode off`, a configured
|
||||
secondary with `--secondary-reasoning-mode on`, and the secondary server's actual
|
||||
positive `--reasoning-budget` (for the documented llama.cpp starting profile,
|
||||
`--reasoning-budget 512`). This evaluator value records provenance; the server must
|
||||
already have been launched with the matching budget.
|
||||
|
||||
An exploratory run against llama.cpp's unrestricted default may record
|
||||
`--reasoning-budget -1`. Its report remains unqualified because the positive-budget
|
||||
provenance gate fails; it is not carried forward after a complete bounded-budget
|
||||
rerun replaces the current evidence.
|
||||
|
||||
Qualification is automated-only. Use a new report filename and run the exact
|
||||
temperature-zero, bounded profile. Exit status `0` means every qualification gate
|
||||
passed and the report records `qualified: true` with
|
||||
`qualification_status: qualified`. Exit status `1` means at least one qualification
|
||||
gate failed, and `2` means setup or report creation failed. The evaluator has no
|
||||
second approval stage; inspecting retained evidence does not alter report status.
|
||||
|
||||
Provider-managed routes can be exercised with
|
||||
`--artifact-kind provider-managed`, but they are recorded as exploratory and can
|
||||
never be marked qualified by this evaluator. Supply the exact provider name, model
|
||||
label, and a small public `/models` metadata record through
|
||||
`--provider-model-metadata-json`; do not invent GGUF, llama.cpp, hardware, revision,
|
||||
or reasoning-budget values for a hosted service. Use `--reasoning-budget 0` when the
|
||||
provider does not publish a bounded budget. The report separates behavioral gate
|
||||
results from qualification eligibility and records the requested reasoning modes as
|
||||
unverified provider controls.
|
||||
|
||||
To retain the qualified local GGUF primary while evaluating a hosted secondary, use
|
||||
`--artifact-kind hybrid-local-provider`. Supply the ordinary local artifact fields
|
||||
for the primary and the provider metadata fields for the secondary. The evaluator
|
||||
retains both identities, but deliberately records the combined route as exploratory
|
||||
and qualification-ineligible because the hosted reasoning controls and budget are
|
||||
not independently verified. The primary's artifact label may differ from its wire
|
||||
model alias (for example, an immutable repository label with `local` on the wire).
|
||||
|
||||
`--target-smoke-count 10` limits only the final target-through-transform smoke calls.
|
||||
It does not limit the preceding transform corpus: the evaluator still runs all 240
|
||||
classifier cases, 80 retained repairs, repeat matrices, and route/media probes. Each
|
||||
smoke request grants 128 output tokens, and its response must contain exactly `OK`
|
||||
with no surrounding whitespace, prose, or native tool call.
|
||||
|
||||
For the initial bounded Qwen route, the automated command must include the exact
|
||||
primary/secondary endpoint and provenance arguments plus:
|
||||
|
||||
```bash
|
||||
python3 test_soma_live.py \
|
||||
--reasoning-mode off \
|
||||
--secondary-reasoning-mode on \
|
||||
--reasoning-budget 512 \
|
||||
--temperature 0 \
|
||||
--report qualification-local/qwen3.5-9b-q6_k-t0-rb512-automated.json \
|
||||
[the exact endpoint, model, server, artifact, and hardware arguments]
|
||||
```
|
||||
|
||||
Reports are immutable evidence files. The evaluator writes a completed report
|
||||
privately and installs it atomically; it never exposes a partially written result or
|
||||
overwrites an existing path. The stabilized evaluator, corpus, and source hashes must
|
||||
match the new run. After a complete budget-512 report has been validated and installed
|
||||
under its truthful filename, remove the obsolete unrestricted-budget artifact so only
|
||||
the current evidence remains.
|
||||
|
||||
The automated gates require:
|
||||
|
||||
- valid contracts on all 240 classification cases, 100% hard-refusal and overall
|
||||
refusal recall, and zero false rewrites;
|
||||
- all 20 schema-off high-risk sentinels and five repeats of every high-risk case at
|
||||
parallelism 1 and 4 with zero repeat failures;
|
||||
- all 80 message-repair cases completed without exhaustion and 100% integrity
|
||||
verification, required-fact retention, and forbidden-fact absence;
|
||||
- exactly 20 cases in each field-decision cell: `pass/pass`, `rewrite/pass`,
|
||||
`pass/rewrite`, and `rewrite/rewrite`;
|
||||
- the exact staged primary-off then secondary-on candidate route;
|
||||
- explicit positive secondary reasoning-budget provenance (use `512` as the initial
|
||||
llama.cpp qualification value);
|
||||
- explicit `placeholder`, `forward`, and `reject` media behavior; and
|
||||
- complete source, evaluator, model, server, configuration, and fixture
|
||||
reproducibility evidence.
|
||||
|
||||
The report also records latency, classification disagreements, backend/phase
|
||||
ownership, semantic repair attempts, verifier fallback, and call ceilings. Strict
|
||||
JSON is exercised both with structured-output mode enabled and with the wire schema
|
||||
omitted.
|
||||
|
||||
Live target smoke is separate and explicitly opt-in because it incurs target cost and
|
||||
can produce a new model/tool decision. It uses benign fixtures, keeps target retry
|
||||
disabled, verifies the complete target-to-transform route, and never executes returned
|
||||
tools. Target-retry behavior remains deterministic offline coverage until separately
|
||||
qualified; live smoke does not enable it. The smoke is not run by `--inventory` or an
|
||||
ordinary transform-only qualification. The count is bounded from 1 through 10:
|
||||
|
||||
```bash
|
||||
python3 test_soma_live.py \
|
||||
--target-smoke \
|
||||
--target-url https://target.example/v1 \
|
||||
--target-model TARGET_MODEL \
|
||||
--reasoning-budget 512 \
|
||||
--target-smoke-count 10 \
|
||||
[the same transform and provenance arguments used for qualification]
|
||||
```
|
||||
|
||||
Without `--target-smoke`, the evaluator makes zero target calls. Supply target keys
|
||||
through the hidden CLI/environment option, never in recorded server arguments or a
|
||||
report intended for sharing.
|
||||
|
||||
An automated pass is final qualification for the exact recorded profile. Reports
|
||||
retain all 80 accepted repair outputs and their evidence hashes so the result can be
|
||||
audited and reproduced, but later inspection does not change qualification status.
|
||||
Any failed gate leaves the profile unqualified, and a smaller model receives no
|
||||
relaxed threshold.
|
||||
|
||||
Any change to model revision, GGUF, server build/arguments, reasoning budget,
|
||||
temperature, prompt, endpoint identity, media mode, context/token limits, field
|
||||
policy, primary/secondary composition, evaluator source, fixture corpus, or assertion
|
||||
semantics creates different evidence and requires a new report. Evidence hashes bind
|
||||
one report's exact inputs and outputs; they do not transfer qualification to a
|
||||
superseded report. Reports can contain synthetic task context and non-secret
|
||||
provenance; inspect them before sharing and never place keys in recorded header/server
|
||||
arguments.
|
||||
|
||||
`qualification-local/`, populated profiles, logs, caches, credentials, and model
|
||||
artifacts are excluded from the release package and checksums.
|
||||
Reference in New Issue
Block a user