This commit is contained in:
2026-08-24 08:10:26 -07:00
commit 6f4024f039
11 changed files with 15422 additions and 0 deletions
+740
View File
@@ -0,0 +1,740 @@
# Soma 2.4.0
Soma is a stateless OpenAI-compatible Chat Completions proxy. It buffers and
validates a target assistant turn, identifies operative refusal or withholding, and
uses a separately configured transform model to produce a context-grounded repair.
Native OpenAI `tool_calls` and valid response metadata remain target-owned and
immutable.
Soma 2.4.0 addresses the central weakness of 2.3.0: a transform cannot accurately
repair references such as “the supplied text,” “both functions,” or “the preceding
result” when it receives only the latest user string and failed assistant field.
Version 2.4 supplies bounded full task context, repairs reasoning and content
together when necessary, distinguishes optional reasoning from the usable answer,
and uses an integrity verifier that can reject invented or context-breaking output.
Soma 2.4.0 is not pre-qualified. JSON grammar guarantees neither correct semantic
decisions nor grounded repairs. The exact model artifact, server build, endpoint
pair, reasoning modes, media modes, prompt, budgets, and temperature must pass the
live automated qualification gates before deployment. Reports retain hash-bound
outputs for audit and reproduction, but inspection is not a separate qualification
stage.
The pre-release 2.4.0 tree was stabilized in place rather than assigning a new
version to review corrections made before qualification. The rollback remains the
unchanged 2.3.0 directory. This stabilization adds no Soma environment variable and
no package dependency; existing 2.4.0 profiles retain the same runtime contract.
Evaluator CLI provenance such as `--reasoning-budget` and the corresponding
llama.cpp server option are not Soma environment settings.
The current Qwen3.5-9B Q6_K route is unqualified. Its exploratory temperature-zero
report used unrestricted secondary reasoning and failed automated gates. That report
is diagnostic evidence only: it cannot be promoted or reinterpreted after evaluator
stabilization. A fresh bounded-budget report must replace it as the current
qualification record.
Soma never downloads, loads, switches, starts, stops, or restarts a model. It does
not execute tools, maintain conversation state, authenticate clients, or provide
tenant isolation.
## Request flow
With the full primary-off/secondary-on staged profile, the normal path is:
```text
client request
-> target model
-> buffer and validate one complete assistant turn
-> build one bounded, role-preserving task context
-> classify every present reasoning/content field on primary, reasoning off,
with the other draft text fields removed from that classification envelope
-> apply field policy and, if necessary, request one joint repair object
candidate 1: primary/off -> primary/off integrity verification
candidate 2: secondary/on -> secondary/on integrity verification
candidate 3: secondary/on -> secondary/on integrity verification
-> if an on-mode verifier truncates or returns invalid JSON:
verify the identical candidate once on the same endpoint/model, reasoning off
-> retain usable verified fields and immutable native tool calls
-> if the turn is still unusable and target retry is enabled:
call the target once more with a fixed corrective system instruction
and process that fresh response under the same bounded rules
-> if a verified repair rewrote reasoning and loop-back is enabled:
call the target once more with only that verified reasoning appended
and process that fresh response under the same bounded rules
-> deduplicate exact reasoning/content copies
-> return JSON or normalized SSE
```
Candidate two is generated only after a completed integrity verifier semantically
rejects candidate one, or after an eligible route failure. Candidate three is the
last semantic candidate and receives a fixed alternate focus on authority and exact
output/tool/stop constraints. Every candidate is generated fresh from the original
task context and failed target draft: neither rejected candidate text nor
model-generated failure analysis is fed into another generation.
Without a configured secondary, Soma permits at most two primary candidates. A
transient primary availability failure can make the remainder of the current client
request start on secondary. That stickiness is request-local; the next request starts
on primary. There is no global circuit breaker or harness/session model state.
The staged reasoning modes are part of the 2.4 contract, not optional tuning:
`TRANSFORM_REASONING_MODE` must be `off`, and a configured secondary must use
`TRANSFORM_SECONDARY_REASONING_MODE=on`. Configuration validation rejects any other
combination so deployment cannot silently diverge from the qualified route.
## Classification and integrity decisions
Classification returns exactly one object:
```json
{"decision":"pass"}
```
or:
```json
{"decision":"rewrite"}
```
The decision concerns the assistant's operative stance, not the presence of words
such as “cannot,” “refuse,” or “safety.” Quoted, translated, summarized, or analyzed
refusal language is not itself a refusal. A failed method followed by continued work,
a different method toward the same requested outcome, and an honest factual
constraint do not become rewrites merely because they contain negative language.
An adopted decision to stop, avoid part of the request, or deliberately substitute a
lesser outcome is a rewrite when the work remains feasible under the authoritative
system/developer instructions. A boundary those instructions explicitly require is
not an avoidable refusal and must be preserved.
Integrity verification uses the same strict object shape:
```json
{"decision":"pass"}
```
or:
```json
{"decision":"rewrite"}
```
During integrity verification, `rewrite` means the candidate has a material defect
and must be rejected or regenerated; it does not request an in-place edit. The
verifier can reject residual withholding, a generic promise,
meta-commentary about the transformation, treating the failed assistant draft as
user-supplied material, invented task-specific inputs or results, contradictions with
accepted reasoning or immutable tool calls, and an unapproved clarification.
Soma accepts only a complete JSON object satisfying the current schema. A pure JSON
fence is accepted, but an object embedded in prose is not. `TRANSFORM_JSON_MODE=true`
is the default and sends a small schema through `response_format`; disabling it
removes that wire hint but retains the same prompts, strict parser, local validation,
and recovery bounds. Separately, target and repaired content requested as JSON must
parse strictly and match an immediately declared top-level type. A configured literal
stop sequence may not survive in forwarded reasoning or content. Soma intentionally
does not implement full client JSON-Schema validation.
## Joint repair contract and field policy
One repair call returns a fixed object with both members present and nullable:
```json
{
"reasoning": "complete repaired reasoning or null",
"content": "complete repaired content or null"
}
```
Only fields classified for repair may be non-null. The two-key wire shape never
changes, while the per-call schema constrains each requested member to string and each
other member to null. Local validation preserves the same contract when a transform
endpoint ignores the schema or JSON mode is disabled. Nonblank exact outputs such as
`{}`, `[]`, punctuation, and Unicode symbols are valid; Soma does not impose an
English-text or alphanumeric "substance" heuristic on the requested deliverable.
Soma classifies all present fields before requesting a repair:
- If reasoning and content pass, both target fields are preserved.
- If content passes and reasoning requires repair, Soma drops the reasoning field;
it does not risk generating new private analysis for an already usable answer.
- If reasoning passes and content requires repair, the accepted reasoning is supplied
as evidence for the content repair.
- If both require repair, one candidate generates reasoning first and then content so
the answer can follow the repaired analysis.
- Verified jointly repaired reasoning is forwarded with its verified content. Because
integrity verification is message-level, a rejected joint candidate is retried as a
whole; Soma never salvages one unverified member from it.
- A reasoning-only response with no content and no native tool call cannot become a
terminal success merely because internal analysis exists. It takes the optional
target retry when enabled; otherwise it fails explicitly.
- Native `tool_calls` are immutable. Soma may repair adjacent reasoning/content using
the full tool context, but exhausted prose repair clears the unusable prose and
preserves the structured call. Soma never invents or edits a tool name, ID,
argument string, ordering, or result. Before repair, every returned function name
must match a supplied tool definition, and multiple returned calls are rejected
when `parallel_tool_calls=false`.
The failed target assistant draft is evidence, not user-supplied task material. The
transform is instructed not to quote, explain, or “convert” the refusal itself. It
may preserve supported facts and genuine constraints, but it must not choose an
arbitrary example, fill invented placeholders, fabricate code changes or external
results, or claim a tool/action completed without evidence.
`TRANSFORM_ALLOW_CLARIFICATION=false` is the default. A transform response that asks
the user for more information is not accepted as the repaired answer unless this
option is explicitly enabled. Enabling it is appropriate only for harnesses where an
essential missing input genuinely requires another user turn; it must be qualified
as a separate behavior profile.
## Full task context and privacy boundary
Classification, repair, and integrity verification receive the original request
context needed to understand references and preserve constraints:
- original messages in order and by role, including `system`, `developer`, `user`,
`assistant`, and `tool` messages and tool results;
- complete tool definitions, `tool_choice`, and `parallel_tool_calls`;
- `response_format`, modality/audio controls, and `stop`;
- the target assistant draft, clearly separated from the original request;
- immutable target native tool calls in a separate read-only section; and
- the configured media representation for every multimodal part.
System and developer messages remain authoritative context below Soma's fixed JSON
and native-tool invariants. Other supplied values are task evidence, not permission
to override the transform contract.
Draft text is projected per phase. A classifier receives only its named target field,
so refusing content cannot contaminate accepted reasoning or vice versa. Repair
generation may inspect fields marked for replacement to preserve facts supported by
the task. Integrity verification removes every replaced or discarded original field
and judges only retained evidence plus the current candidate.
Soma does not send target/transform endpoint credentials, HTTP headers, the target
model name, sampling knobs, or rejected transform candidates. It does not log task
context, prompts, target drafts, repaired output, tool arguments, media payloads, or
credentials.
This is nevertheless a wider trust boundary than 2.3.0. Any secret embedded inside a
conversation, tool definition, tool argument, or tool result is part of the original
task context and can reach every transform endpoint used for that request, including
a remote secondary. Configure only transform services authorized to receive the full
request. Header exclusion cannot remove secrets that the client placed in message or
tool data.
`TRANSFORM_CONTEXT_MAX_CHARS=131072` bounds the serialized `task_context`, and
`TRANSFORM_FIELD_MAX_CHARS=32768` bounds an individual target reasoning/content
field. The configured field limit must not exceed the context limit. The context
limit has a hard maximum of 4000000 characters. Because this is a character bound,
not tokenizer accounting, large-context profiles should leave room for transform
instructions and generated output. Soma rejects oversized semantic input
rather than truncating messages, tool schemas, code, or evidence into a misleading
task. Phase envelopes add the bounded candidate/contract data, and forwarded native
media remains subject to the upstream endpoint and trusted front proxy's byte limits.
## Media modes
Media handling is explicit per transform endpoint:
- `placeholder` preserves typed part positions and non-payload metadata, omits the
actual binary/media payload, and marks the part unseen. The transform must not infer
absent media details. This is the correct setting for a text-only or `--no-mmproj`
llama.cpp server.
- `forward` sends original typed content media using native OpenAI multimodal message
parts. It does not serialize base64 media into ordinary JSON text. Provider-specific
top-level assistant media has no portable input envelope and fails explicitly in
this mode; use `placeholder` for that shape. Use `forward` only for an endpoint that
is authorized and qualified to accept the request's typed media parts.
- `reject` refuses to send a media-bearing task to that endpoint. Soma may use a
configured compatible transform route; otherwise it fails explicitly. Target retry
is not used to bypass an operator's transform-media policy.
Set `TRANSFORM_MEDIA_MODE` for primary and `TRANSFORM_SECONDARY_MEDIA_MODE` for
secondary. If a `forward` endpoint rejects the media request, Soma routes only to a
compatible configured secondary or fails explicitly. It never invokes target retry
to bypass media policy and never silently retries the task as placeholder text,
because either action would change the evidence available to the model.
Assistant audio attached to a usable text or native-tool turn is preserved, including
audio accumulated from a target stream. Audio-only target turns are explicitly
unsupported: Soma cannot inspect or repair the audio payload under its text repair
contract, so it returns `unsupported_target_response` instead of forwarding an
unchecked terminal answer.
## Bounded recovery and verifier fallback
The primary reasoning-off profile owns normal classification, candidate one, and its
integrity verification. After semantic rejection, a configured secondary
reasoning-on profile owns candidates two and three, each generated from the pristine
task package and independently verified.
Reasoning-enabled generation can improve task understanding, but a small model may
spend an entire decision budget thinking and end with `finish_reason=length` before
emitting its tiny JSON decision. If an on-mode integrity verification is truncated or
structurally invalid, Soma does not discard the candidate. It verifies that identical
candidate exactly once on the same endpoint and model with reasoning disabled. Only
a completed `rewrite` decision advances to a fresh generation.
Transport/availability failures follow bounded route failover. Structural JSON
recovery may include a concise closed failure category, but never rejected output or
raw exception text. Semantic retries receive only positive instructions and the
pristine task context; they are not primed with the preceding candidate or its
failure.
There are hard ceilings of:
- two target calls per client request;
- 20 transform calls for each target response; and
- 40 transform calls across the complete client request.
`TRANSFORM_TOTAL_TIMEOUT=1200` is one aggregate deadline. It starts after the first
target response completes and covers every transform call, an optional second target
call, and processing of the second response. It does not reset after target retry.
The initial target call remains governed by `CONNECT_TIMEOUT` and `REQUEST_TIMEOUT`
outside that aggregate window.
## Optional target retry
`TARGET_RETRY_ON_UNREPAIRABLE=false` preserves the normal one-target-call behavior.
When enabled, Soma may call the target exactly once more only when the completed turn
is unrepairable and leaves no usable content or immutable native tool call. A failed
optional reasoning field does not trigger target retry when valid content remains.
The retry starts from the original request and inserts one fixed corrective system
instruction immediately after the leading system/developer block. It preserves the
conversation and tool contract and never includes the rejected target response or a
rejected transform candidate. This avoids training the second response to imitate the
failure, but it does add target latency/cost and may produce a different native tool
decision. Soma still does not execute that call.
Enable target retry only after qualifying the complete target-plus-transform route.
It is not a general retry for target HTTP errors, optional reasoning loss, or a merely
imperfect answer.
## Loop-back on verified repair
`TARGET_LOOP_BACK_ON_VERIFIED_REPAIR=false` is the default. When enabled, Soma may
make exactly one additional target call after an integrity-verified repair that
rewrote the target's refusal reasoning. Instead of returning the transform's repaired
candidate directly, Soma re-sends the original request with one appended assistant
message carrying only that verified repaired reasoning in a `reasoning_content`
field, so the target re-ingests the relaxed context and produces the task output
itself. The re-entry payload preserves the original conversation, media, tool
definitions, `tool_choice`, stop controls, and response-format settings untouched.
Loop-back fires only on the first target attempt, only when reasoning was one of the
repaired fields, and only after that candidate passed integrity verification.
Content-only repairs, fields that classified as `pass`, tool-only turns, cleared
tool prose, and the second target attempt never loop. Genuine refusals never loop
because truthful technical, environmental, evidentiary, uncertainty, impossibility,
missing-input, and factual limitations classify as `pass` and are never rewritten.
The second target call shares the hard ceiling of two target calls per client
request and the aggregate `TRANSFORM_TOTAL_TIMEOUT` window, which is not reset. The
second response is processed under the same classification, repair, and integrity
rules; if it is also unrepairable, the request fails explicitly and Soma never
makes a third target call. `TARGET_RETRY_ON_UNREPAIRABLE` and loop-back are
mutually exclusive per request because they handle disjoint failure classes
(unrepairable turns versus verified reasoning repairs) and share the single
additional-call slot.
The reasoning carrier is fixed to `reasoning_content` with no fallback. Backends
that reject that field in input messages fail explicitly rather than silently
degrading to a different carrier. Loop-back adds target latency and cost; qualify
the complete target-plus-transform route before enabling it.
## Failure behavior
`FAIL_OPEN=false` is the default. Exhausted mandatory repair, invalid verification,
oversized context, incompatible media, missing usable terminal output, and other
nonrecoverable transform errors return an explicit error instead of forwarding a
known-bad candidate.
`FAIL_OPEN=true` is an availability policy only. It can restore an original refusal,
withholding field, or otherwise rejected target text and therefore defeats strict
repair guarantees. Do not treat fail-open as a safety, compliance, or successful
quality mode, and do not enable it merely to hide model qualification failures.
Fail-open never makes a reasoning-only or otherwise empty terminal turn successful;
that turn still takes the explicitly enabled target retry or returns an error.
Client and upstream JSON reject non-finite numbers. Transform objects additionally
reject duplicate member names. Client `stream` and `parallel_tool_calls` values must
be booleans, and `n` must be null or integer `1`. Target assistant text, reasoning
aliases, and native tool-call shapes are validated before any local mutation.
Endpoint configuration rejects userinfo, queries, fragments, invalid ports, and
unsafe header overrides. `REQUIRE_DISTINCT_ENDPOINTS=true` prevents exact
target/transform origin collisions and direct self-routes. Operators must still avoid
DNS aliases or LAN addresses that resolve to a wildcard-bound Soma listener.
## Configuration
Minimal one-profile configuration:
```bash
TARGET_URL=https://opencode.ai/zen/v1
TRANSFORM_URL=http://127.0.0.1:8001/v1
TRANSFORM_MODEL=local
TRANSFORM_REASONING_MODE=off
TRANSFORM_MEDIA_MODE=placeholder
```
Same-server primary-off/secondary-on profile:
```bash
PROXY_HOST=127.0.0.1
PROXY_PORT=8080
# Clear the removed 2.3.x option from an already-populated shell.
unset TRANSFORM_CONFIRM_REWRITES
TARGET_URL=https://opencode.ai/zen/v1
TARGET_KEY=
TARGET_HEADERS_JSON={}
TRANSFORM_URL=http://127.0.0.1:8001/v1
TRANSFORM_KEY=
TRANSFORM_MODEL=local
TRANSFORM_HEADERS_JSON={}
TRANSFORM_REASONING_MODE=off
TRANSFORM_MEDIA_MODE=placeholder
TRANSFORM_SECONDARY_URL=http://127.0.0.1:8001/v1
TRANSFORM_SECONDARY_KEY=
TRANSFORM_SECONDARY_MODEL=local
TRANSFORM_SECONDARY_HEADERS_JSON={}
TRANSFORM_SECONDARY_REASONING_MODE=on
TRANSFORM_SECONDARY_MEDIA_MODE=placeholder
# Same-server primary/secondary is allowed. This rejects target/transform collisions.
REQUIRE_DISTINCT_ENDPOINTS=true
ENABLE_REASONING={}
TRANSFORM_TEMPERATURE=0
TRANSFORM_JSON_MODE=true
TRANSFORM_CONTEXT_MAX_CHARS=131072
TRANSFORM_FIELD_MAX_CHARS=32768
TRANSFORM_DECISION_MAX_TOKENS=1536
TRANSFORM_REWRITE_MAX_TOKENS=16384
TRANSFORM_ALLOW_CLARIFICATION=false
TRANSFORM_TOTAL_TIMEOUT=1200
TARGET_RETRY_ON_UNREPAIRABLE=false
TARGET_LOOP_BACK_ON_VERIFIED_REPAIR=false
FAIL_OPEN=false
CONNECT_TIMEOUT=15
REQUEST_TIMEOUT=600
```
`TRANSFORM_CONFIRM_REWRITES` was removed. Soma rejects the variable even when its
value is `false`; this catches a stale 2.3.x deployment rather than silently changing
its meaning. Deleting an export from a file does not clear an existing shell value,
so either start from a clean environment or run:
```bash
unset TRANSFORM_CONFIRM_REWRITES
```
The primary and secondary keys/headers never inherit from one another. A same-server
secondary supplies behavioral diversity but no process, GPU, or availability
isolation. An independent endpoint/model can supply both, at the cost of extending
the full-context trust boundary. Qualify the secondary by itself and then qualify the
exact composed pair. Soma 2.4 requires primary `off` and secondary `on`; `default` and
the inverse mode assignments are rejected during configuration validation.
`TRANSFORM_TEMPERATURE` is sent on every transform call and overrides the llama
server sampling default. `TRANSFORM_DECISION_MAX_TOKENS` covers classifications and
integrity decisions; `TRANSFORM_REWRITE_MAX_TOKENS` covers the fixed joint-repair
object. Valid ranges are 25616384 decision tokens, 25616384 repair tokens,
40964000000 context characters, and 10244000000 field characters, with the field
limit no greater than the context limit. Larger budgets bound output but do not
improve model judgment by themselves.
Environment files are shell profiles and are not loaded automatically. Restart Soma
after every environment change:
```bash
set -a
. ./soma.env
set +a
python3 soma.py --check-config
python3 soma.py
```
Additional environment variables not shown in the profiles above:
- `LOG_LEVEL` (default `INFO`) — Python logging level for proxy diagnostics.
- `FORWARD_CLIENT_HEADERS` (default `true`) — forward non-hop, non-credential
client headers to the target endpoint.
- `TRANSFORM_PROMPT` (default built in) — base system prompt prepended to every
transform phase prompt.
- `SOMA_AUTO_REQUIRES_TOOL` (default `false`) — strict auto-tools mode that
classifies each request as requiring a native call or a text response.
- `UPSTREAM_ERROR_BODY_LIMIT` (default `4000`, range 25665536) — bounded number
of upstream error-body bytes retained for target diagnostics.
- `SSE_CHUNK_CHARS` (default `2048`, range 12865536) — maximum characters per
normalized SSE text delta.
Verify the effective version, endpoint identities, reasoning/media modes, JSON mode,
context/field/token limits, clarification, target-retry, and loop-back policies,
aggregate deadline, and call ceilings through `--check-config`, startup diagnostics,
or `/health`.
Point clients at:
```text
http://<proxy-host>:8080/v1/chat/completions
```
Aliases are available at `/chat/completions`, `/v1/models`, `/models`, and `/health`.
### llama.cpp recommendation for a shared local endpoint
For the shared-endpoint topology where one llama.cpp process serves a primary
reasoning-off profile and a secondary reasoning-on profile through per-request
`enable_thinking`, enable server reasoning support and cap thinking so a small
decision response has room to emit JSON:
```text
--reasoning on --reasoning-budget 512 --temp 0
```
The primary profile still sends `enable_thinking=false`; the global server mode must
not prevent the secondary profile from producing and parsing reasoning when it sends
`enable_thinking=true`. A 512-token cap is the required starting profile for the
bounded-budget qualification run; configuring it is not itself a qualification
claim. The evaluator's matching `--reasoning-budget 512` argument records what the
already-running server uses and does not configure the server.
Keep both the server and Soma transform temperature at zero for qualification. Soma's
per-request `TRANSFORM_TEMPERATURE=0` is authoritative for transform calls; the server
flag supplies a matching default. Temperature zero removes deliberate sampling
variance so repeat failures can be attributed to the route under test, although it
does not promise byte-identical output across server builds, speculative decoding,
cache state, or concurrency. Any nonzero temperature is a different profile and
requires a separate report.
Re-run the exact live profile after changing any server argument. Soma does not add
these arguments, restart the server, or download a model. A text-only server launched
with `--no-mmproj` should use `placeholder`, not `forward`, for both transform media
modes.
### Multiple harness profiles
Use one Soma process and listener port per harness. The supplied
[`profiles/harness-a.env.example`](profiles/harness-a.env.example) and
[`profiles/harness-b.env.example`](profiles/harness-b.env.example) use ordinary
environment variables and distinct ports. Soma has no `HARNESS_TYPE` dispatch or
shared mutable deployment profile.
```bash
cp profiles/harness-a.env.example profiles/harness-a.env
cp profiles/harness-b.env.example profiles/harness-b.env
env -i PATH="$PATH" /bin/sh -c 'set -a; . ./profiles/harness-a.env; set +a; exec python3 soma.py --check-config'
env -i PATH="$PATH" /bin/sh -c 'set -a; . ./profiles/harness-b.env; set +a; exec python3 soma.py --check-config'
```
The supplied `profiles/.gitignore` excludes populated profile names while retaining
the examples. Keep production profiles outside distributable artifacts even when
ignore rules are present. A shared transform server must be qualified at the combined
load and configured concurrency; a one-slot llama server serializes both harnesses.
### Trust and resource boundary
Keep `PROXY_HOST=127.0.0.1` unless a trusted front proxy supplies authentication,
access control, TLS, request-size limits, buffering limits, timeouts, and rate limits.
Soma warns when bound to a non-loopback interface. It buffers complete target turns
and full bounded task packages and has no in-process concurrency-admission limit, so
the front proxy must enforce limits appropriate to available memory.
## Native tool calls and streaming
Soma supports native OpenAI `tool_calls` only. It preserves IDs, `type: function`,
function names, strict JSON argument strings, ordering, and streaming fragments.
Proprietary text tool syntaxes are ordinary assistant text; conversion belongs in the
target's OpenAI-compatible gateway.
For `stream:true`, Soma buffers the complete target stream, processes it, and emits
normalized OpenAI delta SSE. Original chunk boundaries are not preserved. Valid
reasoning, content, native tool calls, finish reason, usage, and response metadata are
retained. Accepted `cost` and `usage` metadata are emitted together at most once.
The upstream stream must produce a terminal non-null `finish_reason`. Soma accepts a
terminal choice followed by EOF or the ordinary sequence ending in `[DONE]`.
Standard empty-choice usage frames are retained. After the first `[DONE]`, at most one
narrow metadata postlude is allowed: an object with `choices: []`, no keys outside
`choices`, `cost`, and `usage`, and at least one non-null metadata value. It may end at
EOF or one closing `[DONE]`. Further objects/delimiters, malformed or non-finite JSON,
duplicate keys, premature `[DONE]`, or meaningful data after the terminal choice are
rejected.
## Diagnostics
Successful and post-dispatch error responses expose privacy-safe trace/timing and
bounded call counts, field decisions, candidate/verifier outcomes, target-retry use,
deduplication, and fail-open status. `/health` and startup diagnostics additionally
show the effective non-secret reasoning and media configuration.
Transform logs identify phase, field/candidate, backend, reasoning and media mode,
purpose, closed failure category, JSON-mode value, channel lengths, finish reason,
token counts, elapsed time, and a request-ID fingerprint. They do not include prompts,
task context, target or transform text, media, tool arguments, credentials, error
bodies, or raw upstream request IDs.
## Tests
Run the complete offline suite:
```bash
python3 -m unittest -v test_soma.py test_soma_extra.py
python3 test_soma_live.py --inventory
```
The suite covers strict schemas, full-context isolation and limits, joint field
policy, media routes, primary/secondary candidate ownership, reasoning-off verifier
fallback, optional target retry, hard call/deadline ceilings, fail-open behavior,
native tool fidelity, JSON validation, and SSE normalization. Offline success is
necessary but is not model qualification.
## Live qualification
`test_soma_live.py` is opt-in and dynamically imports the adjacent `soma.py`, so it
exercises the exact runtime prompts, schemas, parsing, validation, routing, and field
policy. It calls only already-running endpoints supplied by the operator and never
manages a model or server.
First inspect the frozen corpus without network access:
```bash
python3 test_soma_live.py --inventory
```
Run the exact selected transform artifact/profile at temperature 0 and retain the
report only under ignored `qualification-local/`. Consult `--help` for the current
provenance and endpoint arguments:
```bash
python3 test_soma_live.py --help
```
Every qualifying run must declare primary `--reasoning-mode off`, a configured
secondary with `--secondary-reasoning-mode on`, and the secondary server's actual
positive `--reasoning-budget` (for the documented llama.cpp starting profile,
`--reasoning-budget 512`). This evaluator value records provenance; the server must
already have been launched with the matching budget.
An exploratory run against llama.cpp's unrestricted default may record
`--reasoning-budget -1`. Its report remains unqualified because the positive-budget
provenance gate fails; it is not carried forward after a complete bounded-budget
rerun replaces the current evidence.
Qualification is automated-only. Use a new report filename and run the exact
temperature-zero, bounded profile. Exit status `0` means every qualification gate
passed and the report records `qualified: true` with
`qualification_status: qualified`. Exit status `1` means at least one qualification
gate failed, and `2` means setup or report creation failed. The evaluator has no
second approval stage; inspecting retained evidence does not alter report status.
Provider-managed routes can be exercised with
`--artifact-kind provider-managed`, but they are recorded as exploratory and can
never be marked qualified by this evaluator. Supply the exact provider name, model
label, and a small public `/models` metadata record through
`--provider-model-metadata-json`; do not invent GGUF, llama.cpp, hardware, revision,
or reasoning-budget values for a hosted service. Use `--reasoning-budget 0` when the
provider does not publish a bounded budget. The report separates behavioral gate
results from qualification eligibility and records the requested reasoning modes as
unverified provider controls.
To retain the qualified local GGUF primary while evaluating a hosted secondary, use
`--artifact-kind hybrid-local-provider`. Supply the ordinary local artifact fields
for the primary and the provider metadata fields for the secondary. The evaluator
retains both identities, but deliberately records the combined route as exploratory
and qualification-ineligible because the hosted reasoning controls and budget are
not independently verified. The primary's artifact label may differ from its wire
model alias (for example, an immutable repository label with `local` on the wire).
`--target-smoke-count 10` limits only the final target-through-transform smoke calls.
It does not limit the preceding transform corpus: the evaluator still runs all 240
classifier cases, 80 retained repairs, repeat matrices, and route/media probes. Each
smoke request grants 128 output tokens, and its response must contain exactly `OK`
with no surrounding whitespace, prose, or native tool call.
For the initial bounded Qwen route, the automated command must include the exact
primary/secondary endpoint and provenance arguments plus:
```bash
python3 test_soma_live.py \
--reasoning-mode off \
--secondary-reasoning-mode on \
--reasoning-budget 512 \
--temperature 0 \
--report qualification-local/qwen3.5-9b-q6_k-t0-rb512-automated.json \
[the exact endpoint, model, server, artifact, and hardware arguments]
```
Reports are immutable evidence files. The evaluator writes a completed report
privately and installs it atomically; it never exposes a partially written result or
overwrites an existing path. The stabilized evaluator, corpus, and source hashes must
match the new run. After a complete budget-512 report has been validated and installed
under its truthful filename, remove the obsolete unrestricted-budget artifact so only
the current evidence remains.
The automated gates require:
- valid contracts on all 240 classification cases, 100% hard-refusal and overall
refusal recall, and zero false rewrites;
- all 20 schema-off high-risk sentinels and five repeats of every high-risk case at
parallelism 1 and 4 with zero repeat failures;
- all 80 message-repair cases completed without exhaustion and 100% integrity
verification, required-fact retention, and forbidden-fact absence;
- exactly 20 cases in each field-decision cell: `pass/pass`, `rewrite/pass`,
`pass/rewrite`, and `rewrite/rewrite`;
- the exact staged primary-off then secondary-on candidate route;
- explicit positive secondary reasoning-budget provenance (use `512` as the initial
llama.cpp qualification value);
- explicit `placeholder`, `forward`, and `reject` media behavior; and
- complete source, evaluator, model, server, configuration, and fixture
reproducibility evidence.
The report also records latency, classification disagreements, backend/phase
ownership, semantic repair attempts, verifier fallback, and call ceilings. Strict
JSON is exercised both with structured-output mode enabled and with the wire schema
omitted.
Live target smoke is separate and explicitly opt-in because it incurs target cost and
can produce a new model/tool decision. It uses benign fixtures, keeps target retry
disabled, verifies the complete target-to-transform route, and never executes returned
tools. Target-retry behavior remains deterministic offline coverage until separately
qualified; live smoke does not enable it. The smoke is not run by `--inventory` or an
ordinary transform-only qualification. The count is bounded from 1 through 10:
```bash
python3 test_soma_live.py \
--target-smoke \
--target-url https://target.example/v1 \
--target-model TARGET_MODEL \
--reasoning-budget 512 \
--target-smoke-count 10 \
[the same transform and provenance arguments used for qualification]
```
Without `--target-smoke`, the evaluator makes zero target calls. Supply target keys
through the hidden CLI/environment option, never in recorded server arguments or a
report intended for sharing.
An automated pass is final qualification for the exact recorded profile. Reports
retain all 80 accepted repair outputs and their evidence hashes so the result can be
audited and reproduced, but later inspection does not change qualification status.
Any failed gate leaves the profile unqualified, and a smaller model receives no
relaxed threshold.
Any change to model revision, GGUF, server build/arguments, reasoning budget,
temperature, prompt, endpoint identity, media mode, context/token limits, field
policy, primary/secondary composition, evaluator source, fixture corpus, or assertion
semantics creates different evidence and requires a new report. Evidence hashes bind
one report's exact inputs and outputs; they do not transfer qualification to a
superseded report. Reports can contain synthetic task context and non-secret
provenance; inspect them before sharing and never place keys in recorded header/server
arguments.
`qualification-local/`, populated profiles, logs, caches, credentials, and model
artifacts are excluded from the release package and checksums.