Files
soma/README.md
T
2026-08-24 08:10:26 -07:00

37 KiB
Raw Blame History

Soma 2.4.0

Soma is a stateless OpenAI-compatible Chat Completions proxy. It buffers and validates a target assistant turn, identifies operative refusal or withholding, and uses a separately configured transform model to produce a context-grounded repair. Native OpenAI tool_calls and valid response metadata remain target-owned and immutable.

Soma 2.4.0 addresses the central weakness of 2.3.0: a transform cannot accurately repair references such as “the supplied text,” “both functions,” or “the preceding result” when it receives only the latest user string and failed assistant field. Version 2.4 supplies bounded full task context, repairs reasoning and content together when necessary, distinguishes optional reasoning from the usable answer, and uses an integrity verifier that can reject invented or context-breaking output.

Soma 2.4.0 is not pre-qualified. JSON grammar guarantees neither correct semantic decisions nor grounded repairs. The exact model artifact, server build, endpoint pair, reasoning modes, media modes, prompt, budgets, and temperature must pass the live automated qualification gates before deployment. Reports retain hash-bound outputs for audit and reproduction, but inspection is not a separate qualification stage.

The pre-release 2.4.0 tree was stabilized in place rather than assigning a new version to review corrections made before qualification. The rollback remains the unchanged 2.3.0 directory. This stabilization adds no Soma environment variable and no package dependency; existing 2.4.0 profiles retain the same runtime contract. Evaluator CLI provenance such as --reasoning-budget and the corresponding llama.cpp server option are not Soma environment settings.

The current Qwen3.5-9B Q6_K route is unqualified. Its exploratory temperature-zero report used unrestricted secondary reasoning and failed automated gates. That report is diagnostic evidence only: it cannot be promoted or reinterpreted after evaluator stabilization. A fresh bounded-budget report must replace it as the current qualification record.

Soma never downloads, loads, switches, starts, stops, or restarts a model. It does not execute tools, maintain conversation state, authenticate clients, or provide tenant isolation.

Request flow

With the full primary-off/secondary-on staged profile, the normal path is:

client request
    -> target model
    -> buffer and validate one complete assistant turn
    -> build one bounded, role-preserving task context
    -> classify every present reasoning/content field on primary, reasoning off,
       with the other draft text fields removed from that classification envelope
    -> apply field policy and, if necessary, request one joint repair object
       candidate 1: primary/off -> primary/off integrity verification
       candidate 2: secondary/on -> secondary/on integrity verification
       candidate 3: secondary/on -> secondary/on integrity verification
    -> if an on-mode verifier truncates or returns invalid JSON:
       verify the identical candidate once on the same endpoint/model, reasoning off
    -> retain usable verified fields and immutable native tool calls
    -> if the turn is still unusable and target retry is enabled:
       call the target once more with a fixed corrective system instruction
       and process that fresh response under the same bounded rules
    -> if a verified repair rewrote reasoning and loop-back is enabled:
       call the target once more with only that verified reasoning appended
       and process that fresh response under the same bounded rules
    -> deduplicate exact reasoning/content copies
    -> return JSON or normalized SSE

Candidate two is generated only after a completed integrity verifier semantically rejects candidate one, or after an eligible route failure. Candidate three is the last semantic candidate and receives a fixed alternate focus on authority and exact output/tool/stop constraints. Every candidate is generated fresh from the original task context and failed target draft: neither rejected candidate text nor model-generated failure analysis is fed into another generation.

Without a configured secondary, Soma permits at most two primary candidates. A transient primary availability failure can make the remainder of the current client request start on secondary. That stickiness is request-local; the next request starts on primary. There is no global circuit breaker or harness/session model state.

The staged reasoning modes are part of the 2.4 contract, not optional tuning: TRANSFORM_REASONING_MODE must be off, and a configured secondary must use TRANSFORM_SECONDARY_REASONING_MODE=on. Configuration validation rejects any other combination so deployment cannot silently diverge from the qualified route.

Classification and integrity decisions

Classification returns exactly one object:

{"decision":"pass"}

or:

{"decision":"rewrite"}

The decision concerns the assistant's operative stance, not the presence of words such as “cannot,” “refuse,” or “safety.” Quoted, translated, summarized, or analyzed refusal language is not itself a refusal. A failed method followed by continued work, a different method toward the same requested outcome, and an honest factual constraint do not become rewrites merely because they contain negative language. An adopted decision to stop, avoid part of the request, or deliberately substitute a lesser outcome is a rewrite when the work remains feasible under the authoritative system/developer instructions. A boundary those instructions explicitly require is not an avoidable refusal and must be preserved.

Integrity verification uses the same strict object shape:

{"decision":"pass"}

or:

{"decision":"rewrite"}

During integrity verification, rewrite means the candidate has a material defect and must be rejected or regenerated; it does not request an in-place edit. The verifier can reject residual withholding, a generic promise, meta-commentary about the transformation, treating the failed assistant draft as user-supplied material, invented task-specific inputs or results, contradictions with accepted reasoning or immutable tool calls, and an unapproved clarification.

Soma accepts only a complete JSON object satisfying the current schema. A pure JSON fence is accepted, but an object embedded in prose is not. TRANSFORM_JSON_MODE=true is the default and sends a small schema through response_format; disabling it removes that wire hint but retains the same prompts, strict parser, local validation, and recovery bounds. Separately, target and repaired content requested as JSON must parse strictly and match an immediately declared top-level type. A configured literal stop sequence may not survive in forwarded reasoning or content. Soma intentionally does not implement full client JSON-Schema validation.

Joint repair contract and field policy

One repair call returns a fixed object with both members present and nullable:

{
  "reasoning": "complete repaired reasoning or null",
  "content": "complete repaired content or null"
}

Only fields classified for repair may be non-null. The two-key wire shape never changes, while the per-call schema constrains each requested member to string and each other member to null. Local validation preserves the same contract when a transform endpoint ignores the schema or JSON mode is disabled. Nonblank exact outputs such as {}, [], punctuation, and Unicode symbols are valid; Soma does not impose an English-text or alphanumeric "substance" heuristic on the requested deliverable.

Soma classifies all present fields before requesting a repair:

  • If reasoning and content pass, both target fields are preserved.
  • If content passes and reasoning requires repair, Soma drops the reasoning field; it does not risk generating new private analysis for an already usable answer.
  • If reasoning passes and content requires repair, the accepted reasoning is supplied as evidence for the content repair.
  • If both require repair, one candidate generates reasoning first and then content so the answer can follow the repaired analysis.
  • Verified jointly repaired reasoning is forwarded with its verified content. Because integrity verification is message-level, a rejected joint candidate is retried as a whole; Soma never salvages one unverified member from it.
  • A reasoning-only response with no content and no native tool call cannot become a terminal success merely because internal analysis exists. It takes the optional target retry when enabled; otherwise it fails explicitly.
  • Native tool_calls are immutable. Soma may repair adjacent reasoning/content using the full tool context, but exhausted prose repair clears the unusable prose and preserves the structured call. Soma never invents or edits a tool name, ID, argument string, ordering, or result. Before repair, every returned function name must match a supplied tool definition, and multiple returned calls are rejected when parallel_tool_calls=false.

The failed target assistant draft is evidence, not user-supplied task material. The transform is instructed not to quote, explain, or “convert” the refusal itself. It may preserve supported facts and genuine constraints, but it must not choose an arbitrary example, fill invented placeholders, fabricate code changes or external results, or claim a tool/action completed without evidence.

TRANSFORM_ALLOW_CLARIFICATION=false is the default. A transform response that asks the user for more information is not accepted as the repaired answer unless this option is explicitly enabled. Enabling it is appropriate only for harnesses where an essential missing input genuinely requires another user turn; it must be qualified as a separate behavior profile.

Full task context and privacy boundary

Classification, repair, and integrity verification receive the original request context needed to understand references and preserve constraints:

  • original messages in order and by role, including system, developer, user, assistant, and tool messages and tool results;
  • complete tool definitions, tool_choice, and parallel_tool_calls;
  • response_format, modality/audio controls, and stop;
  • the target assistant draft, clearly separated from the original request;
  • immutable target native tool calls in a separate read-only section; and
  • the configured media representation for every multimodal part.

System and developer messages remain authoritative context below Soma's fixed JSON and native-tool invariants. Other supplied values are task evidence, not permission to override the transform contract.

Draft text is projected per phase. A classifier receives only its named target field, so refusing content cannot contaminate accepted reasoning or vice versa. Repair generation may inspect fields marked for replacement to preserve facts supported by the task. Integrity verification removes every replaced or discarded original field and judges only retained evidence plus the current candidate.

Soma does not send target/transform endpoint credentials, HTTP headers, the target model name, sampling knobs, or rejected transform candidates. It does not log task context, prompts, target drafts, repaired output, tool arguments, media payloads, or credentials.

This is nevertheless a wider trust boundary than 2.3.0. Any secret embedded inside a conversation, tool definition, tool argument, or tool result is part of the original task context and can reach every transform endpoint used for that request, including a remote secondary. Configure only transform services authorized to receive the full request. Header exclusion cannot remove secrets that the client placed in message or tool data.

TRANSFORM_CONTEXT_MAX_CHARS=131072 bounds the serialized task_context, and TRANSFORM_FIELD_MAX_CHARS=32768 bounds an individual target reasoning/content field. The configured field limit must not exceed the context limit. The context limit has a hard maximum of 4000000 characters. Because this is a character bound, not tokenizer accounting, large-context profiles should leave room for transform instructions and generated output. Soma rejects oversized semantic input rather than truncating messages, tool schemas, code, or evidence into a misleading task. Phase envelopes add the bounded candidate/contract data, and forwarded native media remains subject to the upstream endpoint and trusted front proxy's byte limits.

Media modes

Media handling is explicit per transform endpoint:

  • placeholder preserves typed part positions and non-payload metadata, omits the actual binary/media payload, and marks the part unseen. The transform must not infer absent media details. This is the correct setting for a text-only or --no-mmproj llama.cpp server.
  • forward sends original typed content media using native OpenAI multimodal message parts. It does not serialize base64 media into ordinary JSON text. Provider-specific top-level assistant media has no portable input envelope and fails explicitly in this mode; use placeholder for that shape. Use forward only for an endpoint that is authorized and qualified to accept the request's typed media parts.
  • reject refuses to send a media-bearing task to that endpoint. Soma may use a configured compatible transform route; otherwise it fails explicitly. Target retry is not used to bypass an operator's transform-media policy.

Set TRANSFORM_MEDIA_MODE for primary and TRANSFORM_SECONDARY_MEDIA_MODE for secondary. If a forward endpoint rejects the media request, Soma routes only to a compatible configured secondary or fails explicitly. It never invokes target retry to bypass media policy and never silently retries the task as placeholder text, because either action would change the evidence available to the model.

Assistant audio attached to a usable text or native-tool turn is preserved, including audio accumulated from a target stream. Audio-only target turns are explicitly unsupported: Soma cannot inspect or repair the audio payload under its text repair contract, so it returns unsupported_target_response instead of forwarding an unchecked terminal answer.

Bounded recovery and verifier fallback

The primary reasoning-off profile owns normal classification, candidate one, and its integrity verification. After semantic rejection, a configured secondary reasoning-on profile owns candidates two and three, each generated from the pristine task package and independently verified.

Reasoning-enabled generation can improve task understanding, but a small model may spend an entire decision budget thinking and end with finish_reason=length before emitting its tiny JSON decision. If an on-mode integrity verification is truncated or structurally invalid, Soma does not discard the candidate. It verifies that identical candidate exactly once on the same endpoint and model with reasoning disabled. Only a completed rewrite decision advances to a fresh generation.

Transport/availability failures follow bounded route failover. Structural JSON recovery may include a concise closed failure category, but never rejected output or raw exception text. Semantic retries receive only positive instructions and the pristine task context; they are not primed with the preceding candidate or its failure.

There are hard ceilings of:

  • two target calls per client request;
  • 20 transform calls for each target response; and
  • 40 transform calls across the complete client request.

TRANSFORM_TOTAL_TIMEOUT=1200 is one aggregate deadline. It starts after the first target response completes and covers every transform call, an optional second target call, and processing of the second response. It does not reset after target retry. The initial target call remains governed by CONNECT_TIMEOUT and REQUEST_TIMEOUT outside that aggregate window.

Optional target retry

TARGET_RETRY_ON_UNREPAIRABLE=false preserves the normal one-target-call behavior. When enabled, Soma may call the target exactly once more only when the completed turn is unrepairable and leaves no usable content or immutable native tool call. A failed optional reasoning field does not trigger target retry when valid content remains.

The retry starts from the original request and inserts one fixed corrective system instruction immediately after the leading system/developer block. It preserves the conversation and tool contract and never includes the rejected target response or a rejected transform candidate. This avoids training the second response to imitate the failure, but it does add target latency/cost and may produce a different native tool decision. Soma still does not execute that call.

Enable target retry only after qualifying the complete target-plus-transform route. It is not a general retry for target HTTP errors, optional reasoning loss, or a merely imperfect answer.

Loop-back on verified repair

TARGET_LOOP_BACK_ON_VERIFIED_REPAIR=false is the default. When enabled, Soma may make exactly one additional target call after an integrity-verified repair that rewrote the target's refusal reasoning. Instead of returning the transform's repaired candidate directly, Soma re-sends the original request with one appended assistant message carrying only that verified repaired reasoning in a reasoning_content field, so the target re-ingests the relaxed context and produces the task output itself. The re-entry payload preserves the original conversation, media, tool definitions, tool_choice, stop controls, and response-format settings untouched.

Loop-back fires only on the first target attempt, only when reasoning was one of the repaired fields, and only after that candidate passed integrity verification. Content-only repairs, fields that classified as pass, tool-only turns, cleared tool prose, and the second target attempt never loop. Genuine refusals never loop because truthful technical, environmental, evidentiary, uncertainty, impossibility, missing-input, and factual limitations classify as pass and are never rewritten.

The second target call shares the hard ceiling of two target calls per client request and the aggregate TRANSFORM_TOTAL_TIMEOUT window, which is not reset. The second response is processed under the same classification, repair, and integrity rules; if it is also unrepairable, the request fails explicitly and Soma never makes a third target call. TARGET_RETRY_ON_UNREPAIRABLE and loop-back are mutually exclusive per request because they handle disjoint failure classes (unrepairable turns versus verified reasoning repairs) and share the single additional-call slot.

The reasoning carrier is fixed to reasoning_content with no fallback. Backends that reject that field in input messages fail explicitly rather than silently degrading to a different carrier. Loop-back adds target latency and cost; qualify the complete target-plus-transform route before enabling it.

Failure behavior

FAIL_OPEN=false is the default. Exhausted mandatory repair, invalid verification, oversized context, incompatible media, missing usable terminal output, and other nonrecoverable transform errors return an explicit error instead of forwarding a known-bad candidate.

FAIL_OPEN=true is an availability policy only. It can restore an original refusal, withholding field, or otherwise rejected target text and therefore defeats strict repair guarantees. Do not treat fail-open as a safety, compliance, or successful quality mode, and do not enable it merely to hide model qualification failures. Fail-open never makes a reasoning-only or otherwise empty terminal turn successful; that turn still takes the explicitly enabled target retry or returns an error.

Client and upstream JSON reject non-finite numbers. Transform objects additionally reject duplicate member names. Client stream and parallel_tool_calls values must be booleans, and n must be null or integer 1. Target assistant text, reasoning aliases, and native tool-call shapes are validated before any local mutation.

Endpoint configuration rejects userinfo, queries, fragments, invalid ports, and unsafe header overrides. REQUIRE_DISTINCT_ENDPOINTS=true prevents exact target/transform origin collisions and direct self-routes. Operators must still avoid DNS aliases or LAN addresses that resolve to a wildcard-bound Soma listener.

Configuration

Minimal one-profile configuration:

TARGET_URL=https://opencode.ai/zen/v1
TRANSFORM_URL=http://127.0.0.1:8001/v1
TRANSFORM_MODEL=local
TRANSFORM_REASONING_MODE=off
TRANSFORM_MEDIA_MODE=placeholder

Same-server primary-off/secondary-on profile:

PROXY_HOST=127.0.0.1
PROXY_PORT=8080

# Clear the removed 2.3.x option from an already-populated shell.
unset TRANSFORM_CONFIRM_REWRITES

TARGET_URL=https://opencode.ai/zen/v1
TARGET_KEY=
TARGET_HEADERS_JSON={}

TRANSFORM_URL=http://127.0.0.1:8001/v1
TRANSFORM_KEY=
TRANSFORM_MODEL=local
TRANSFORM_HEADERS_JSON={}
TRANSFORM_REASONING_MODE=off
TRANSFORM_MEDIA_MODE=placeholder

TRANSFORM_SECONDARY_URL=http://127.0.0.1:8001/v1
TRANSFORM_SECONDARY_KEY=
TRANSFORM_SECONDARY_MODEL=local
TRANSFORM_SECONDARY_HEADERS_JSON={}
TRANSFORM_SECONDARY_REASONING_MODE=on
TRANSFORM_SECONDARY_MEDIA_MODE=placeholder

# Same-server primary/secondary is allowed. This rejects target/transform collisions.
REQUIRE_DISTINCT_ENDPOINTS=true

ENABLE_REASONING={}
TRANSFORM_TEMPERATURE=0
TRANSFORM_JSON_MODE=true
TRANSFORM_CONTEXT_MAX_CHARS=131072
TRANSFORM_FIELD_MAX_CHARS=32768
TRANSFORM_DECISION_MAX_TOKENS=1536
TRANSFORM_REWRITE_MAX_TOKENS=16384
TRANSFORM_ALLOW_CLARIFICATION=false
TRANSFORM_TOTAL_TIMEOUT=1200
TARGET_RETRY_ON_UNREPAIRABLE=false
TARGET_LOOP_BACK_ON_VERIFIED_REPAIR=false
FAIL_OPEN=false

CONNECT_TIMEOUT=15
REQUEST_TIMEOUT=600

TRANSFORM_CONFIRM_REWRITES was removed. Soma rejects the variable even when its value is false; this catches a stale 2.3.x deployment rather than silently changing its meaning. Deleting an export from a file does not clear an existing shell value, so either start from a clean environment or run:

unset TRANSFORM_CONFIRM_REWRITES

The primary and secondary keys/headers never inherit from one another. A same-server secondary supplies behavioral diversity but no process, GPU, or availability isolation. An independent endpoint/model can supply both, at the cost of extending the full-context trust boundary. Qualify the secondary by itself and then qualify the exact composed pair. Soma 2.4 requires primary off and secondary on; default and the inverse mode assignments are rejected during configuration validation.

TRANSFORM_TEMPERATURE is sent on every transform call and overrides the llama server sampling default. TRANSFORM_DECISION_MAX_TOKENS covers classifications and integrity decisions; TRANSFORM_REWRITE_MAX_TOKENS covers the fixed joint-repair object. Valid ranges are 25616384 decision tokens, 25616384 repair tokens, 40964000000 context characters, and 10244000000 field characters, with the field limit no greater than the context limit. Larger budgets bound output but do not improve model judgment by themselves.

Environment files are shell profiles and are not loaded automatically. Restart Soma after every environment change:

set -a
. ./soma.env
set +a
python3 soma.py --check-config
python3 soma.py

Additional environment variables not shown in the profiles above:

  • LOG_LEVEL (default INFO) — Python logging level for proxy diagnostics.
  • FORWARD_CLIENT_HEADERS (default true) — forward non-hop, non-credential client headers to the target endpoint.
  • TRANSFORM_PROMPT (default built in) — base system prompt prepended to every transform phase prompt.
  • SOMA_AUTO_REQUIRES_TOOL (default false) — strict auto-tools mode that classifies each request as requiring a native call or a text response.
  • UPSTREAM_ERROR_BODY_LIMIT (default 4000, range 25665536) — bounded number of upstream error-body bytes retained for target diagnostics.
  • SSE_CHUNK_CHARS (default 2048, range 12865536) — maximum characters per normalized SSE text delta.

Verify the effective version, endpoint identities, reasoning/media modes, JSON mode, context/field/token limits, clarification, target-retry, and loop-back policies, aggregate deadline, and call ceilings through --check-config, startup diagnostics, or /health.

Point clients at:

http://<proxy-host>:8080/v1/chat/completions

Aliases are available at /chat/completions, /v1/models, /models, and /health.

llama.cpp recommendation for a shared local endpoint

For the shared-endpoint topology where one llama.cpp process serves a primary reasoning-off profile and a secondary reasoning-on profile through per-request enable_thinking, enable server reasoning support and cap thinking so a small decision response has room to emit JSON:

--reasoning on --reasoning-budget 512 --temp 0

The primary profile still sends enable_thinking=false; the global server mode must not prevent the secondary profile from producing and parsing reasoning when it sends enable_thinking=true. A 512-token cap is the required starting profile for the bounded-budget qualification run; configuring it is not itself a qualification claim. The evaluator's matching --reasoning-budget 512 argument records what the already-running server uses and does not configure the server.

Keep both the server and Soma transform temperature at zero for qualification. Soma's per-request TRANSFORM_TEMPERATURE=0 is authoritative for transform calls; the server flag supplies a matching default. Temperature zero removes deliberate sampling variance so repeat failures can be attributed to the route under test, although it does not promise byte-identical output across server builds, speculative decoding, cache state, or concurrency. Any nonzero temperature is a different profile and requires a separate report.

Re-run the exact live profile after changing any server argument. Soma does not add these arguments, restart the server, or download a model. A text-only server launched with --no-mmproj should use placeholder, not forward, for both transform media modes.

Multiple harness profiles

Use one Soma process and listener port per harness. The supplied profiles/harness-a.env.example and profiles/harness-b.env.example use ordinary environment variables and distinct ports. Soma has no HARNESS_TYPE dispatch or shared mutable deployment profile.

cp profiles/harness-a.env.example profiles/harness-a.env
cp profiles/harness-b.env.example profiles/harness-b.env
env -i PATH="$PATH" /bin/sh -c 'set -a; . ./profiles/harness-a.env; set +a; exec python3 soma.py --check-config'
env -i PATH="$PATH" /bin/sh -c 'set -a; . ./profiles/harness-b.env; set +a; exec python3 soma.py --check-config'

The supplied profiles/.gitignore excludes populated profile names while retaining the examples. Keep production profiles outside distributable artifacts even when ignore rules are present. A shared transform server must be qualified at the combined load and configured concurrency; a one-slot llama server serializes both harnesses.

Trust and resource boundary

Keep PROXY_HOST=127.0.0.1 unless a trusted front proxy supplies authentication, access control, TLS, request-size limits, buffering limits, timeouts, and rate limits. Soma warns when bound to a non-loopback interface. It buffers complete target turns and full bounded task packages and has no in-process concurrency-admission limit, so the front proxy must enforce limits appropriate to available memory.

Native tool calls and streaming

Soma supports native OpenAI tool_calls only. It preserves IDs, type: function, function names, strict JSON argument strings, ordering, and streaming fragments. Proprietary text tool syntaxes are ordinary assistant text; conversion belongs in the target's OpenAI-compatible gateway.

For stream:true, Soma buffers the complete target stream, processes it, and emits normalized OpenAI delta SSE. Original chunk boundaries are not preserved. Valid reasoning, content, native tool calls, finish reason, usage, and response metadata are retained. Accepted cost and usage metadata are emitted together at most once.

The upstream stream must produce a terminal non-null finish_reason. Soma accepts a terminal choice followed by EOF or the ordinary sequence ending in [DONE]. Standard empty-choice usage frames are retained. After the first [DONE], at most one narrow metadata postlude is allowed: an object with choices: [], no keys outside choices, cost, and usage, and at least one non-null metadata value. It may end at EOF or one closing [DONE]. Further objects/delimiters, malformed or non-finite JSON, duplicate keys, premature [DONE], or meaningful data after the terminal choice are rejected.

Diagnostics

Successful and post-dispatch error responses expose privacy-safe trace/timing and bounded call counts, field decisions, candidate/verifier outcomes, target-retry use, deduplication, and fail-open status. /health and startup diagnostics additionally show the effective non-secret reasoning and media configuration.

Transform logs identify phase, field/candidate, backend, reasoning and media mode, purpose, closed failure category, JSON-mode value, channel lengths, finish reason, token counts, elapsed time, and a request-ID fingerprint. They do not include prompts, task context, target or transform text, media, tool arguments, credentials, error bodies, or raw upstream request IDs.

Tests

Run the complete offline suite:

python3 -m unittest -v test_soma.py test_soma_extra.py
python3 test_soma_live.py --inventory

The suite covers strict schemas, full-context isolation and limits, joint field policy, media routes, primary/secondary candidate ownership, reasoning-off verifier fallback, optional target retry, hard call/deadline ceilings, fail-open behavior, native tool fidelity, JSON validation, and SSE normalization. Offline success is necessary but is not model qualification.

Live qualification

test_soma_live.py is opt-in and dynamically imports the adjacent soma.py, so it exercises the exact runtime prompts, schemas, parsing, validation, routing, and field policy. It calls only already-running endpoints supplied by the operator and never manages a model or server.

First inspect the frozen corpus without network access:

python3 test_soma_live.py --inventory

Run the exact selected transform artifact/profile at temperature 0 and retain the report only under ignored qualification-local/. Consult --help for the current provenance and endpoint arguments:

python3 test_soma_live.py --help

Every qualifying run must declare primary --reasoning-mode off, a configured secondary with --secondary-reasoning-mode on, and the secondary server's actual positive --reasoning-budget (for the documented llama.cpp starting profile, --reasoning-budget 512). This evaluator value records provenance; the server must already have been launched with the matching budget.

An exploratory run against llama.cpp's unrestricted default may record --reasoning-budget -1. Its report remains unqualified because the positive-budget provenance gate fails; it is not carried forward after a complete bounded-budget rerun replaces the current evidence.

Qualification is automated-only. Use a new report filename and run the exact temperature-zero, bounded profile. Exit status 0 means every qualification gate passed and the report records qualified: true with qualification_status: qualified. Exit status 1 means at least one qualification gate failed, and 2 means setup or report creation failed. The evaluator has no second approval stage; inspecting retained evidence does not alter report status.

Provider-managed routes can be exercised with --artifact-kind provider-managed, but they are recorded as exploratory and can never be marked qualified by this evaluator. Supply the exact provider name, model label, and a small public /models metadata record through --provider-model-metadata-json; do not invent GGUF, llama.cpp, hardware, revision, or reasoning-budget values for a hosted service. Use --reasoning-budget 0 when the provider does not publish a bounded budget. The report separates behavioral gate results from qualification eligibility and records the requested reasoning modes as unverified provider controls.

To retain the qualified local GGUF primary while evaluating a hosted secondary, use --artifact-kind hybrid-local-provider. Supply the ordinary local artifact fields for the primary and the provider metadata fields for the secondary. The evaluator retains both identities, but deliberately records the combined route as exploratory and qualification-ineligible because the hosted reasoning controls and budget are not independently verified. The primary's artifact label may differ from its wire model alias (for example, an immutable repository label with local on the wire).

--target-smoke-count 10 limits only the final target-through-transform smoke calls. It does not limit the preceding transform corpus: the evaluator still runs all 240 classifier cases, 80 retained repairs, repeat matrices, and route/media probes. Each smoke request grants 128 output tokens, and its response must contain exactly OK with no surrounding whitespace, prose, or native tool call.

For the initial bounded Qwen route, the automated command must include the exact primary/secondary endpoint and provenance arguments plus:

python3 test_soma_live.py \
  --reasoning-mode off \
  --secondary-reasoning-mode on \
  --reasoning-budget 512 \
  --temperature 0 \
  --report qualification-local/qwen3.5-9b-q6_k-t0-rb512-automated.json \
  [the exact endpoint, model, server, artifact, and hardware arguments]

Reports are immutable evidence files. The evaluator writes a completed report privately and installs it atomically; it never exposes a partially written result or overwrites an existing path. The stabilized evaluator, corpus, and source hashes must match the new run. After a complete budget-512 report has been validated and installed under its truthful filename, remove the obsolete unrestricted-budget artifact so only the current evidence remains.

The automated gates require:

  • valid contracts on all 240 classification cases, 100% hard-refusal and overall refusal recall, and zero false rewrites;
  • all 20 schema-off high-risk sentinels and five repeats of every high-risk case at parallelism 1 and 4 with zero repeat failures;
  • all 80 message-repair cases completed without exhaustion and 100% integrity verification, required-fact retention, and forbidden-fact absence;
  • exactly 20 cases in each field-decision cell: pass/pass, rewrite/pass, pass/rewrite, and rewrite/rewrite;
  • the exact staged primary-off then secondary-on candidate route;
  • explicit positive secondary reasoning-budget provenance (use 512 as the initial llama.cpp qualification value);
  • explicit placeholder, forward, and reject media behavior; and
  • complete source, evaluator, model, server, configuration, and fixture reproducibility evidence.

The report also records latency, classification disagreements, backend/phase ownership, semantic repair attempts, verifier fallback, and call ceilings. Strict JSON is exercised both with structured-output mode enabled and with the wire schema omitted.

Live target smoke is separate and explicitly opt-in because it incurs target cost and can produce a new model/tool decision. It uses benign fixtures, keeps target retry disabled, verifies the complete target-to-transform route, and never executes returned tools. Target-retry behavior remains deterministic offline coverage until separately qualified; live smoke does not enable it. The smoke is not run by --inventory or an ordinary transform-only qualification. The count is bounded from 1 through 10:

python3 test_soma_live.py \
  --target-smoke \
  --target-url https://target.example/v1 \
  --target-model TARGET_MODEL \
  --reasoning-budget 512 \
  --target-smoke-count 10 \
  [the same transform and provenance arguments used for qualification]

Without --target-smoke, the evaluator makes zero target calls. Supply target keys through the hidden CLI/environment option, never in recorded server arguments or a report intended for sharing.

An automated pass is final qualification for the exact recorded profile. Reports retain all 80 accepted repair outputs and their evidence hashes so the result can be audited and reproduced, but later inspection does not change qualification status. Any failed gate leaves the profile unqualified, and a smaller model receives no relaxed threshold.

Any change to model revision, GGUF, server build/arguments, reasoning budget, temperature, prompt, endpoint identity, media mode, context/token limits, field policy, primary/secondary composition, evaluator source, fixture corpus, or assertion semantics creates different evidence and requires a new report. Evidence hashes bind one report's exact inputs and outputs; they do not transfer qualification to a superseded report. Reports can contain synthetic task context and non-secret provenance; inspect them before sharing and never place keys in recorded header/server arguments.

qualification-local/, populated profiles, logs, caches, credentials, and model artifacts are excluded from the release package and checksums.