parse_decision is the pure validation boundary (no network), so it — and the
fakes that stand in for a real client in tests — are trivially testable.
Classes
AskUser
Decision to ask the user a question (blocks in a non-interactive run) (§6).
DecisionParseError
Malformed model output — recoverable; fed back to the model (§6), never fatal.
Carries usage when the client exhausts its in-client retries, so a lost turn is
still billed — the expensive failure mode is exactly the one that must not be
undercounted (PR-#31 review).
Args:
message: The parse-failure description fed back to the model.
usage: Tokens spent across the failed attempts, or None if unreported.
DecisionRetryNote
One malformed in-client attempt: what was wrong, and a capped raw excerpt.
DecisionUsage
Provider-reported token usage for one decision (all in-client attempts summed).
EmptyResponseError
The provider returned an empty / whitespace-only / all-NUL body (a 200 with no content).
The OpenAI SDK does not retry this (it is a successful HTTP response), so it must be caught
explicitly and treated as a transport failure — not routed into the model parse-retry, which
would re-prompt the model for what is really a dead/stalled provider reply (ADR-0028 R2).
FinalAnswer
Decision claiming the task is complete — a proposal for the verifier (§6, §12).
ModelClient
Anything that turns a context packet into a validated decision (§6).
The real implementation calls an OpenAI-compatible endpoint and runs the
result through parse_decision; tests substitute a scripted fake.
ModelClient.adecide(self, context: ContextPacket) -> ModelDecision
Async entry point for one decision (ADR-0029 R5); defaults to offloading sync decide.
OpenAIModelClient overrides this with a cancellable streaming path; every fake inherits
this bridge unchanged — its sync decide runs in a worker thread (uncancellable mid-call,
but fast and deterministic, so the runner’s cancel-race still resolves promptly).
Args:
context: The assembled context packet.
Returns:
The validated decision for the current turn.
ModelClient.decide(self, context: ContextPacket) -> ModelDecision
Turn a context packet into a validated decision for the current turn.
Args:
context: The assembled context packet.
Returns:
The validated decision for the current turn.
ModelDecision
One validated model decision: a thought plus exactly one action (§6).
retry_trace is a harness-owned diagnostics channel: the model client annotates
the decision with any malformed attempts it recovered from in-client, so the runner
can record them as evidence and journal them (invariant #5). It is never accepted
from raw model output — parse_decision clears it.
OpenAIModelClient
Calls an OpenAI-compatible endpoint and validates the reply (§6, §18).
A malformed reply is fed back to the model for a bounded number of retries
before surfacing as a DecisionParseError (which the runner treats as a
recoverable, model-correctable error).
Args:
config: The harness configuration.
client: An injected OpenAI-compatible client, or None to build one lazily on
first use — so construction needs no credentials; the optional openai
extra and an API key are required only when decide() is first called.
max_parse_retries: Number of retries on malformed model output.
transport_max_retries: Transport-layer retries on a NUL/empty body or a request
failure (ADR-0028 R3); None takes config.transport_max_retries.
sleep: Backoff sleeper, injectable so tests exercise retries without real delay.
aclient: An injected async OpenAI-compatible client (the ADR-0029 R5 streaming path),
or None to build an AsyncOpenAI lazily on first adecide().
asleep: Async backoff sleeper, injectable so async tests skip real delay.
OpenAIModelClient.adecide(self, context: ContextPacket) -> ModelDecision
Call the endpoint asynchronously, streaming by default for idle-timeout + cancellation.
Streams native tool-calls when enabled (ADR-0029 R5) so a stall is caught at the idle
timeout regardless of generation length; a provider that can’t stream trips
_streaming_unsupported and the SAME request is re-issued non-streaming for the rest of the
session (D4). Either path is cancellable mid-call (the runner races a cancel against it) and
raises TransportError/DecisionParseError exactly like the sync decide.
Args:
context: The assembled context packet.
Returns:
The validated decision for the current turn.
OpenAIModelClient.decide(self, context: ContextPacket) -> ModelDecision
Call the endpoint and validate the reply, retrying on malformed output (§6).
The default transport is native provider tool-calling (ADR-0003 A) — the
provider owns the call envelope, so a large patch can’t die in hand-escaping;
config.native_tool_calls=False restores the legacy single-JSON-object protocol.
Either path raises DecisionParseError when every attempt is malformed.
Args:
context: The assembled context packet.
Returns:
The validated decision for the current turn.
StreamingUnsupportedError
The provider can’t stream tool-calls — trip the per-instance flag and fall back (ADR-0029 D4).
Deliberately NOT a TransportError: re-issuing the same streaming request would fail
identically, so it must not flow into the transport-retry. The client catches it once, flips
_streaming_unsupported, and re-issues the SAME request non-streaming for the rest of the
session. A capability verdict (a streaming-rejection 4xx, or unusable tool-call framing), never
a transient fault — when in doubt the discrimination defaults to TransportError (ADR-0029 D).
ToolCall
Decision to invoke a named tool with validated input (§6).
TransportError
A model call failed at the transport layer — NOT model-correctable (§16, ADR-0028).
A request timeout, connection reset, or an empty/NUL body (EmptyResponseError) means the
provider returned nothing usable. Unlike DecisionParseError, this is never fed back to
the model: it is retried in-client at the transport layer (re-issue the same request with
backoff), and on exhaustion surfaced to the runner as a system failure. Carries usage so
the billed-but-lost attempts are not undercounted.
Args:
message: A short description of the transport failure.
usage: Tokens spent across the failed attempts, or None if unreported.
Functions
build_messages(context: ContextPacket, *, native_tools: bool = False) -> list[dict[str, str]]
Assemble the system + user messages for one decision (§9 packet → prompt).
Args:
context: The assembled context packet.
native_tools: True for the native tool-calling transport (ADR-0003 A) — the
provider carries the tool schemas, so the prompt drops the JSON-envelope
contract and the prose tool list; False keeps the legacy protocol verbatim.
Returns:
The system + user messages for one decision.
build_tool_schemas(context: ContextPacket) -> list[dict]
The function schemas for one decision: the advertised tools + the decision actions.
Each phase-admitted tool rides its real pydantic input_schema; final_answer and
ask_user become functions too, so every §6 decision shape is a structured call the
provider validates — never a hand-escaped JSON envelope (ADR-0003 A).
Args:
context: The assembled context packet (its allowed_tools are advertised).
Returns:
OpenAI-style tools= entries.
parse_decision(raw: str) -> ModelDecision
Validate raw model output into a ModelDecision, or raise a recoverable error.
Args:
raw: The raw model output to validate.
Returns:
The validated ModelDecision.
Raises:
DecisionParseError: If raw is not valid JSON or not a valid decision.