LLM Step
LLMStepDescription is the default step type: a chain-of-thought reasoning step
backed by an LLM call.
Fields
Section titled “Fields”| Field | Type | Default | Purpose |
|---|---|---|---|
aim | str | "" (required, non-empty) | The primary objective of the step. |
reasoning_questions | str | "" | Key questions the step should answer. |
stage_action | str | "" | The specific action to perform. |
example_reasoning | str | "" | An example of the expert reasoning you want. |
step_context_queries | list[ContextQuery | str] | [] | RAG-like queries that pull relevant context from outer_context into this step’s prompt. |
llm_config | LLMStepConfig | None | None | Per-step model / temperature / mode override. |
retry_max | int | None | None | Retry attempts for this step (None = context default). |
timeout | float | None | None | Per-step timeout in seconds. |
Plus all the shared fields.
Basic example
Section titled “Basic example”from mmar_carl import LLMStepDescription
LLMStepDescription( number=1, title="Assess data quality", aim="Assess the quality and completeness of the input data.", reasoning_questions="What data patterns and anomalies are present?", step_context_queries=["missing values", "data consistency"], stage_action="Evaluate reliability and flag issues.", example_reasoning="High-quality data enables reliable downstream analysis.",)Per-step LLM config
Section titled “Per-step LLM config”Override the model, temperature, or other settings for one step with
LLMStepConfig:
| Field | Type | Default | Purpose |
|---|---|---|---|
model | str | None | None | Model id (e.g. anthropic/claude-3.5-sonnet). |
temperature | float | None | None | 0.0–2.0. |
max_tokens | int | None | None | Output cap. |
timeout | float | None | None | Per-step timeout. |
token_budget_warning | int | None | None | Warn when the step exceeds this many tokens. |
execution_mode | ExecutionMode | FAST | FAST or SELF_CRITIC. |
use_message_history | bool | False | Send structured multi-turn messages instead of a flat prompt. |
from mmar_carl import LLMStepDescription, LLMStepConfig
LLMStepDescription( number=2, title="Deep analysis", aim="Perform a rigorous analysis.", llm_config=LLMStepConfig(model="anthropic/claude-3.5-sonnet", temperature=0.2),)Two extras worth calling out:
token_budget_warning— emits a warning when the step’s total tokens exceed the threshold. It only fires with a client that reports usage (e.g.OpenAICompatibleClient); see the actuals in cost estimation.use_message_history— sends a structuredsystem/user/assistantmessage list instead of a flat prompt: the system prompt + outer context become the firstsystemmessage, prior turns fromcontext.messagesare included, and the step’s reply is appended back tocontext.messages. Requires a client that implementsget_response_with_messages(e.g.OpenAICompatibleClient).
Execution modes
Section titled “Execution modes”FAST(default) — a single LLM pass.SELF_CRITIC— the step generates an answer, runs it past one or more evaluators, and regenerates up toself_critic_max_revisionstimes until the evaluators approve.
from mmar_carl import ExecutionMode
LLMStepConfig( execution_mode=ExecutionMode.SELF_CRITIC, self_critic_evaluators=["llm"], self_critic_max_revisions=2, self_critic_instruction="Reject answers that aren't backed by the data.",)self_critic_evaluators names evaluators in order; all must approve. "llm"
is the built-in LLM reviewer. Register your own (non-LLM or custom) evaluator with
context.register_self_critic_evaluator(name, evaluator) — subclass
SelfCriticEvaluatorBase and return a SelfCriticDecision(verdict=..., review_text=...).
Example
Section titled “Example”The repo’s execution-modes example
(mock client, no API key) runs a 3-step chain: a FAST pass, a SELF_CRITIC step
with the default "llm" evaluator, and a SELF_CRITIC step whose evaluator chain
adds a custom keyword guard.
class KeywordGuardEvaluator(SelfCriticEvaluatorBase): def __init__(self, required_keyword: str): self.required_keyword = required_keyword.lower()
async def evaluate(self, step, candidate, base_prompt, context, llm_client, retries): if self.required_keyword in candidate.lower(): return SelfCriticDecision(verdict="APPROVE", review_text="found") return SelfCriticDecision(verdict="DISAPPROVE", review_text="missing keyword")
context.register_self_critic_evaluator("keyword_guard", KeywordGuardEvaluator("mitigation"))
LLMStepConfig( execution_mode=ExecutionMode.SELF_CRITIC, self_critic_evaluators=["llm", "keyword_guard"], # both must approve self_critic_max_revisions=2,)Per-step mode telemetry (execution_mode, llm_calls, revision rounds,
quality_warning) lands in context.metadata["execution_mode_details"].
See also
Section titled “See also”- Dynamic references for wiring step outputs together.
- Context extraction for how
step_context_querieswork.