Skip to content

LLM Step

LLMStepDescription is the default step type: a chain-of-thought reasoning step backed by an LLM call.

FieldTypeDefaultPurpose
aimstr"" (required, non-empty)The primary objective of the step.
reasoning_questionsstr""Key questions the step should answer.
stage_actionstr""The specific action to perform.
example_reasoningstr""An example of the expert reasoning you want.
step_context_querieslist[ContextQuery | str][]RAG-like queries that pull relevant context from outer_context into this step’s prompt.
llm_configLLMStepConfig | NoneNonePer-step model / temperature / mode override.
retry_maxint | NoneNoneRetry attempts for this step (None = context default).
timeoutfloat | NoneNonePer-step timeout in seconds.

Plus all the shared fields.

from mmar_carl import LLMStepDescription
LLMStepDescription(
number=1,
title="Assess data quality",
aim="Assess the quality and completeness of the input data.",
reasoning_questions="What data patterns and anomalies are present?",
step_context_queries=["missing values", "data consistency"],
stage_action="Evaluate reliability and flag issues.",
example_reasoning="High-quality data enables reliable downstream analysis.",
)

Override the model, temperature, or other settings for one step with LLMStepConfig:

FieldTypeDefaultPurpose
modelstr | NoneNoneModel id (e.g. anthropic/claude-3.5-sonnet).
temperaturefloat | NoneNone0.0–2.0.
max_tokensint | NoneNoneOutput cap.
timeoutfloat | NoneNonePer-step timeout.
token_budget_warningint | NoneNoneWarn when the step exceeds this many tokens.
execution_modeExecutionModeFASTFAST or SELF_CRITIC.
use_message_historyboolFalseSend structured multi-turn messages instead of a flat prompt.
from mmar_carl import LLMStepDescription, LLMStepConfig
LLMStepDescription(
number=2,
title="Deep analysis",
aim="Perform a rigorous analysis.",
llm_config=LLMStepConfig(model="anthropic/claude-3.5-sonnet", temperature=0.2),
)

Two extras worth calling out:

  • token_budget_warning — emits a warning when the step’s total tokens exceed the threshold. It only fires with a client that reports usage (e.g. OpenAICompatibleClient); see the actuals in cost estimation.
  • use_message_history — sends a structured system/user/assistant message list instead of a flat prompt: the system prompt + outer context become the first system message, prior turns from context.messages are included, and the step’s reply is appended back to context.messages. Requires a client that implements get_response_with_messages (e.g. OpenAICompatibleClient).
  • FAST (default) — a single LLM pass.
  • SELF_CRITIC — the step generates an answer, runs it past one or more evaluators, and regenerates up to self_critic_max_revisions times until the evaluators approve.
from mmar_carl import ExecutionMode
LLMStepConfig(
execution_mode=ExecutionMode.SELF_CRITIC,
self_critic_evaluators=["llm"],
self_critic_max_revisions=2,
self_critic_instruction="Reject answers that aren't backed by the data.",
)

self_critic_evaluators names evaluators in order; all must approve. "llm" is the built-in LLM reviewer. Register your own (non-LLM or custom) evaluator with context.register_self_critic_evaluator(name, evaluator) — subclass SelfCriticEvaluatorBase and return a SelfCriticDecision(verdict=..., review_text=...).

The repo’s execution-modes example (mock client, no API key) runs a 3-step chain: a FAST pass, a SELF_CRITIC step with the default "llm" evaluator, and a SELF_CRITIC step whose evaluator chain adds a custom keyword guard.

class KeywordGuardEvaluator(SelfCriticEvaluatorBase):
def __init__(self, required_keyword: str):
self.required_keyword = required_keyword.lower()
async def evaluate(self, step, candidate, base_prompt, context, llm_client, retries):
if self.required_keyword in candidate.lower():
return SelfCriticDecision(verdict="APPROVE", review_text="found")
return SelfCriticDecision(verdict="DISAPPROVE", review_text="missing keyword")
context.register_self_critic_evaluator("keyword_guard", KeywordGuardEvaluator("mitigation"))
LLMStepConfig(
execution_mode=ExecutionMode.SELF_CRITIC,
self_critic_evaluators=["llm", "keyword_guard"], # both must approve
self_critic_max_revisions=2,
)

Per-step mode telemetry (execution_mode, llm_calls, revision rounds, quality_warning) lands in context.metadata["execution_mode_details"].