Skip to content

Reflection

After a run, chain.reflect(...) asks an LLM to analyse how well the chain accomplished the task — surfacing weak steps and concrete suggestions.

result = chain.execute(context)
reflection = chain.reflect(
task_description="Analyze customer sentiment and extract key themes",
)
print(reflection)
ParameterTypeDefaultPurpose
task_descriptionstr— (required)The original goal to judge against.
contextReasoningContext | NoneNoneDefaults to the last execution context.
languageLanguage | NoneNoneDefaults to the context language.
optionsReflectionOptions | NoneNoneControls what goes into the prompt.

It raises RuntimeError if no run has happened yet.

ReflectionOptions tunes the reflection prompt:

from mmar_carl import ReflectionOptions
chain.reflect(
task_description="...",
options=ReflectionOptions(include_dependency_analysis=False),
)

When include_metric_scores=True (default), any metric scores are fed into the prompt as concrete quality signals — e.g. a low keyword_coverage on a step nudges the model to suggest better step_context_queries.

Pair reflection with DatasetEvaluator: a SelectionStrategy (ThresholdStrategy / TopKWorstStrategy) picks the worst-scoring cases so you reflect on the failures that matter.