Skip to content

Cost Estimation

chain.estimate_cost(...) does a dry run: it walks the chain and projects token usage and cost per LLM-calling step — without making a single API call.

estimate = chain.estimate_cost(
context,
pricing={"qwen/qwen3-8b": (0.00002, 0.00006)}, # {model: (input_per_1k, output_per_1k)}
default_output_tokens=512,
)
print(estimate.format_table())
ParameterTypeDefaultPurpose
contextReasoningContext—Provides the input the estimate is sized against.
pricingdict[str, tuple[float, float]] | NoneNonePer-model (input_per_1k_usd, output_per_1k_usd).
default_output_tokensint512Assumed output length per step.
char_per_tokenint4Heuristic for input token counting.

It returns a CostEstimate (with a StepCostEstimate per step). Models missing from pricing are reported so you know what’s uncounted.

print(estimate.format_table()) # per-step token / USD table

In Jupyter, type estimate — _repr_markdown_ renders a banner + table.

To project the spend of a whole evolution (smoke + population × generations × cases), use evolver.estimate_cost(context_factory, pricing=...) instead — it multiplies a single chain estimate by the run size.

The repo’s token-usage example (mock client, no API key) walks three granularities: a pre-flight estimate, the actual per-step token usage after a run, and aggregation across a batch.

# 1. Pre-flight — no LLM calls made.
estimate = chain.estimate_cost(ctx, pricing=PRICING)
print(estimate.format_table())
# 2. Actuals — after execution, read per-step + chain-total usage.
result = await chain.execute_async(ctx)
for sr in result.step_results:
print(sr.step_number, sr.token_usage) # {"prompt", "completion", "total"} per step
print(result.token_usage) # chain totals
print(result.get_profiling_summary()) # peak/history bytes + total time

Actual usage requires a client that reports it (e.g. via get_response_with_usage); chain totals are also available as the result.token_usage_by_step property. Set token_budget_warning on a step’s LLMStepConfig to warn when a step exceeds a token budget.

  • Visualization — format_cost_by_model, profiling tables.
  • Tracing — real token usage after a run.