Overview
The mental model behind Kitaru — run, replay, improve.
Kitaru is the runtime for production AI agents. It records every run as durable checkpoints, lets you replay a real run with one thing changed, and helps you roll the winning change across recent runs. The loop is run → replay → improve.
A Kitaru flow is a dynamic ZenML pipeline and a checkpoint is like a step, so agents run on the same stacks, server, and dashboard as your ZenML pipelines.
Run (durable). Every model call and tool call is recorded as a checkpoint. This is the enabler for everything below, not the headline.
Replay (the differentiator). Re-execute a real run from a checkpoint. A rerun with no change reproduces the original — that faithful baseline is your control. Replay again with one input changed (a different model, a different prompt) and diff the two. Because the baseline reproduced, the difference is your change, not replay noise. This re-executes the real run; it is not re-scoring outputs like an eval.
Improve. Apply the same change across a cohort of recent runs, measure cost, latency, and quality, and keep the winner.
Durable execution is the mechanism that makes replay faithful. Start with Harness, Runtime, Platform for where Kitaru fits in an agent stack, or How It Works for the three-planes model (control / orchestration / execution) and what runs where in local dev vs production.
Core ideas
Flow
The outer durable boundary around your workflow
Checkpoint
A unit of work inside a flow whose output is persisted
Execution
A single run of a flow, identified by a unique ID
Structured metadata
Key-value data you attach to executions and checkpoints with kitaru.log()
Runtime log storage
Where runtime logs are sent (configured separately from structured metadata)
Active stack
The default execution target used when no per-run stack=... override is passed
What you can use today
Kitaru's current release includes:
@flow— mark a function as a durable workflow@checkpoint— mark a function as a persisted work unitflow.run(...).wait()— run a flow to completion; the handle carries.exec_idflow.replay(exec_id, at="<checkpoint>", flow_overrides={...})— re-execute a recorded run from a checkpoint, optionally overriding flow inputs such asmodelorprompt_profilekitaru.log()— attach structured metadata to the current scopekitaru.wait()— pause a flow until external input is suppliedkitaru.llm()— make tracked model calls with prompt/response capturekitaru.connect()— connect to a Kitaru serverkitaru.configure()— set process-local runtime defaultskitaru.save()/kitaru.load()— persist and load named artifacts in checkpointskitaru.list_stacks()/kitaru.current_stack()/kitaru.use_stack()— manage the default stackKitaruClient— inspect executions, fetch logs, resolve waits, retry, replay, and browse artifactsFlowHandle— interact with a running or finished execution
Replay and diff are also exposed over an MCP server and the kitaru CLI (kitaru executions replay <id> --at <checkpoint> --flow-overrides <json>), so a coding agent can drive the run → replay → improve loop directly.
All of the primitives listed here ship today. Some capabilities are backend-dependent — runtime log retrieval, for example, requires a server-backed connection — but they are part of the supported Kitaru surface.
Explore the concepts
Last updated
Was this helpful?