Complete returns agent tutorial
Investigate supplied returns agent traces and test one evidence-led improvement.
This tutorial applies Kitaru's complete method to a small customer-support agent. The agent looks up orders, return policies, and shipments, then chooses whether to refund, replace, or escalate each request.
You begin with ten recorded PydanticAI sessions exported from Langfuse. The walkthrough does not reveal or use the example's test-only expected outcomes. You will survey the population, inspect complete traces, record your own judgments, define one observable behavior, freeze its reviewed evidence, and test one bounded agent change.
The tutorial is intentionally more detailed than the Quickstart. It explains what each resource preserves and why each command is part of the evidence chain. Your exact sessions, questions, evaluator, candidate, and result will depend on what you observe.
Meet the example agent
Each session starts with one synthetic customer ticket. The agent looks up the order, gathers the relevant policy or shipping evidence, chooses one terminal outcome, and returns a structured resolution with a customer reply.

For return and refund requests, get_return_policy supplies the rules for the order's product category. Final sale means the item is not eligible for an ordinary return. A reported defect can still qualify when the category has a final-sale defect exception.
Footwear
30 days
Yes
$150
Apparel
30 days
Yes
$150
Accessories
14 days
No
$100
Luggage
45 days
Yes
$200
These values are evidence available to the agent, not guarantees enforced by Kitaru. The tutorial asks you to inspect whether the agent used that evidence correctly and whether the recorded action agrees with its final response.
What you will build
A verified agent version, ten imported sessions, and descriptive evaluations
Confirm what the example preserved and select a bounded, varied review worklist.
An investigation, evidence-linked annotations, and verdicts
Store what a human concluded without rewriting the trace.
One accepted behavior, an immutable cohort version, and an evaluator version
Turn reviewed evidence into a repeatable measurement.
A candidate agent version, experiment, and experiment run
Run one bounded change against the frozen population under an explicit tool policy.
Paired baseline and replay evidence
Decide whether the result is improved, regressed, a trade-off, or inconclusive.
Each page begins with the same five-step map. The first four phase pages end with a Checkpoint, and the final page summarizes the complete evidence chain. Because this is evidence-led, placeholders such as YOUR_SESSION_UUID are deliberate: substitute IDs produced by your own review rather than copying a predetermined ticket list.
Prepare the PydanticAI returns agent
Install jq, then open the PydanticAI returns agent README and complete its setup through the ten-session confirmation. That README is the source of truth for cloning, entering the example directory, the frozen environment, workspace selection, agent registration, worker startup, and the checked-in Langfuse import. Keep running the commands below from the example directory.
The example uses synthetic customers, orders, shipments, and actions. Refund and replacement tools modify only an isolated in-memory store. No model-provider or Langfuse credentials are needed for setup, import, or the deterministic parts of this tutorial.
Before continuing, confirm these conditions from the example README:
the selected workspace does not already contain tutorial resources named
returns-resolver,returns-discovery,returns-regression,returns-behavior, orreturns-candidate;returns-resolver@1is registered from the example directory;ten imported sessions have the
returns-baselinetag; andthe example worker remains running in the second terminal.
Stop and select another workspace if those resource names already exist. Do not delete an existing workspace merely to make its names available.
Some tutorial commands create jobs. The worker claims those jobs and performs the work in your environment, so the Kitaru server does not receive your agent code or model credentials. Keep the example worker running while you Observe, Judge, and Define. In the Replay phase you will restart it with OPENAI_API_KEY before any paid model call.
Prefer a coding agent?
The pages that follow teach the manual path so you can see each object and boundary. If you want an agent to guide the same evidence loop, install the Kitaru skills and use the guided-tour prompt in the PydanticAI returns agent README.
Start the investigation
Continue to 1. Observe the recorded behavior.
Last updated
Was this helpful?