kullback is the open-source harness behind Leibler.

Your traces in.
An environment out.

Traces are the logs your agent already writes. Kullback reads them, rebuilds the tools, data and rules the agent used, and replays the logs to check the rebuild. The claim we are working toward, not made yet: a 2B parameter model post-trained in that environment performs on the real tasks.

new traces, fixes, disputes tracesyour runs buildtools, state, checks runany model verdictcode, end state reportyou decide

The loop. Code decides pass or fail from the final data, not from the transcript. A judge model can take a pass away, never give one.

What a week of traces turns into.

Every value points back to the line in your trace it came from. Nothing is invented.

What has been measured so far.

Measured on Sierra's public retail traces (tau2), where the real tools and database exist to compare against. Seen: runs used for the build. Held out: runs the build never saw.

checkseenheld out
tool signatures0/15
starting rows0/252
writes0/370/20
reads0/1250/65
errors0/60/4

Runs where a model had to stand in for a missing tool are shown in the report, never counted as passes.

All of it is public.

One claim to earn.
A 2B model, post-trained in this.

Not claimed yet. Next: the same rebuild on domains it has never seen, then a 2B parameter model post-trained on trajectories the verifier passed, measured on the real held-out tasks. The numbers get published either way.