kairn

Evidence

What the results show.

A closer look at task completion and token use in earlier Kairn campaigns. Historical beta results, with their limits in view.

More tasks completed.

Codex, with and without Kairn · 25 official tasks

With Kairn

20/ 25
tasks passed

Without Kairn

15/ 25
tasks passed

Passed Did not passEach square represents one task in the totals.

Combined historical SWE-Pro campaigns. These totals are not a prediction for every repository.

Fewer tokens for the same task.

Three selected comparisons where both runs passed

Click endurance

Both runs passed

32.2%fewer tokens

About this comparison

10-turn debugging session; reduction is calculated from the displayed aggregate token counts. See the result in context

Node SemVer endurance

Both runs passed

16.4%fewer tokens

About this comparison

Fresh long-session JavaScript package test. See the result in context

OpenLibrary SWE-Pro

Both runs passed

54.2%fewer tokens

About this comparison

Official evaluator row. See the result in context

How these results were measured

The completion totals combine 25 official SWE-Pro task evaluations from earlier Kairn campaigns. The token examples above are separate selected comparisons from Click, Node SemVer, and OpenLibrary.

Token reductions compare observed model usage only when both runs pass. They are not estimated savings or a prediction for your repository. Versions and test conditions vary across the historical campaigns.

Read the historical methods and limitations
Full results, failures & caveats

See what changes in your repository.

Install once. Try it on your next task.

Install Kairn