Click endurance
Both runs passed
32.2%fewer tokens
About this comparison
10-turn debugging session; reduction is calculated from the displayed aggregate token counts. See the result in context
Evidence
A closer look at task completion and token use in earlier Kairn campaigns. Historical beta results, with their limits in view.
Codex, with and without Kairn · 25 official tasks
With Kairn
20/ 25
tasks passed
Without Kairn
15/ 25
tasks passed
Combined historical SWE-Pro campaigns. These totals are not a prediction for every repository.
Three selected comparisons where both runs passed
Both runs passed
32.2%fewer tokens
10-turn debugging session; reduction is calculated from the displayed aggregate token counts. See the result in context
Both runs passed
16.4%fewer tokens
Fresh long-session JavaScript package test. See the result in context
Both runs passed
54.2%fewer tokens
Official evaluator row. See the result in context
The completion totals combine 25 official SWE-Pro task evaluations from earlier Kairn campaigns. The token examples above are separate selected comparisons from Click, Node SemVer, and OpenLibrary.
Token reductions compare observed model usage only when both runs pass. They are not estimated savings or a prediction for your repository. Versions and test conditions vary across the historical campaigns.
Read the historical methods and limitationsInstall once. Try it on your next task.