kairn

Evidence details

The full historical record.

Explore the tasks behind the headline results, including successful comparisons, failures, and cases where Kairn used more tokens. “Pass/pass” means both runs completed the task successfully.

Historical evidence — not calibrated-ranker results

These rows predate the calibrated evidence ranker and the frozen four-arm comparison protocol. They remain visible as historical, quality-gated evidence and are not presented as current ranker results.

The three-file blocklist example

The homepage walkthrough illustrates one historical official coding task in qutebrowser/qutebrowser. A blocked parent domain needed to cover its subdomains, while allowlisted sites and disabled blocking still had to work. The required fix involved three production files:

In the recorded comparison, Codex alone changed tests/unit/components/test_hostblock.py and failed the official evaluator. The final Kairn-assisted run changed the three production files, did not edit the test, and passed. An earlier Kairn attempt missed the third source file and also failed; a general routing fix preceded the passing run.

This is one historical quality result, not proof that every agent needs help with this task. Because the baseline failed, its token count is not included as a clean savings comparison.

The homepage uses this task to illustrate what Kairn can do across a session: find a focused starting set, avoid repeating unchanged code, and respond when changed files or a reported failure provide new evidence. Those follow-up events are illustrative possibilities, not a transcript or measured sequence from this qutebrowser run. The recorded pass/fail outcome above is separate.

The earlier Click illustration

The previous homepage example used a Click task where colored text must have visible length three. Kairn identified src/click/_compat.py as the starting file. That example showed file guidance, but the agent without Kairn also passed its recorded check, so it did not demonstrate a quality difference. The Click endurance result remains in the historical ledger below.

How the historical results were measured

The completion comparison combines 25 official SWE-Pro evaluations from earlier campaigns: 20 passed with Kairn and 15 without it. It is an aggregate across those campaigns, not one newly controlled study.

The three token examples are separate selected successful comparisons. Click measures a 10-turn debugging session; Node SemVer covers a JavaScript package session; OpenLibrary is an official evaluator row. Reductions use the displayed observed model-token counts, calculated as (without Kairn − with Kairn) / without Kairn.

These campaigns used different versions and test conditions. The selected examples do not represent all outcomes or establish universal savings. The ledger below preserves failures, regressions, run types, and scope limitations. No uncertainty interval is claimed for the historical headline totals.

Clean savings

Counted only when both baseline and Kairn pass the quality gate.

Quality/source focus

Reported separately when Kairn passes and baseline misses quality.

MCP evidence

Requires actual MCP calls plus useful source guidance, not just a configured server.

Fresh repeat-1 checks

Useful new signal, but repeat-3 is needed before it becomes a headline benchmark.

The planned comparison protocol describes a separate future experiment. It is not the methodology behind the historical headline figures.

Clean pass/pass savings

Rows where both baseline and Kairn completed the task successfully, so token reduction can be counted directly.

WorkflowTypeBaselineKairnReductionQualityStatusNotes
Click enduranceLong session
repeat-3
90,73361,52832.2%pass/passclean savings10-turn debugging session with zero scope violations; this percentage is the signed delta of the displayed aggregate counts.
Node SemVer enduranceLong session
repeat-3
69,51358,14216.4%pass/passclean savingsFresh JavaScript package endurance row outside the earlier Python-heavy set.
OpenLibrary SWE-ProOfficial evaluator
official row
21,93210,04854.2%pass/passclean official savingsOfficial evaluator pass/pass row; counted separately from quality-rescue rows.

Latest fresh checks

Newer post-launch checks that are useful signal, but not promoted to headline claims until repeated.

WorkflowTypeBaselineKairnReductionQualityStatusNotes
Fresh source-routing checkSource routing
repeat-1 fresh check
122,98075,33938.7%3/3 pass/passfresh repeat-1Active-valid assistance with scope discipline passing. Needs repeat-3 before it becomes a headline benchmark.
Fresh SWE-Pro governance sliceOfficial evaluator
fresh official slice
4/8 known local passes9/9 valid official passes9.3% clean median10/10 local passfresh officialPost-governor branch check. Strongest signal is source/scope discipline; one row excluded for evaluator image availability.

Endurance and MCP evidence

Long-session and MCP rows where Kairn had to stay useful beyond the first turn or through actual MCP calls.

WorkflowTypeBaselineKairnReductionQualityStatusNotes
httpcore MCP enduranceMCP endurance
repeat-3
recordedrecorded18.5% median100%clean MCPActual MCP calls, useful returned files, and zero scope violations.
Requests MCP post-fixMCP endurance
repeat-3
recordedrecorded28.1% median3/3 passclean MCPUseful MCP file returns in all three runs; zero scope violations.
Click MCP-firstMCP endurance
repeat-3
recordedrecorded52.6% medianpass/passclean MCPMCP-first variant of the Click endurance suite.

Official evaluator evidence

Rows run through the official evaluator path. Clean savings are counted only when both sides pass.

WorkflowTypeBaselineKairnReductionQualityStatusNotes
SWE-Pro official campaignsOfficial evaluator
25 official rows
15/25 passed20/25 passed6 clean rowsofficial evaluatorquality/source-focus + clean savingsClean rows saved 63,393 observed tokens with a 22.1% median reduction; rescue rows are reported separately.
SWE-Pro new expansionOfficial evaluator
20 new rows
13/20 passed15/20 passed5 clean rowsofficial evaluatorfresh official expansionNew expansion rows saved 51,509 observed tokens across clean pass/pass savings rows.
SWE-Pro governance slice clean rowsOfficial evaluator
3 clean pass/pass rows
174,620154,62711.4%official pass/passclean official savingsLatest slice clean rows had positive savings; source/scope discipline was the stronger result.

What does not count as clean savings

These rows are useful for product learning, but they are not used as headline savings claims.

WorkflowTypeBaselineKairnReductionQualityStatusNotes
Fresh live canaryLive canary
repeat-1
missed qualitypassedsupporting onlyquality rescuenot clean savingsUseful because Kairn improved the outcome, but not a pass/pass savings claim.
SWE-Pro token regressionsOfficial evaluator
tracked regressions
official passofficial pass4 regressionspass/passnot clean savingsSome official pass/pass rows used more tokens with Kairn; these are tracked as optimization targets.
Suppressed or passive rowsGovernor behavior
varies
variessilent or tinysupporting onlyvariesnot active savingsCorrect silence is product evidence, but it is reported separately from active token wins.

Claims to keep straight

HistoricalPre-ranker campaigns show that Kairn can reduce observed model tokens on quality-passing source-rescue, debug, MCP, and endurance workflows.
SafeKairn may stay silent when confidence is low; that is intentional suppression.
SafeCodex CLI/session is the most-tested path; MCP is the portable editor path.
CarefulSavings vary by task and model; historical SWE-Pro rows separate clean savings from quality/source-focus evidence.
PlannedThe calibrated ranker and four-arm ReasonBlocks comparison are v0.3 work and have no published outcome yet.
AvoidDo not claim universal 20-70% savings or public works-anywhere reliability yet.
Back to overviewInstall KairnPlanned comparison protocol