Benchmarks
Real measurements, real tasks
10 runs × 2 conditions × 2 task types, all with Claude Opus 4.6. Every run is shown, including the ones that went the other way.
What ten runs can and cannot show. Two of these results are clear: during a task ghfs makes no API calls at all, and multi-issue work costs more. The single-issue token and time differences are not clear — they average in ghfs’s favour, but four of the ten runs went the other way. Each section below says which kind it is.
Single-Issue Task
Investigate one issue, implement a fix
Given an issue requesting i18n (internationalization) support, the AI investigates the codebase and produces an implementation plan. A focused task that represents typical day-to-day AI-assisted development.
Token consumption per run
Every run, in full
| Run | Tokens without | Tokens with | Time without | Time with | API calls without | API calls with |
|---|---|---|---|---|---|---|
| #1 | 1,234K | 1,451K | 127s | 116s | 1 | 0 |
| #2 | 1,520K | 1,257K | 185s | 105s | 1 | 0 |
| #3 | 2,390K | 1,147K | 227s | 144s | 1 | 0 |
| #4 | 889K | 885K | 115s | 112s | 1 | 0 |
| #5 | 1,497K | 864K | 169s | 124s | 1 | 0 |
| #6 | 1,892K | 1,667K | 148s | 184s | 1 | 0 |
| #7 | 770K | 1,086K | 104s | 120s | 1 | 0 |
| #8 | 806K | 1,081K | 111s | 145s | 1 | 0 |
| #9 | 1,204K | 1,289K | 133s | 147s | 1 | 0 |
| #10 | 1,359K | 919K | 190s | 115s | 1 | 0 |
The API call goes away. The rest is not separable at ten runs.
With ghfs the agent reads the issue from disk instead of calling the GitHub API, and that call is gone in all ten runs. Tokens came out 14% lower and wall-clock 20 seconds faster on average, but both flipped direction in four of the ten runs (Welch’s t = 1.07 and 1.32), and the drop in run-to-run variance falls just short of the usual two-sided 5% threshold (F = 3.84 against 4.03). We show the averages because they are what we measured. We are not claiming a saving from ten runs.
Multi-Issue Task
Cross-cutting research across all open issues
The AI reads all open issues, analyzes priorities, and produces a roadmap report. With ghfs the issues are already on disk, so the task runs without a single API call. The agent also reads more of them, and that costs more tokens and more time.
Token consumption per run
Every run, in full
| Run | Tokens without | Tokens with | Time without | Time with | API calls without | API calls with | Issues without | Issues with |
|---|---|---|---|---|---|---|---|---|
| #1 | 871K | 925K | 118s | 122s | 6 | 0 | 27 | 32 |
| #2 | 424K | 1,007K | 96s | 117s | 1 | 0 | 23 | 14 |
| #3 | 890K | 1,014K | 126s | 141s | 10 | 0 | 18 | 31 |
| #4 | 869K | 1,107K | 122s | 123s | 1 | 0 | 31 | 24 |
| #5 | 481K | 821K | 70s | 113s | 1 | 0 | 16 | 30 |
| #6 | 739K | 1,173K | 99s | 101s | 8 | 0 | 12 | 23 |
| #7 | 687K | 1,097K | 92s | 129s | 9 | 0 | 16 | 28 |
| #8 | 477K | 1,340K | 86s | 201s | 1 | 0 | 31 | 26 |
| #9 | 709K | 1,296K | 100s | 158s | 7 | 0 | 32 | 28 |
| #10 | 1,055K | 948K | 129s | 100s | 11 | 0 | 21 | 19 |
Multi-issue work costs more. Here is all of it.
ghfs puts every open issue on disk, so the agent reads all of them instead of fetching a subset through the CLI. Across ten runs that came to 49% more tokens (higher in 9 of 10, t = 4.20) and 26% more wall-clock time (slower in 9 of 10, t = 2.35). It also read 26 issues on average against 23 without, but that gap sits inside the run-to-run spread (t = 0.96), so we are not presenting it as fuller coverage. What is left is the API calls: 1 to 11 without ghfs, 0 in every run with it. Whether that trade is worth taking depends on your task, and we would rather show you both sides than argue it.
Coming Next
The best of both worlds
The multi-issue cost above is the thing we most want to reduce, and loading only the issues a task needs is how we expect to do it. These are what we are building toward. We are not putting dates on them.
Query Files
Load only the issues you need with condition-based filtering. Query for specific labels, states, or keywords — solving the multi-issue token increase.
Semantic Search
Find related issues intelligently using embeddings and vector search. Instead of reading every issue to find connections, the AI asks a question and gets the relevant issues back — no full scan needed.
Methodology
| Model | claude-opus-4-6 (Claude Code default) |
| Version | Pre-release build, measured February 2026 (before v1.0.0) |
| Measured | 2026-02 |
| Runs | 10 per condition |
| Conditions | ghfs enabled vs disabled |
| Task 1 | Single-issue: given an i18n request, investigate and produce implementation plan |
| Task 2 | Multi-issue: analyze all open issues and produce priority roadmap |
| Fair conditions | Clean working state, identical prompts, no cached conversations, prompt cache warmed before measurement |