ghfs

Benchmarks

Real measurements, real tasks

10 runs × 2 conditions × 2 task types, all with Claude Opus 4.6. Every run is shown, including the ones that went the other way.

What ten runs can and cannot show. Two of these results are clear: during a task ghfs makes no API calls at all, and multi-issue work costs more. The single-issue token and time differences are not clear — they average in ghfs’s favour, but four of the ten runs went the other way. Each section below says which kind it is.

Single-Issue Task

Investigate one issue, implement a fix

Given an issue requesting i18n (internationalization) support, the AI investigates the codebase and produces an implementation plan. A focused task that represents typical day-to-day AI-assisted development.

Tokens (avg)
1,356K 1,164K
lower in 6 of 10 runs
AI API calls (avg)
1.0 0.0
0 in 10 of 10 runs
Time (avg)
151s 131s
faster in 6 of 10 runs
Stability (CV)
37.3% 22.2%
not separable at 10 runs

Token consumption per run

Without ghfs With ghfs
#1
1,234K
1,451K
#2
1,520K
1,257K
#3
2,390K
1,147K
#4
889K
885K
#5
1,497K
864K
#6
1,892K
1,667K
#7
770K
1,086K
#8
806K
1,081K
#9
1,204K
1,289K
#10
1,359K
919K

Every run, in full

Run Tokens without Tokens with Time without Time with API calls without API calls with
#1 1,234K 1,451K 127s 116s 1 0
#2 1,520K 1,257K 185s 105s 1 0
#3 2,390K 1,147K 227s 144s 1 0
#4 889K 885K 115s 112s 1 0
#5 1,497K 864K 169s 124s 1 0
#6 1,892K 1,667K 148s 184s 1 0
#7 770K 1,086K 104s 120s 1 0
#8 806K 1,081K 111s 145s 1 0
#9 1,204K 1,289K 133s 147s 1 0
#10 1,359K 919K 190s 115s 1 0

The API call goes away. The rest is not separable at ten runs.

With ghfs the agent reads the issue from disk instead of calling the GitHub API, and that call is gone in all ten runs. Tokens came out 14% lower and wall-clock 20 seconds faster on average, but both flipped direction in four of the ten runs (Welch’s t = 1.07 and 1.32), and the drop in run-to-run variance falls just short of the usual two-sided 5% threshold (F = 3.84 against 4.03). We show the averages because they are what we measured. We are not claiming a saving from ten runs.

Multi-Issue Task

Cross-cutting research across all open issues

The AI reads all open issues, analyzes priorities, and produces a roadmap report. With ghfs the issues are already on disk, so the task runs without a single API call. The agent also reads more of them, and that costs more tokens and more time.

0
AI API calls
10 of 10 runs, against 1 to 11 without
+49%
tokens
higher in 9 of 10 runs
+26%
wall-clock time
slower in 9 of 10 runs

Token consumption per run

Without ghfs With ghfs
#1
871K
925K
#2
424K
1,007K
#3
890K
1,014K
#4
869K
1,107K
#5
481K
821K
#6
739K
1,173K
#7
687K
1,097K
#8
477K
1,340K
#9
709K
1,296K
#10
1,055K
948K

Every run, in full

Run Tokens without Tokens with Time without Time with API calls without API calls with Issues without Issues with
#1 871K 925K 118s 122s 6 0 27 32
#2 424K 1,007K 96s 117s 1 0 23 14
#3 890K 1,014K 126s 141s 10 0 18 31
#4 869K 1,107K 122s 123s 1 0 31 24
#5 481K 821K 70s 113s 1 0 16 30
#6 739K 1,173K 99s 101s 8 0 12 23
#7 687K 1,097K 92s 129s 9 0 16 28
#8 477K 1,340K 86s 201s 1 0 31 26
#9 709K 1,296K 100s 158s 7 0 32 28
#10 1,055K 948K 129s 100s 11 0 21 19

Multi-issue work costs more. Here is all of it.

ghfs puts every open issue on disk, so the agent reads all of them instead of fetching a subset through the CLI. Across ten runs that came to 49% more tokens (higher in 9 of 10, t = 4.20) and 26% more wall-clock time (slower in 9 of 10, t = 2.35). It also read 26 issues on average against 23 without, but that gap sits inside the run-to-run spread (t = 0.96), so we are not presenting it as fuller coverage. What is left is the API calls: 1 to 11 without ghfs, 0 in every run with it. Whether that trade is worth taking depends on your task, and we would rather show you both sides than argue it.

Coming Next

The best of both worlds

The multi-issue cost above is the thing we most want to reduce, and loading only the issues a task needs is how we expect to do it. These are what we are building toward. We are not putting dates on them.

Planned

Query Files

Load only the issues you need with condition-based filtering. Query for specific labels, states, or keywords — solving the multi-issue token increase.

Exploring

Semantic Search

Find related issues intelligently using embeddings and vector search. Instead of reading every issue to find connections, the AI asks a question and gets the relevant issues back — no full scan needed.

Methodology

Model claude-opus-4-6 (Claude Code default)
Version Pre-release build, measured February 2026 (before v1.0.0)
Measured 2026-02
Runs 10 per condition
Conditions ghfs enabled vs disabled
Task 1 Single-issue: given an i18n request, investigate and produce implementation plan
Task 2 Multi-issue: analyze all open issues and produce priority roadmap
Fair conditions Clean working state, identical prompts, no cached conversations, prompt cache warmed before measurement