I compared AOCI + Codex with native Codex in a production repository of roughly 360,000 lines. I saw no efficiency gain in this test: elapsed time and uncached input tokens increased in all four scenarios. One task was blocked by the onboarding validation gate.
In the most comparable completed full-repository pair, both answers scored 6/6. AOCI took 2.93× as long and used 3.40× the uncached input tokens. That excludes nearly four hours of one-time indexing.
So far, I have not found a way for AOCI to pay off in my own work.
My principle: every added layer needs evidence of its benefit
I prefer giving the agent the problem and the necessary tools, then letting it choose how to use them.
Indexes, workflows and other scaffolding are worth trying. Keeping them depends on a concrete question: for the same task and comparable quality, is the result faster, cheaper or more reliable?
GitHub stars and recommendations help me discover tools. Comparisons on real tasks determine whether I use them. As models improve, I want clearer evidence that an added layer earns its place.
Test details
A code index organizes a repository in advance to help an agent find relevant material. Think of it as a book’s table of contents. Building and reading that table also costs something.
Both setups used gpt-6.1-sol / low. Each task started in a fresh session. Two tasks covered an eight-file slice; two covered the full repository.
| Task | Seconds: native → AOCI | Uncached input tokens: native → AOCI | Outcome |
|---|---|---|---|
| Call-chain trace, 8 files | 69 → 105 | 36,479 → 44,068 | Both answered |
| Bug diagnosis, 8 files | 26 → 58 | 9,445 → 28,210 | Both fixes passed 15 tests |
| Source-null trace, full repository | 106 → 230 | 60,803 → 169,731 | AOCI gate blocked; no business answer |
| Forecast to release, full repository | 111 → 326 | 61,495 → 209,236 | Both answers scored 6/6 |

One-time full-repository indexing took about 3 hours 55 minutes and used 6.07 million uncached input tokens, including failed attempts and recovery. Task times exclude indexing. Uncached input tokens are not a direct billing estimate.
Possible explanations
- A large repository can still contain small tasks. With a clear entry point, failing example or call chain, native Codex can locate the relevant code through search and a few source reads.
- The index adds a fixed onboarding cost. Both full-repository tasks loaded all 22 index chunks and performed delivery confirmation and cognitive challenges. Precise answers still required source verification, potentially adding index reading to code reading.
- Reuse remains untested. Fresh sessions repeated the onboarding cost. Several tasks in one session might produce different results.
These are possible explanations. I did not separately time index loading, validation and business reasoning, so I cannot attribute the slowdown precisely.
These were single runs, with only one completed full-repository pair. Long-session reuse and amortization remain untested. The results support my current usage decision, without establishing how AOCI performs across all repositories and tasks.
If you have measured a clear gain, I would like to know which task, how you integrated the tool, and where the benefit came from.