Retire used research
Your agent searches the codebase to find a function.
Once it finds it, the full search trail is no longer needed while writing code.
Winglet sends the model a short note and keeps the full output on disk if needed.
Results may vary by workload and workflow.
An AI usage saver that actually works. Backed by research.
Winglet seamlessly integrates with your coding agent. It cuts stale reads, repeated tool output, and wasted context before they reach the model. Everything stays on your machine.
3-day free trial · no credit card required
agent-winglet
33% saved
Your plan goes ~49% further
14.7 MiB
Directly measured
13.2M(est.)
Scaled from bytes saved
$54.55(est.)
Uses API pricing
WHERE THE SAVINGS COME FROM
Old investigation output archived
Long output trimmed
Repeat output skipped
SESSIONS (47)
Measured on August 8, 2026 while building Winglet with Winglet running.
Winglet is a desktop app. It integrates with Claude Code and Codex through hooks, so it can optimize what gets sent to the model while your agent thinks and uses tools. It runs locally, no data leaves your machine.
It's simple: Send less junk to the model.
Your agent searches the codebase to find a function.
Once it finds it, the full search trail is no longer needed while writing code.
Winglet sends the model a short note and keeps the full output on disk if needed.
80% of useful information is usually in the first and last 15 lines.
Winglet sends that head and tail, and keeps the full output on disk if needed.
This also implicitly steers the model toward focused searches.
Sometimes the agent asks for the same thing twice.
If the output is unchanged, it is already in context.
Winglet does not send it to the model again.
Winglet offers to compact after the agent uses what it found.
This reduces context and saves more usage.
Note: Compacting is optional. It is not part of Winglet's "usage saved" calculation.
Papers that inspired Winglet
"If I have seen further, it is by standing on the shoulders of giants." - Sir Isaac Newton
Criterion
Winglet
Caveman
RTK
Headroom
Winglet
Trajectory reduction
Drops useless command output, repeated file reads, and older observations that no longer affect the next step.
Caveman
Caveman-style replies
Rewrites prose like 'few word enough' while leaving tool calls, code, diffs, and error text unchanged.
RTK
Bash output rewriting
A PreToolUse hook rewrites eligible shell output; Read/Grep, python3, scripts, loops, and tricky shell syntax bypass it.
Headroom
Compress-then-retrieve
Stores the original output behind retrieval IDs. Claude sees compressed text first and has to ask for missing details.
Winglet
Your plan goes 26.7%-56.0% further
Based on computational cost savings of 21.1%-35.9% and input token savings of 39.9%-59.7%.
Caveman
Your plan goes 9.3% further
Based on JetBrains' evaluation measuring 8.5% cost saved (vs. 65% advertised).
RTK
Your plan goes 7.1% less far
Based on JetBrains' evaluation measuring 7.6% higher cost (vs. 60-90% advertised savings).
Headroom
Your plan goes 3.8% further
Based on a community evaluation measuring 3.7% saved, while Headroom advertises 15-20% token savings for coding agents (not cost savings).
Winglet
Remains the same
Parity with the original agent (ranging between -1.0% and +2.0%).
Caveman
-1.5% lower average task score
Average task score moved from 0.326 to 0.311 in JetBrains' evaluation.
RTK
No detectable quality difference
Quality remains the same despite cost increasing 7.6%.
Headroom
Remains the same on internal evals
Parity with the original agent on internal evals based on pure compression and Q&A tasks.
Caveman
Chat Q&A tasks
Despite being advertised as a coding-focused tool, its benchmark is chat Q&A.
RTK
CLI command examples
cargo test, pytest, go test, git diff, git status, git log, cat, grep, find, ls, deps, and one rtk gain screenshot.
Headroom
Generic compression suites
Benchmarks include GSM8K, TruthfulQA, SQuAD v2, BFCL, JSON logs, and search-output workloads.
One plan covers every agent.
Every tier is included.
Your Winglet plan covers all plans across Claude Code and Codex. You do not have to switch to a different Winglet plan if you use something like Max 5x.
Any issues with installation? Contact us at hi@agentwinglet.com.
Your usage and savings data stays on your machine. Suitable for private repos, client work, and internal codebases.