2026-09-24 · America/Los_Angeles · 社区动态 · #3
I benchmarked repowise, CodeGraph, Serena, Graphify, code-review-graph, cocoindex and codebase-memory-mcp across Codex, Claude Code and a local model. The 60-90% token-saving claims didn't hold up
Scroll to bottom for tldr In July, JetBrains reran the headline claims of two token-saving tools on real agent workloads. Caveman claimed 65% and measured 8.5%. RTK claimed 60–90% and ended up slightly more expensive than using nothing. How the tools fool you?? It felt like every tool out there was overclaiming, so I benchmarked 5 token saving tools with conditions closer to how agents actually use them my setup was: 48 Django questions drawn from SWE-bench Five question types, selected before running anything Same agent, prompt, repository commit and tool access Fresh index for every tool…
热度 49.3 / 100;排名与评分保留该期记录。