Rendered at 02:30:17 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
solenoid0937 2 hours ago [-]
So this is just delegating certain work to dumber models? I certainly wouldn't use Gemini 2.5 Flash (!!?) for code writing as suggested.
I've never had an issue with Codex or Claude reading massive files, they're really good at precise greps.
jampa 16 minutes ago [-]
> I've never had an issue with Codex or Claude reading massive files
Reading files isn't a problem they want to solve. The idea seems to be using a cheaper model to "scout" for the intended code, instead of an expensive one that reads all the things (and spends more tokens / thinks about them).
I think this might be useful because Opus 5 especially tends to over-read. So this looks like an "LLM Bloom filter", telling "hey this is the code you might want to read".
bensyverson 55 minutes ago [-]
Yes, this makes little sense. It looks like it's a way to avoid having Claude read or write your code.
And why stop at 90%? I have this one weird trick to reduce Claude Code token use by 100%: use a different harness and model!
1 hours ago [-]
14u2c 46 minutes ago [-]
This does seem to just be a subagents implementation.
Banditoz 37 minutes ago [-]
Oh dear, why does this website override scrolling behavior?
orliesaurus 33 minutes ago [-]
glad im not the only one that enabled screen reader mode to scan the article for some goodies
jnwatson 2 hours ago [-]
It cuts token usage because they are using a different service with a different token budget for the reader/code writer tasks.
You can also just delegate this to subagents with Claude Code (though you have a more limited choice of models unless you swap the cheaper models via OpenRouter).
I'm OK using a dumb model as a smart grep, but the whole point of using the frontier models is using their intelligence for the hard stuff like coding.
FelineStateMach 17 minutes ago [-]
I sometimes get jumpscaped at the thought of older or less proven models used in enterprise settings. I understand the devex ergonomics argument; I'm not a fan of profiles concepts typically if trodding into delegation.
gruez 55 minutes ago [-]
>The benchmarks
>Tested against a Java monorepo across four scenarios, measuring tokens Claude would consume reading files directly vs. consuming the bulk-reader's summary or writing code via the code-writer. Mean bulk-read savings were around a whopping 90%.
>The code-write scenario is harder to measure in tokens because without shunt, Claude both reads the reference files and generates the output as expensive output tokens. With shunt, the code goes straight to disk, Claude never sees it.
So nothing about accuracy or actual performance? At least run against DeepSWE bench or something.
ryuuseijin 17 minutes ago [-]
Here is another technique to save tokens: allow the model to read a skeleton of the source code before reading the code, to give it an index into the code so it can read targeted chunks.
There is a tool that uses ripgrep and treesitter that does this [1], adapted from the maki coding agent.
Aider pioneered this with the "repo map" which works tremendously well.
tolugenius 2 hours ago [-]
Isn't this a somewhat standard multi-model setup? there's nothing ground breaking here, just delegate claude to plan -> smaller model for implementation.
cute_boi 20 minutes ago [-]
STOP hijacking my scroll. I don't know why chrome even allow such behavior?
And, I can't believe this is from official spotify.... What a joke.
I've never had an issue with Codex or Claude reading massive files, they're really good at precise greps.
Reading files isn't a problem they want to solve. The idea seems to be using a cheaper model to "scout" for the intended code, instead of an expensive one that reads all the things (and spends more tokens / thinks about them).
I think this might be useful because Opus 5 especially tends to over-read. So this looks like an "LLM Bloom filter", telling "hey this is the code you might want to read".
And why stop at 90%? I have this one weird trick to reduce Claude Code token use by 100%: use a different harness and model!
You can also just delegate this to subagents with Claude Code (though you have a more limited choice of models unless you swap the cheaper models via OpenRouter).
I'm OK using a dumb model as a smart grep, but the whole point of using the frontier models is using their intelligence for the hard stuff like coding.
>Tested against a Java monorepo across four scenarios, measuring tokens Claude would consume reading files directly vs. consuming the bulk-reader's summary or writing code via the code-writer. Mean bulk-read savings were around a whopping 90%.
>The code-write scenario is harder to measure in tokens because without shunt, Claude both reads the reference files and generates the output as expensive output tokens. With shunt, the code goes straight to disk, Claude never sees it.
So nothing about accuracy or actual performance? At least run against DeepSWE bench or something.
There is a tool that uses ripgrep and treesitter that does this [1], adapted from the maki coding agent.
[1]: https://github.com/ninjaxtools/treesitter-index
And, I can't believe this is from official spotify.... What a joke.