27 September 2026·1 min read
Jev vs Laya. Who gets to throw your agent's context away?
Coding agents get worse as old tool output piles up in their context. Jev and Laya both decide what to throw away with a typed answer instead of a summary. One runs as a hosted API, and the other runs on your own GPU.

Coding agents get worse the longer they run, because their context fills with old tool output that they reread every turn. Something has to decide what to throw away. Jev and Laya both answer that as a typed decision instead of prose. The difference is where the judge lives. Jev is a hosted API, and Laya is open weights you run yourself.
We went down this hole this week because the usual fix is a summary. When the window fills, a language model rewrites the history shorter, which is roughly how Claude Code's /compact works. Summaries are paraphrases, and paraphrases are bad at exact things. A file path becomes "the config file". An error code becomes "an auth error".
Decision models skip the paraphrase. You give them a state and a closed question, and they return an answer with a probability. That's the shape of the garbage question. fast-jev-compaction, a Claude Code plugin built on Jev, asks two things about every tool call. Is the call worth keeping, and is its result worth keeping verbatim? That gives three outcomes: keep everything, keep the call with a truncated result, or drop both. Your messages and the agent's replies are never touched, and the first and latest messages are pinned.
Jev comes from TypeSafe AI, which released it on 15 September and calls it a System One model. You call it over the network, and it takes up to 64k tokens per request. Laya you download, run on your own GPU and fine-tune. The Laya checkpoint in the comparison we read takes 512 tokens per question, though its project lists different limits for different checkpoints.