What it is: a new projection head for Qwen3-8B.
github: https://github.com/Contrastive-LM/CLM
hf: https://huggingface.co/Contrastive-LM
The whole non-autoregressive calibrated decision making was underutilized. The entire point of Jev is the Zero-Shot Broad Knowledge.
At the API and functional interface level, CLM supports everything Jev does—it is not a subset. However, there are important trade-offs in generalization, context scale, and architecture between the two.
1. Functional Parity (Same Primitives)
CLM was specifically engineered as an open-weights, self-hostable alternative to TypeSafe AI's Jev. It implements the exact same "System One" decision interface and supports all three of Jev’s core question primitives:
Choice: Evaluates a discrete set of candidates and returns a categorical probability distribution.
Noul: Outputs a calibrated true/false probability for a proposition or guardrail check.
Score: Scores an input against an ordered rubric or scale.
Code written for the TypeSafe Jev client can be pointed directly at a clm-serve endpoint with drop-in compatibility (from clm import CLMClient, Choice, Noul, Score).
2. Where CLM Outperforms Jev
Latency and Disaggregated Caching: Jev is a proprietary cloud model that evaluates state and question choices jointly. CLM separates the state head from the action head. If an agent has a persistent set of tools or actions, CLM embeds those actions once and caches them. In benchmarks like interactive browser agents and gaming (T-Rex, Super Mario), CLM is 4× to 13× faster than Jev.
Open Weights & Fine-Tunability: Jev is a closed API with no user fine-tuning (you can only prompt it via state and question instructions). Because CLM’s heads are tiny open weights (~75 MB), you can fine-tune them on your own agent trajectories.
Coding Benchmark Verifiers: When fine-tuned on agent trajectories, CLM achieves state-of-the-art verifier performance on Terminal-Bench 2.1 (87.6%) and DeepSWE (81.6%), whereas zero-shot Jev struggled on those exact benchmarks (scoring ~71% on DeepSWE).
3. Where Jev Still Has the Edge (CLM-8B Limitations)
While CLM covers the entire feature surface of Jev, the current CLM-v0.1-8B release trails Jev in a few areas:
Zero-Shot Broad Knowledge: Jev is backed by a larger, proprietary model On zero-shot open-domain tasks, Jev still holds an edge in edge-case accuracy (e.g., Berkeley Function Calling Leaderboard v4: Jev scored 99.2% vs. CLM-8B’s 95.2%; WikiRacing: Jev 30/30 vs. CLM-8B 26/30).
Context Budget: Jev accepts requests up to a 64K token context out-of-the-box. CLM-8B was tested and calibrated at 2K to 8K context. While its Qwen3 backbone can accept longer prompts, representations past 8K haven't been calibrated for the reference head.
Probability Normalization: CLM calculates probabilities via dot products and softmax over the candidates passed in that request Its probabilities are inherently relative to the candidate set provided, whereas Jev’s scoring is calibrated internally against absolute criteria.
Summary
If you are asking if you will lose API features by using CLM instead of Jev: No, you get the full primitive set (Choice, Noul, Score) with massive latency gains and zero API costs. You only sacrifice some zero-shot generalization on niche out-of-domain tasks compared to TypeSafe's hosted service.
Use cases - You can use it to build classifiers like Jev being used to babysit background processes being used by agentic LLMs, such a builds or network operations, etc. Being able to determine if a background process should be given more time/iterations, or control flow returned to the supervising LLM could save a lot of tokens, especially for long-running tasks where the KV cache TTL keeps going over between iterations causing input tokens to be uncached on status checks, or edge devices with slow LLMs..
Classifiers like Jev seem well-suited to determine if the system logs have anything actionable in them to take action on.
The idea that Jev's superpower is zeroshot broad knowledge is exactly why it's backwards. Its very nature begs to be fine tuned.
You always have to set it up for success by correctly engineering its options and instructions. And if you can already do this, then you likely can also finetune a small model to perform even better.
You can tell qwen to play chess, and then in the next prompt tell it to create a flappy bird. But you are never gonna be able to do this with Jev.
Why would anyone who is serious about solving a real world scenario would choose Jev over BERT or the dozens of other alternatives?
Because it works on the fly for any problem. You can put completely arbitrary data along with whatever background context + related retrieved data in state and it just works.
32k shared context in state + 32k isolated per question.
It's not even remotely comparable to BERT style classification. You can solve obscenely hard problems with this.
Why not fine tune
Because of the system design implications of a model that just works.
If you have to do a fine-tuning pass every time you wanna account for some unknown case or rewrite your data pipeline it puts a big blocker on idea iteration.
Credits - https://www.reddit.com/r/LocalLLaMA/comments/1wouby6/jev_almost_dead_clm_vs_jev/
If we offload the decision making to Jev, then the LLM doesnt have context on why this action was performed. You essentially seperate the decision and reasoning from the model. And just tell the model just go do X. You also lose the chain of thought on why things occurred in this workflow dont you?
I dont see this being a problem in simple agentic tasks. But couldnt this be more of a problem in deep agentic and general purpose tasks?
Jev is meant as a quick classifier, not a reasoner. It checks things like tool calls or old tool results, and the main LLM still plans and answers. The context loss mostly shows up in compaction, where results get replaced by a "removed" placeholder. For gating, just pass the verdict back to the model as the tool result, so it sees why a call was blocked.
The problem isn't "can the LLM do this decision making for me" it is that the LLM is orders of magnitude more expensive.
If I have to call an LLM a million times a day to route an opening conversation of a chat bot "what does this sentence imply for a user workflow from this list...." Then I am spending a non-trivial amount to do that.
However, routing that many through jev is much less expensive and can help route to the correct workflow handler.
One of the best use cases I can see for these style of classifiers is to help the LLM decide if it should be reading something.
If you have the proper folder structure, and documentation to support it. Your LLM can "phone home" with a prompt to the classifier and say tell me what's relevant to this prompt per our documents and it'll fly through it.
Llms take time and often speed is the name of the game. Id strongly suggest fine tuning a model and not relying on a 3P model for this simply because you are exposing your data to another provider. You can train a BERT model on a 1070 for cheap and it's 2-10x faster than Jev. I trained a BERT on my email it's amazing. Out performed JEV for my use case.
Anyway, if the LLM needs to spend 20 seconds per file against idk 30 files that's just down time. Your classifier can handle each one in a out a second. Then you get into better model selections if you use it as the agent spawner. Is this really a Opus level task? Is it Sonnet for example, and it can help better pick agents to spawn and remove that from the orchestrator helping you increase speed and reduce tokens.
You can also use it to validate the outcome of gates as well. Though I would still use the LLM there to validate till you get confident on outcomes.
Another really cool use case is you can run simulations with them and see if the outcome of a function is fixed or not, and if it isn't what percentages are the options leaning too.
You can use it as your auto tool classifier, that already uses another llm. You can also return it's result into the llm so it knows what it chose from. Some people use it in compact... i would use it in a limited way there personally. You can use it to pick the right sub agent.
You can use it for possible efficiency gains in things like skill selection / loading and dynamic model routing.
We're exploring Cloudflare's Clef instead of Jev as a lightweight decision layer within the agent harness.
Clef's specific role would be to handle bounded decisions like intent classification, risk triage, and scoring predefined options. It returns structured results and probabilities, not tool executions. The primary LLM still handles planning and reasoning, while deterministic policy controls handle authorization and approvals.
I agree with your context concern. The key is feeding Clef's decision results and relevant context back into the workflow rather than letting it operate as a separate black box.
We also like Clef's open weights and self-hosting flexibility.
From https://www.reddit.com/r/AI_Agents/comments/1x0cw8r/is_jev_or_similar_actually_useful_in_a_harness/