TL;DR - Jev is a classifier. You give it a state and a set of questions, each constrained to yes/no, a named choice, or a score, and it returns probabilities and a confidence value from a single forward pass. TypeSafe claims it makes decisions 193.6× faster and 444.6× cheaper than an LLM [1]. It needs no labeled data, so an agent can offload any clear-cut decision to it at runtime. The agent should still check each answer’s confidence before acting on it, and keep arithmetic and date comparisons in its own code, where Jev is still unreliable [5].
Table of Contents
- What is Jev?
- Why it matters
- Using Jev
- Confidence vs. probability
- Failure modes
- Use cases
- Community applications
- Open-source alternatives
- References
What is Jev?
Jev is fundamentally a classifier. Instead of predicting the next word like GPT models, it predicts probabilities and a confidence level, kind of like BERT. It works like filling out a multiple-choice form instead of writing an essay.
Why it matters
Because multiple decisions can be made in a single forward pass, Jev is fast. It doesn’t spend tokens on thinking or on generating long responses. TypeSafe claims it is 193.6× faster and 444.6× cheaper than an LLM making the same decisions [1].
Jev also operates zero-shot: you define the categories or questions at runtime without labeled examples. That is its main advantage over both fine-tuned BERT classifiers and logistic regression on frozen embeddings, which both need labeled data up front [2]. So whenever your LLM needs to make a clear-cut decision, it can save time and tokens by calling Jev instead.
Using Jev
Jev answers in one of three formats:
- Noul: a yes/no answer
- Choice: one of a set of named options
- Score: a numeric rating on a scale you define
Because the output is constrained to these formats and defined categories, people describe it as “solving the hallucination problem”. Which is a silly claim, because it can still pick the wrong option; see Failure modes. Additionally, its decision quality now depends on the options and contexts you provide. If a customer complaint concerns fraud but the available answers are “billing,” “sales” and “technical,” the model cannot return the correct category.”
Example API call
Each API call needs a state (the context your questions are about) and one or more questions.
E.g., A support ticket triage request:
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"
]
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity"
}
}
}
And the API response:
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": {
"technical": 0.85,
"sales": 0.0,
"billing": 0.15
}
},
"frustration": {
"type": "score",
"score": 1.0,
"confidence": 1.0,
"legend": {
"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language"
},
"probabilities": {
"0": 0.0,
"1": 1.0,
"2": 0.0
}
},
"is_urgent": {
"type": "noul",
"noul": 1.0
}
},
"usage": {
"input_tokens": 392,
"output_tokens": 65
}
}
This single call returns three decisions with 65 output tokens. See the quickstart [3] for other API details.
Integrating with Claude Code
Use a skill to tell Claude:
- which kinds of questions suit Jev,
- how to shape the API call, and
- how to interpret the response.
TypeSafe publishes skills for this:
claude plugin marketplace add typesafe-ai/skills
claude plugin install typesafe@typesafe-ai
Alternatively, use one of the MCP connectors the community has already created:
Confidence vs. probability
Each choice and score answer includes a probabilities value (0–1) for every option and a single confidence value (0–1) that summarizes the spread of that distribution [4]. An even distribution gives a confidence of 0; a clear skew toward one option gives a high confidence. In the example above, department puts 0.85 on technical and 0.15 on billing, and reports a confidence of 0.78.
You can then pick the next step based on confidence:
- High: proceed.
- Medium: gather more context, or act and flag for review.
- Low: route back to a human.
The thresholds for high/medium/low confidence needs to be adjusted based on the number of options available and risk risk of whatever action follows the decision.
Failure modes
The failure modes documented in the offical website demonstrates the importance of treating Jev like a decision-making LLM. We must structure our calls in a self-consistent way.
| # | Failure mode | Do this instead |
|---|---|---|
| 1 | Literal reading | Write the exact condition, criteria for each available options |
| 2 | Math and Numbers | Keep the arithmetic in code |
| 3 | Date and time comparison | Extract components; compare in code |
| 4 | Indirection | Reduce hops; point to the relevant state |
| 5 | Large state full of irrelevant detail | Filter first; send only what the question needs |
| 6 | Adversarial content | Write precise prompts, and test edge cases before deploying |
| 7 | Contradictory instructions and criteria | Align the criteria and instruction |
| 8 | Common-sense structural invariants | Ask each decision one way; enforce identities in code |
| 9 | Generation | Use a generative model |
Use cases
TypeSafe’s docs include cookbooks: full working examples you can point Claude Code at. These four are the ones worth opening first:
- Guardrails for LLMs. One Jev call screens every message going into and out of an LLM app. Nouls check for hazards like a jailbreak attempt, and a Score rates how much harm complying would do. Your code decides whether to pass, review, or block.
- Double-checking citations. One Choice checks whether a quote’s source actually supports the claim, and low confidence flags it for a human.
- Classifying RAG passages. Score every retrieved passage before it reaches the answering model, and drop the ones carrying a hidden instruction or prompt injection.
- Skill suggestion. Against the 182 skills in Nous Research’s Hermes catalog, a single request ranks every skill and asks whether the turn needs one at all. The agent loads only the winner.
Community applications
Beyond classifying and ranking, people have already built:
- jev-ultrafast: decides which DOM element to click in a browser.
- jev-router: a model router that saves latency and tokens.
- winnow: context filtering, similar in spirit to RTK. Where RTK compacts context, winnow uses Jev to decide whether a piece of context is relevant before passing it back to your LLM.
- Jev-as-a-Judge: Replacing LLM-as-a-Judge for cheaper and faster evals.
More are listed on jevai.org/apps [6] and in awesome-jev.
Open-source alternatives
The idea behind Jev isn’t fundamentally new, and several open-source alternatives appeared within two weeks of its announcement [7] (star counts at time of writing):
- Laya (19,301★, Apache-2.0): 421M parameters on ModernBERT-large,
pip install laya, 32.8–39.5 ms per question on a T4. Its base checkpoints score near chance zero-shot, so it’s a base to fine-tune, not a drop-in replacement, though it claims to slightly beat Jev on some applications. - Kev (5,385★, Apache-2.0): LoRA adapters on Qwen3.5 at 0.8B, 4B, and 9B, serving TypeSafe’s
/v1/systemonecontract so the official SDK works unchanged. - SemIf (4,023★, MIT; formerly OpenJev): reads option logits off a frozen Qwen3.5-4B, taking 1.023 s for 21 decisions versus 5.332 s for the same decisions via generated JSON.
- NanoJev (2,086★, MIT): a 0.6B model for real-time control loops, beating Jev 128/128 to 56/128 on ViZDoom Basic. Not a general classifier.
- jevlike (1,255★, MIT): ships a trainer, not a model. Bring labeled options and train an option-attention head on a laptop.
The caveat: on an independent 49-task classifier benchmark, Jev scores 0.966 macro accuracy against 0.704 for the best open entrant [7]. These projects win on latency, price, and control, not out-of-the-box accuracy.
