veasy

Jev · TypeSafe · AI models · AI agents · research

Trying Jev: a model that picks an answer and tells you how sure it is

Jev from TypeSafe does not write text. It answers typed questions with probabilities. I used it to triage 111 prompt-audit findings and to score 17 song hooks. Here is what it got right and what it could not do.

Bar chart of Jev scores for song hooks. Emotion, 0 to 2: a hook I did not like from an earlier song 1.96, the four best new hooks 1.92, 1.90, 1.90 and 1.89. A cliché control line scored 0.42 for natural against an average of 1.26, and 0.73 for sounds like AI against an average of 0.50.

Most of the AI I use writes something: code, a reply, a plan. Jev, from TypeSafe, doesn't write anything. You give it some text and a few questions with fixed answers, and it gives back the answer plus how likely each option is. I tried it twice this week on real work. Here is what it is, the two tests with their numbers, and where it fell short.

What Jev is

TypeSafe calls Jev a "System One" model, after Daniel Kahneman's fast, intuitive System 1. You send a state (a string, a JSON object or an array of text) and one or more questions of three kinds:

  • Choice: pick one option from a list you define. You get the choice, a probability for every option, and a confidence number.
  • Score: place the state on levels you describe, for example 0 = calm, 1 = frustrated, 2 = very frustrated. The score can land between two levels.
  • Noul: a yes/no question. You get one number, the probability that the answer is yes.

The answer is always one of the options you gave it, so there is no JSON to fish out of a paragraph. The docs are clear that Jev is not a chat model and not something you plug into a coding agent. It is meant to sit inside your own code and make narrow, quick judgments while the code stays in charge.

Price is by input tokens only: $0.042 per million for jev-1.13.0, and output tokens are free.

One more thing from the docs that matters for me: English is the language it was trained on first. Other languages work but not equally well, and both of my tests below were in Vietnamese.

When I reach for it instead of an LLM

When the answer is a decision my code will act on, and I want to know how sure the model is before it acts. The docs suggest splitting confidence into three bands: act on your own when it's high, ask for confirmation when it's in the middle, and hand it to a person when it's low.

I don't use it for anything that needs writing or slow reasoning. The docs say to ask for "a judgment a knowledgeable person makes in a second", and to split bigger judgments into small questions and combine the answers in code.

Test 1: which fixes can my agents apply on their own?

On September 27 I ran a prompt audit on the Claude Code agents I use every day. The audit went through each agent's instruction files and reported problems: rules that point to tools that no longer exist, outdated notes, instructions that contradict each other. That came to 111 findings from 12 reports.

For each finding I asked Jev a Choice (apply automatically, ask me, or skip) plus a second question about how sensitive the change was. Then a rule in code decided:

  • apply automatically only if Jev picked "apply", with confidence of 0.7 or more, and the sensitivity was under 0.3;
  • never touch files shared by several agents without asking me.

Jev picked "apply" for 74 findings, "ask" for 27 and "skip" for 10. The code rule held back 37 of the 74: 30 because Jev's confidence was under 0.7, and 7 because of the other rules, mostly because they touched files shared by several agents. In the end 25 changes went in. One had to be dropped, because applying it would have made an agent's instructions say something that wasn't true. Everything else came to me as a short list, grouped by topic.

"I'm not sure" was a usable answer here: 30 of the 74 findings Jev wanted to apply came back with confidence under 0.7, and the rule sent those to me instead of applying them. I didn't measure the cost of this run.

Test 2: scoring song hooks

On September 29 I had 17 candidate hooks for a sad Vietnamese love song (the line of the chorus people remember). I also added 5 control lines, including one obvious cliché and one hook from an earlier song that I did not like.

For each line Jev answered five Score questions from 0 to 2 (natural, concrete, not a cliché, easy to remember, emotional) and one Noul: does it sound like AI wrote it? Separately, ACE-Step sang each hook and Whisper checked whether the words came out right, because a hook nobody can make out is useless no matter how it scores.

Two results stood out:

  • The cliché control, "Tim em tan vỡ vì anh mãi mãi" ("my heart is broken for you forever"), got the lowest naturalness score of all 22 lines (0.42) and the highest "sounds like AI" probability (0.73). Jev spots a worn-out line.
  • The hook I did not like from the earlier song got 1.96 for emotion, higher than every one of the 17 new hooks (the best was 1.92). Jev can tell whether a line carries emotion. It can't tell whether I will like it.

The whole run was 22 requests, about 22,000 input tokens (roughly a tenth of a cent at the list price) and about 2,000 output tokens, which aren't charged.

Where it helps and where it doesn't

It helps when:

  • my code needs a decision it can branch on;
  • I want the model to be able to say "I don't know";
  • I want to ask many small questions about the same text in one call.

It doesn't help with taste, with anything that needs writing, or with a judgment that takes real thinking. The docs also point out that calibration is measured across many predictions, so it doesn't guarantee that one particular answer is right. That's why the thresholds in test 1 matter as much as the model.

Next: Ăn Gì?

I'm building Ăn Gì? ("what should we eat?"), a small app where you type the ingredients you have and it suggests Vietnamese dishes. People type ingredients in many ways, with or without accents, and with regional names for the same thing. The plan is to check a dictionary first, and only when that fails, ask Jev a Choice among the standard ingredient names. If the top probability is too low, the app asks the user instead of guessing.

None of that is built yet, so I have no numbers for it. I'll write them up once the ingredient box is live.

Source: TypeSafe docs, in particular the pages on System One, questions, confidence and models.


ShareFacebookLinkedInX

Comments