Skip to main content
State and typed questions go into Jev, typed answers with confidence come out, and application code decides whether to act, confirm or escalate

TypeSafe Jev: A Model That Makes Decisions Instead of Writing Text

Jev returns typed answers with calibrated confidence in under half a second. Here is where a decision-only model belongs in a production AI system, and where it does not.

A model with no text output

On 15 September 2026, TypeSafe AI came out of stealth with $40 million in seed funding and released Jev, the first model in what it calls the System One class (TypeSafe announcement, The Register). CEO Diogo Almeida is a co-author of the InstructGPT paper behind ChatGPT.

Jev does not write text. You send it some state (an email, a log line, a support ticket, a JSON object) and a set of typed questions. It sends back typed answers. There are three question types:

  • Choice picks one option from a set you define, up to 255 options, and returns a probability for each.
  • Score rates the state against ordered levels you describe, and returns a probability for each level.
  • Noul answers a yes/no question with the probability that the answer is yes.

Choice and Score answers also carry a confidence value between 0 and 1. TypeSafe trains the model with a method it calls Reinforcement Learning for Calibrated Decisions (RLCD), which aims to make those probabilities match real hit rates.

For most production AI teams, that sounds narrow. But look at how many of your LLM calls exist only to produce a label, a score or a yes/no answer. Routing, triage, guardrails, eval grading and extraction checks usually go to a general-purpose model, then through a JSON parser and a retry loop. Jev is built for exactly those calls.

What the API looks like

A single request carries the state and any number of named questions. Using the Python SDK (docs):

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

with TypeSafeClient() as client:
    response = client.system_one(
        state={"ticket": ticket_text, "plan": customer.plan},
        questions={
            "team": Choice(
                instructions="Which team should handle this ticket?",
                criteria={
                    "billing": "Payment or subscription issues",
                    "technical": "Bugs or integration problems",
                    "sales": "Pricing or account questions",
                },
            ),
            "urgency": Score(
                instructions="How urgent is this ticket?",
                criteria=["can wait", "this week", "today"],
            ),
            "refund_request": Noul(instructions="Is the customer asking for a refund?"),
        },
    )

team = response.choices["team"]
if team.confidence < 0.5:
    send_to_triage_queue(ticket_text)
else:
    assign(team.choice)

The answer is already a typed value in your program. You don’t need a parser, a Pydantic validation step, or a retry when the model wraps the JSON in prose. Because output is not billed and several questions share one request, TypeSafe recommends asking speculative questions you may not need and letting your code decide which ones matter.

The headline claims, and how much to trust them

TypeSafe’s launch claims:

Claim TypeSafe’s figure
Latency 70 to 500 ms end to end
Price $0.042 per million input tokens, output free
Speed vs LLMs 40x to 200x faster on comparable tasks
Workflow evals 193.6x faster and 444.6x cheaper than frontier models
Type errors 0%, because schema conformance is guaranteed by design

The type-safety claim is real in a narrow sense. The model cannot return an option you didn’t define or a malformed object. TypeSafe itself says the 0% is “not empirical”, because it is guaranteed by construction. That removes a whole class of production failure. It does not make the answers correct. As Anthony Maio’s analysis puts it, Jev constrains the shape of the output, not the judgment. The model can still pick the wrong route or give a bad answer a high probability.

The benchmark claims need more caution:

  • TypeSafe designed the workflows it measured. The company says the multipliers are at “the higher end of real world gains”.
  • The reference answers came from other models. Accuracy was scored against the average of GPT-6 Astra and Claude Fable 5.1 answers, not against human-labelled ground truth (heise). That tells you how closely Jev agrees with frontier models, not how often it is right.
  • Quality varies by task. On invoice processing, Jev scored 61.8% against 79.1% for the strongest comparison model. A cheaper model that is wrong more often on an expensive decision is not a saving.
  • The price may not last. TypeSafe has said the launch price may be subsidised.

None of this means the claims are false. It means that for your workload they are hypotheses, and you should check them with your own evaluation set. Our LLM evaluation framework covers how to build one from production data.

Where a decision model fits

The pattern TypeSafe recommends, and the one we would use too, keeps code in control and gives the model narrow decisions. Good fits:

  • Intent and ticket routing. Classify, then hand off to deterministic code, a specialist LLM or a person.
  • Guardrails. Policy checks on inputs and outputs before an agent acts, at a latency you can afford on every request.
  • Agent control layers. Choosing the next tool, spotting loops, deciding when to escalate. The agent does the generative work and Jev makes the cheap, frequent control decisions around it.
  • Eval grading and data labelling at volume. At $42 per billion input tokens, labelling a large backlog becomes a routine batch job instead of a budget request.
  • Real-time systems. TypeSafe’s Doom demo ran at 10 queries per second for about $7 an hour, which gives a sense of what a sub-second decision loop costs.

Poor fits are anything that needs language as output: drafting replies, summarising, writing code, or explaining a decision to a user or an auditor. If a regulator or customer can ask “why was this declined?”, you still need a way to produce that explanation. A probability distribution won’t answer it.

Confidence changes how you design the workflow

Confidence is the most useful part of the release for production systems. Most teams gate LLM decisions by trying to parse a “confidence” value the model wrote in text, which is rarely calibrated. Jev gives you a probability distribution and a derived confidence on every Choice and Score answer, and TypeSafe’s docs recommend setting thresholds by risk:

  1. High confidence: act automatically.
  2. Medium confidence: act with a confirmation step or flag it for review.
  3. Low confidence: do not act; route to a person or a larger model.

Set the thresholds per action, not per model. Showing the wrong help article can use a low bar. Approving a refund needs a high one. Then check calibration against your own labelled data, because calibration can drift when inputs shift, policies change or someone crafts adversarial input. Thresholds tuned on one model version also need re-validating when you move to jev-latest’s successor, so pin the version the same way we recommend in our Gemini Flash routing guide.

How to evaluate Jev this month

  • List every LLM call in your system whose output is a label, a score or a boolean. That list is your candidate set. Our guide to reducing LLM inference costs shows how to attribute spend per call site so you can rank it.
  • Pick the highest-volume candidate and build a labelled set from real production traffic, with human labels rather than labels from another model.
  • Run Jev in shadow mode next to your current model. Compare accuracy, calibration (does 0.9 confidence mean right 90% of the time?), latency at your traffic levels and cost.
  • Switch over only where accuracy holds. Keep the existing model as the fallback for low-confidence answers.
  • Because the service is in early access, review data handling and residency (it is currently hosted on the US West Coast) before sending customer data to it.

Jev is a new component, not a new frontier model. For the many small decisions that make an AI product work, it could remove most of the cost, latency and parsing code. For everything that needs language, your LLM stays. If you want help finding which of your calls are really decisions, or building the evaluation to prove the switch, talk to Tensorplay about an AI architecture review.

Frequently asked questions

Answers to common questions about this topic.

What is TypeSafe Jev?
Jev is the first System One model from TypeSafe AI, released in early access on 15 September 2026. It does not generate text. You send it state and a set of typed questions, and it returns typed answers (a choice, a score or a yes probability) with probabilities and a confidence value that your code can use directly.
Does Jev really never hallucinate?
It cannot return an answer outside the schema you defined, so there are no invented fields, malformed JSON or options that do not exist. It can still pick the wrong option or be confidently wrong, so you still need evaluation on your own data and thresholds that match the risk of each action.
How much does Jev cost?
At launch TypeSafe priced Jev at $0.042 per million input tokens with output tokens free. The company has said the price may be subsidised, so budget with headroom rather than treating it as permanent.
Can Jev replace our LLM?
No. It gives up text generation entirely, so it cannot draft replies, summarise, write code or explain its reasoning. It is a replacement for the LLM calls in your system that only ever needed a label, a score or a yes/no answer.
September 2026 in AI for production teams: frontier prices falling, managed agent runtimes, agent security incidents and new AI laws

AI in September 2026: Cheaper Models, Managed Agents and New Rules

The AI news from September 2026 that matters to teams shipping AI: Claude Opus 5.5, GPT-6 Sol and Luna pricing, OpenAI's Agents API, Plugin4Shell, California's new AI laws and a migration checklist.

Release gate for adopting GPT-6 Astra: candidate, evaluate, shadow traffic, then route only where it wins

GPT-6 Astra in Production: What Changes for AI Engineering Teams

How to evaluate, route and govern OpenAI's GPT-6 Astra in production: pricing, refusal behaviour, reduced reasoning visibility, agent risk and a migration plan.

Router sending requests to Flash-Lite, Gemini 3.8 Flash, or a frontier model by measured task quality

Gemini 3.7 and 3.8 Flash: When the Cheap Tier Gets Good at Coding

Gemini 3.7 and 3.8 Flash for production AI teams: benchmarks, pricing that doubles after December 2026, deprecated sampling parameters and tiered routing.

Production AI Architecture

Turn prototypes into reliable, observable, and scalable AI systems.

Discuss your project

Model Optimization & Evaluation

Improve model quality, control costs, and establish repeatable evaluation systems.

Discuss your project