Technical Guide

Jev vs Laya: Which AI Decision Model Fits Your Work

Both models answer typed questions instead of writing text. Jev is a paid API; Laya is free to run yourself. This guide sets out what each one actually is, what the published benchmarks do and do not show, and how to choose without being swayed by a headline number.

Short answer

Jev is TypeSafe's paid decision model; Laya is Convai Innovations' free open-weights model that runs on your own hardware. Pick Jev to start today, cover varied tasks without training, or choose among many options; pick Laya when data must stay on your servers, volume makes per-call fees add up, and you have engineers to fine-tune it. Laya's claim to be ten times faster and more accurate than Jev holds only under narrow conditions.

Jev or Laya: which should you choose?

Choose Jev if you want to start today, handle a wide range of tasks without training anything, or need to pick from many options at once. Choose Laya if data must stay on your own servers, request volume is high enough that a per-call fee adds up, and you have the engineers to fine-tune and run a model.

The claim that Laya is ten times faster and more accurate than Jev is true only under narrow conditions. The accuracy lead appears only after Laya was fine-tuned on the benchmark's own training data, and the speed figure compares a local GPU with a call over the internet. Test both on your own data before believing either.

What is a decision model, and why does it matter?

A decision model is an AI model that answers questions in a format you define, with a probability attached, instead of writing text. Much of what businesses call AI work is not writing at all. It is deciding: which team should get this support ticket, is this invoice complete, is this message complaining, which of these five categories does this document belong to. Using a chat model such as ChatGPT or Claude for that means paying for text generation you will throw away, then parsing the text back into an answer.

A decision model skips the text. You give it a state, such as an email, a ticket or a JSON record, and a set of typed questions. It returns an answer for each question with a probability attached. Nothing to parse, and no way for it to invent an answer that is not one of your options.

Both Jev and Laya use the same three question types: choice, which picks one item from a supplied list; score, which places the input on an ordered scale; and a yes-or-no check that returns the probability a statement is true.

What is Jev?

Jev is a paid model from TypeSafe, launched on 15 September 2026. One of the co-founders is Diogo Almeida, an author of the InstructGPT paper. TypeSafe calls it the first of its System One models, meaning fast, reflexive judgement rather than step-by-step reasoning.

Published price is USD 0.042 per million input tokens, with output not charged. The context window is about 32,000 tokens, a choice question can carry up to 255 options, and the vendor states end-to-end latency of 70 to 500 ms. It is reached through the TypeSafe console, OpenRouter or Vercel AI Gateway, using a decisions endpoint rather than a chat endpoint, so ordinary chat SDKs do not work with it.

What is not known matters as much. TypeSafe has not published the architecture, the parameter count or the weights. As of this writing, independent testing is thin. One reviewer planted seven defects across 37 documents; Jev caught six, where a frontier chat model caught all seven, at a small fraction of the time and cost.

What is Laya?

Laya is an open-weights model from Convai Innovations, with its technical write-up published on 18 September 2026, three days after Jev. It is released under Apache 2.0, so commercial use is allowed and it can run entirely on your own hardware.

It ships as three checkpoints. The English model is built on ModernBERT-large at 421 million parameters with a 512-token context. The multilingual model is built on mmBERT-base at 322 million parameters, covers more than 100 languages and has a 1,024-token context. A third checkpoint is tuned for typed-decision workflows. A router picks the right checkpoint by detecting the script of the input.

Laya was one of at least six Jev alternatives that appeared within 48 hours of Jev's launch. It drew the most attention because it published numbers that beat Jev, and it also drew the most scrutiny.

What do the benchmarks really show?

Laya's own README reports a win on typed-decisions, 0.766 to Jev's 0.727. That checkpoint was fine-tuned on the benchmark's training split. The base checkpoints, untuned, score about 0.36, below the 0.461 you would get by always guessing the most common answer. The lead is real for a model trained on that task; it says nothing about zero-shot use on your own work.

The speed figure, 32.8 ms against Jev's 236 to 276 ms for one question, compares Laya running locally on a Tesla T4 GPU with Jev reached as a hosted service. That is roughly a 7.8x gap, not the 10x repeated in headlines, and part of it is simply network distance. Your production latency depends on where your server sits and how you batch requests.

Laya does well on a few public datasets, such as AG News at 0.950 against 0.910 and DAIR Emotion at 0.595 against 0.480. It loses badly when the answer list is long: on Banking77, with 77 labels, Laya scores 0.425 while Jev scores 0.870. Laya's calibration figure of 0.081 is measured after temperature adjustment; the raw model is overconfident. Even Laya's own pages disagree on Jev's calibration score, which is a reason to treat every number here as provisional.

Jev and Laya side by side

This compares the things that decide whether a model can go into a real system, not only the benchmark scores. Figures come from each project's own published material and have not been independently verified.

Jev and Laya compared, as published in September 2026
JevLaya
MakerTypeSafeConvai Innovations
Released15 Sep 202618 Sep 2026
WeightsClosedOpen, Apache 2.0
Model sizeNot disclosed322M to 421M parameters
CostUSD 0.042 per million input tokens, output freeFree; you pay for the GPU and upkeep
Where data goesTypeSafe's serversStays on your own infrastructure
ContextAbout 32,000 tokens512 to 1,024 tokens
Many options at onceUp to 255, holds up wellWeak past about 50 without tuning
Works without trainingYes, general-purposeBase models near chance; fine-tune first
ThaiNot stated by the vendorMultilingual checkpoint only; test before use

Do Jev and Laya work for Thai?

Not proven yet. Neither project has published a Thai-specific result, so any claim that either one works well in Thai is a guess until you test it.

With Laya, the detail that matters most is that the English checkpoint collapses on non-Latin scripts. On Khmer it scored zero accuracy while reporting 95.2% confidence, which is worse than failing loudly because nothing warns you. Thai must go to the multilingual checkpoint. That one averages 0.451 across non-English languages on the MASSIVE intent benchmark, and 45 of 51 languages scored above three times random, which is usable for some jobs but far from settled.

Jev does not publish per-language results. The practical move is the same for both: take a few hundred real Thai cases from your own work, label them, and measure accuracy before anything goes live.

Which one fits which job?

Jev fits when you want a result this week, the tasks vary, and your volume is moderate. With no GPU to buy and no model to maintain, the per-call fee is small next to the cost of an engineer's time.

Laya fits when data cannot leave your organisation, for example medical records or financial documents covered by PDPA, or when volume runs to millions of decisions a month so that per-call fees stop being small, and the task is narrow enough to fine-tune well. Its fine-tuning notebook runs on two T4 GPUs in about four to five hours for 30,000 labelled questions.

Remember that free weights do not mean free operation. A GPU server, monitoring, updates, security and someone to review wrong answers all cost money. Compare the total cost per correctly handled case, not the price per request.

Most importantly, design the system so the model is swappable. This field produced six clones in two days; the best model in six months may be neither of these. Your labelled examples, your test set and a workflow that sends uncertain cases to a person are the parts that keep their value.

Before you pick either one

  • Is the job really a decision with a fixed set of answers, not writing?
  • Do you have a few hundred real, labelled examples to test against?
  • Are those examples in the language your users actually write in?
  • Is the data allowed to leave your organisation under PDPA and your contracts?
  • How many decisions a month, and what would the API fee be at that volume?
  • Who will run and update a self-hosted model, and what will the GPU cost?
  • What happens when the model is unsure: does the case go to a person?
  • Could you replace the model next year without rebuilding the system?

Frequently asked questions

What is Jev?

Jev is a paid AI decision model from TypeSafe, launched on 15 September 2026. It does not generate text. It takes an input plus typed questions and returns a choice, a score or a yes-or-no probability, each with a confidence value. It costs USD 0.042 per million input tokens with output free, and is available through TypeSafe, OpenRouter and Vercel AI Gateway.

What is Laya?

Laya is a free, open-weights decision model from Convai Innovations, released under Apache 2.0 three days after Jev. It answers the same kinds of typed questions and runs on your own hardware. The English model has 421 million parameters and the multilingual model 322 million, small enough to run on a single inexpensive GPU.

Is Laya really ten times faster than Jev?

Only under specific conditions. Laya answered one question in 32.8 ms on a local Tesla T4 GPU, while Jev took 236 to 276 ms reached as an online service. That is about 7.8x, and part of the gap is network travel rather than the model. Measure on your own infrastructure before relying on it.

Is Laya more accurate than Jev?

On some tests and not others. A Laya checkpoint fine-tuned on the typed-decisions benchmark scored 0.766 to Jev's 0.727, but Laya's untuned base models score about 0.36 on the same test. Jev is far stronger when there are many possible answers, scoring 0.870 to Laya's 0.425 on the 77-label Banking77 set.

Can Jev or Laya handle Thai?

Neither has published Thai results. Laya's English checkpoint fails on non-Latin scripts while still reporting high confidence, so Thai must use its multilingual checkpoint, which is usable for some jobs but unproven for Thai. Test either model on a few hundred real Thai cases from your own work before going live.

Should we use a decision model or a chat model like ChatGPT?

Use a decision model when the answer is one item from a fixed list, a score or a yes-or-no, and volume is high. It is faster, far cheaper and cannot invent an answer outside your options. Use a chat model when the output must be written text, or the task needs multi-step reasoning. Many systems use both: a decision model routes, a chat model drafts.