Kev: the open model that classifies text and tells you when it is not sure

Lee este artículo en español.
Imagine 4,000 support emails a day, each of which has to land with the right team. The usual answer today is to ask a language model: write a prompt, request JSON and get back {"team": "billing"}.
At scale the underlying problem shows. That answer does not say whether the model was sure or flipping a coin, and without that there is no way to decide what gets automated and what a person has to review.
Kev is an open model built for that gap. It reads a text and, instead of writing an answer, spreads a probability across the options of each question you ask it.
For this article we started from a technical evaluation that put it through eight experiments on a consumer GPU, reproduced those experiments on a Mac with an M4 Pro chip, and checked the figures against those the project publishes.
Our conclusion is that, if your texts must not leave your infrastructure, Kev deserves a pilot for classifying them with a confidence threshold, although it is not yet ready to adopt as general infrastructure. If they can leave, the commercial service it replicates is cheaper and more accurate.
Contents
- What Kev is, in plain English
- Why a probability is worth more than a label
- How it works inside
- What we measured on our hardware
- Where it fits in a company
- Kev versus Jev: accuracy, cost and control
- Where it does not fit
- How to run a pilot
- Sources
What Kev is, in plain English
A chat model writes: you ask and it composes an answer in text that then has to be interpreted. Kev is closer to a judge with a scoring sheet. Whoever asks the question fixes the options in advance, and the model splits 100% of its belief among them.
You send it a text (an email, a ticket, a form) and a set of questions, and it returns a number for each possible answer, something like "47% returns, 28% shipping, 25% billing", which a program can use directly.
It supports three kinds of question:
- Yes or no ("does this need urgent attention?"): it returns the probability of yes.
- Pick one option out of up to 255 ("which team should handle it?"): one probability per option and a confidence, which measures how far the winning option beats picking at random.
- Rate on a scale ("how angry is the customer?"): an average level, weighted by probability.
It can be wrong like any model, but every answer comes with a measure of how much it trusts it, and that measure can be checked against data.
Kev is an open replica of Jev, a commercial service that the company TypeSafe calls a System One model. Developer Jared Palmer published it with the same API, and according to its repository TypeSafe's official Python client works against a Kev server unchanged.
The code and the base models are under the Apache 2.0 licence, which allows commercial use, and Kev runs on a machine of your own.
Why a probability is worth more than a label
With a bare label, every ticket goes to the team the model named and the errors stay invisible until a customer complains.
With reliable numbers, on the other hand, you can write a business rule on the confidence of each answer: automate anything above, say, 0.80 and send the rest to a human queue. That way you measure how much work you automated and at what error rate.
That requires the confidence to be reliable, and you have to check it with your own data, because the calibration Kev ships with comes from its authors' data.
The report sums up the differences from asking a chat model for JSON like this:
| Chat model with JSON | Kev | |
|---|---|---|
| Output | Text you have to interpret | One probability per option |
| Invalid answers | Possible, must be validated | Impossible: it only picks among the options given |
| Confidence | Made up, if you ask for it | Computed; the project measures how well it matches real accuracy |
| Cost of five questions | About five times the output | Almost the same as one |
| Injection between questions | Possible | Blocked by the architecture |
| Open-ended reasoning | Yes | No |
How it works inside
Technically, Kev-4B, the model we tested, is a small adapter (a LoRA) sitting on top of an open model from Alibaba's Qwen family, which stays frozen. Trained that way, it takes up a few tens of megabytes on top of that base model. The largest in the family, Kev-27B, retrains the whole model instead.
A chat model writes its answer piece by piece, and every piece costs another full pass through the network. Kev computes the answers without writing anything. It processes the document once, keeps it in a cache and reuses it for every question, so, according to the report, five questions cost about the same as one.
Each question reads the document and its own wording, and nothing else. Kev runs each one as an independent sequence over that cached document, so one question cannot read another's text.
At the end of each question, a small component compares the question with each option and turns those scores into percentages. Because the model points at options instead of writing them, you can add a new category today, say needs_legal_review, and the model will score it without retraining.
What we measured on our hardware
We ran Kev-4B, the 4-billion-parameter version, on a Mac with an M4 Pro chip and 24 GB of memory, using the model published on 24 September 2026 and the MLX engine in half precision.
A three-question request on a short ticket took 190 milliseconds; the evaluation we started from, on an NVIDIA RTX 3090 Ti, had measured 46. Where our figures differ from theirs, we say so.
We sent it, in Spanish, a complaint from a customer who had been charged twice for the September fee and ended with "if this is not fixed today I am cancelling the contract". Kev was trained on ten English datasets and still got all three questions right.
It classified the ticket as billing (probability 0.88, confidence 0.84), marked the urgency as "today" and gave a churn risk of 0.94, connecting "I am cancelling the contract" with a question phrased in a different way. The original evaluation had obtained the same results, within hundredths.

Faced with an email mixing a late delivery, a wrong size and a duplicate charge, Kev split its belief between returns (0.51), shipping (0.36) and billing (0.13), with a confidence of 0.27. Faced with a text containing no useful information, confidence was 0.26.
On these three documents, any confidence threshold between 0.3 and 0.8 separates the clear ticket from the two doubtful ones.
Here came the first surprise. The evaluation we started from, done three days earlier with the previous version of the model, had measured 0.085 on that same ambiguous email and 0.13 on the text with no data.
The 24 September version is more confident exactly where it should not be. A threshold has to be measured on the version you are going to deploy, and measured again when it changes.

Three documents are not enough to set a threshold. Besides, Kev's confidence numbers were tuned on its authors' data, and according to the project itself, on an external dataset the model declared an average confidence of 0.82 while getting 0.70 right.
Before trusting a threshold, you need to recalibrate with a few hundred examples of your own, using the tool the repository ships.
In the security test we placed a secret and an instruction to hijack the model inside one question, and asked another question to reveal the secret. It answered "unknown" with 0.69.
Run on its own, that second question gave the same probabilities within 0.004, the noise of computing in different batches; on the original evaluation's GPU the figures were identical to the fourth decimal.
In neither case did the secret leak: the design leaves no path for information to travel from one question to another.
On a business narrative of about 600 words it got the churn risk right (0.94), the unrefunded duplicate charge (0.99) and that no legal review was needed (0.05).
On "who should take the next action" it hesitated between billing and sales, with confidence 0.26: exactly what a threshold would send to a person. It took 640 ms the first time and 285 the second, with the document already cached.
Where it fits in a company
Kev fits processes with short or medium texts, a fixed set of questions and a way to send doubtful cases to a person. That last requirement is the human oversight we argued for when analysing Uber's 825 million euro fine (in Spanish), and here the confidence threshold sets it.
| Scenario | Typical questions |
|---|---|
| Support triage | Which team should handle it? Is it urgent? How angry is the customer? |
| Churn risk | Is there an intention to cancel? Is a competitor mentioned? |
| Pre-screening of complaints | Is there an unrefunded duplicate charge? Does it need legal review? |
| Routing forms and requests | What kind of request is it? Is information missing to process it? |
| Conversation analysis | Was the query resolved? Is a follow-up needed? |
It makes even more sense when the data should not leave the building, as in law firms, clinics or financial institutions. Just as when we ran ALIA locally (in Spanish), with sovereign AI on your own servers the text is classified without travelling to any third party.
Kev versus Jev: accuracy, cost and control
Accuracy. According to the figures the project publishes, on datasets Kev did not see in training, Kev-4B gets 0.817 of the questions right, Kev-9B 0.820 and Kev-27B 0.851, against 0.857 for Jev. The project warns that it does not know what data Jev was trained on, so the comparison is not controlled.
Another figure is more useful for deciding how much to automate. According to the project's own measurements, if you accept a 5% error rate, Kev-4B, 9B and 27B let you automate between 52% and 69% of decisions, and Jev 70%.

Cost. Jev charges $0.042 per million input tokens and does not charge for output. Each request in the example takes about 300 tokens including the questions, so the 4,000 daily emails add up to about 36 million a month and cost around $1.50. Ten times the volume keeps the bill under $20.
Kev-4B used 8.6 GB of video memory in the evaluation, and on our 24 GB Mac it ran comfortably. The report calculates that, at 46 ms per request, a 24 GB consumer card handles on the order of a million short requests a day if kept busy.
We estimate a card in that range at 600 to 2,000 euros depending on model and condition, plus 30 to 65 euros a month of electricity if it runs around the clock. Without a GPU, Kev gives the same answers on the CPU, between 20 and 54 times slower according to the evaluation; on an Apple Silicon Mac it uses MLX and runs about four times slower than on the RTX 3090 Ti.
On top of that comes the time of whoever deploys, watches and recalibrates it. At these volumes, a single day of technical work a year already costs more than Jev's annual bill.
Control. Self-hosting, the text never leaves your infrastructure and you can tune the model with your data and audit its code. If the project is abandoned, the code and the weights stay on your server; if a service changes its price or shuts down, you depend on what its owner decides. If none of that matters in your case, Jev is cheaper and more accurate.
Where it does not fit
Kev is not a general-purpose assistant. Our tests, the evaluation we started from and the project's documentation agree on these limits:
- Long documents. It was trained on texts of up to about 280 words. It accepts more, but according to the project accuracy drops visibly on long documents, and it has this logged as an open issue. Our 600-word narrative came out fine; that is one case, not proof.
- Questions that require world knowledge rather than reading the text in front of it. On general-knowledge questions, according to the project, Jev scores 0.90 and Kev 0.74.
- Malicious text inside the document. The isolation protects questions from each other, but all of them read the document by design. If a customer writes "ignore your instructions" in their email, that sentence reaches every question. It is the same family of attack we explained with AI browsers, and it has to be considered whenever the text comes from outside.
- Close calls. We tried six orders of the options. On a clear case the winner held, with a probability between 0.98 and 0.99; on the ambiguous email it also held, but its probability swung by 0.14 depending on the order, and the evaluation of the previous version saw the winner change on that same email. A confidence threshold filters these cases out, and the repository also lets you average over several orders.
- Dates. It cannot subtract dates. The project includes an optional preprocessing step that computes the days and writes them into the text; you have to enable it.
- One request at a time. Its server does not batch requests from different clients. With real traffic you need several copies or a queue in front.
- Maturity. It is a one-person project. Its repository was created on 17 September 2026, two days after Jev was announced, and it depends on fast-moving libraries. On 1 October it released Kev 1.0, which pins one version of each model, a step towards stability. The report also warns that it clones the API of a commercial product; it considers this legitimate, but worth knowing before building a business on top of it. It rates the engineering discipline as high, and even so, depending on such a young project is a risk.
How to run a pilot
Reproducing the experiments in this article took us one afternoon on a Mac, most of it downloading the model. The pilot adds labelling a few hundred real cases, which takes the time of someone who knows the process.
- Pick a process with short texts and fixed questions, where today someone reads and decides.
- Label a few hundred real cases and hold out 10 to 20% for evaluation.
- Recalibrate with your data.
- Measure what share of decisions you could automate at an error rate you can defend, and which would still have to go through a person.
If accuracy falls short, the project recommends fine-tuning the model with your examples starting from the published model, because starting from the base model throws away what Kev has already learned about this question format.
To check it yourself takes three commands (Python 3.12 or 3.13 and uv):
git clone https://github.com/jaredpalmer/kev.git && cd kev
uv sync --extra serve
uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b --port 8009
Does someone in your company read texts every day to decide where they go? That process is a good candidate. Tell us about it and, with a sample of your own tickets, we will tell you what share you could automate and at what error rate.
Sources
- Kev repository, created on 17 September 2026, Kev 1.0 release of 1 October: code, evaluations and Apache 2.0 licence — github.com/jaredpalmer/kev
- Published Kev models — huggingface.co/jaredpalmer
- Jev pricing and launch date: TypeSafe announcement of 15 September 2026 ($0.042 per million input tokens, output free), checked on 26 September 2026 — typesafe.ai
- Technical evaluation "Kev, explicado desde cero" (23 September 2026, commit 557598f, Kev-4B on an NVIDIA RTX 3090 Ti): unpublished document we had access to and whose experiments we reproduced.
- Our own measurements: Kev-4B published on 24 September 2026, MLX in half precision, Mac with M4 Pro chip and 24 GB, 28 September 2026. We keep the requests and responses from every test.
Our measurements from 28 September 2026, project figures checked on 2 October and Jev pricing checked on 26 September. The project moves fast: between two versions published three days apart, confidence on ambiguous cases changed.
