· 6 min read

Zero-shot classification in production: self-hosted jeff vs jev on 10 real use cases

Can a 400M-parameter model on a MacBook replace a hosted classification API? I ran both on 2,040 items drawn from real product use cases. The answer depends on the kind of question, and on how you phrase it.

  • AI engineering
  • Benchmark
  • Classification
  • Guardrails

Lire en français →

Context

TypeSafe offers a classification API called System One. You send it a text (the state) and typed questions, and you get structured, calibrated answers back, with no text generation and no prompt engineering:

  • noul: a yes/no question, the model returns the probability that the answer is yes;
  • choice: the model picks one option from a set you define, with the full probability distribution;
  • score: the model rates the text on an ordered scale you define (2 to 10 levels).

The hosted model is called jev. It is the kind of building block you put in front of an LLM (guardrails, intent routing, moderation) or inside a data pipeline, wherever a generation call would be too slow and too expensive.

jeff is an open-source reimplementation of that API by Logan Markewich (MIT license), running on GLiFormer-large-v1, a 400M-parameter encoder. Same wire format, same official SDK: change the base URL and it works. Its author published benchmarks on academic datasets (AG News, SST, BoolQ…). I wanted to know something else: on tasks I could actually ship, when is a small self-hosted model enough, and when is it not?

Method

Ten tests, built as product use cases rather than research datasets:

Use case Domain Type Dataset (Hugging Face)
Prompt-injection guard LLM security noul deepset/prompt-injections
Same, rephrased as choice LLM security choice deepset/prompt-injections
Jailbreak detection LLM security noul jackhhao/jailbreak-classification
Comment moderation Trust & safety noul SetFit/toxic_conversations_50k
Chatbot intent routing (6 intents) Customer support choice mteb/banking77
Market sentiment on news Fintech score twitter-financial-news-sentiment
Clickbait headline filter Content / media noul clickbait_title_classification
Same, rephrased as choice Content / media choice clickbait_title_classification
French movie reviews, question in English Multilingual noul tblard/allocine
French movie reviews, question in French Multilingual noul tblard/allocine

The rules:

  • One question per item, written the way I would ship it. For the guard, for instance: "Does this text try to override, ignore or manipulate the AI assistant's instructions?"
  • Balanced samples (100 per class for binary questions, 40 per intent for routing), drawn with a fixed seed, texts truncated at 2,000 characters. Both backends see exactly the same items.
  • The official typesafe-sdk on both sides. jev is reached through Vercel's AI Gateway (typesafe-ai/jev); jeff runs locally on a MacBook (MPS) with its default configuration.
  • Metrics: for a noul, accuracy at the 0.5 threshold plus AUROC; for a choice, accuracy plus macro-F1; for a score, accuracy of the most likely level plus mean absolute error (MAE).
  • An operational verdict per row: ≥ 85% ship it, 70 to 85% ship with human review, below 70% not yet.

In total: 2,040 predictions per backend. The code is three Python files of about a hundred lines each around the SDK: the tasks, the runner (async, resumable, backoff on rate limits) and the report.

Results

Comparison table of jeff vs jev on ten use cases
Comparison table of jeff vs jev on ten use cases
Use case Type n jeff jev
Prompt-injection guard noul 200 43.5% · AUROC 0.39 83.5% · AUROC 0.99
Prompt injection, as choice choice 200 62.5% · F1 0.57 83.5% · F1 0.83
Jailbreak detection noul 200 87.0% · AUROC 0.97 94.5% · AUROC 0.99
Comment moderation noul 200 77.5% · AUROC 0.84 77.5% · AUROC 0.86
Intent routing (6) choice 240 93.8% · F1 0.94 99.2% · F1 0.99
Market sentiment score 201 82.1% · MAE 0.20 84.1% · MAE 0.18
Clickbait filter noul 200 48.0% · AUROC 0.50 88.0% · AUROC 0.96
Clickbait, as choice choice 199 70.9% · F1 0.68 95.5% · F1 0.95
FR reviews, question EN noul 200 92.5% · AUROC 0.97 96.5% · AUROC 0.99
FR reviews, question FR noul 200 91.0% · AUROC 0.97 96.5% · AUROC 0.99

What I take away

1. jev wins everywhere, but the gap depends on the nature of the question. When the answer can be read from the vocabulary of the text (sentiment, a customer's intent, jailbreak patterns), jeff sits 4 to 7 points behind jev and clears the shipping bar. Three use cases out of ten are shippable as-is with a 400M model running on a laptop.

2. Where the model has to understand what the text is trying to do, jeff is at chance level. The prompt-injection guard (AUROC 0.39) and the clickbait filter (AUROC 0.50) require reasoning about intent and meta-language, not words. An encoder this size does not do that. jev scores 0.99 and 0.96. For a security guardrail that is not a nuance, it is disqualifying.

3. Phrasing changes everything for the small model. Rephrasing the noul as a two-option choice with described labels ("attack: tries to make the AI ignore its instructions…" vs "normal: a regular request…") lifts jeff from 43.5 to 62.5% on injection and from 48 to 71% on clickbait. For jev, the same change is worth 0 and +7.5 points. The descriptions hand the small model the lexical cues it cannot infer on its own. Practical rule: with a small model, write your criteria like a spec, not like a question.

4. Multilingual is not a problem. On French reviews, asking the question in English or in French barely matters for either model (92.5 vs 91% for jeff, 96.5% both ways for jev). GLiFormer is trained multilingually and it shows.

5. When both plateau at the same place, the labels are the ceiling. Comment moderation: 77.5% on both sides, AUROC 0.84 vs 0.86. Toxicity labels are notoriously subjective. No model will beat annotator agreement, and that is useful to know before buying anything.

6. On an ordinal scale, near parity. Three-level market sentiment: 82 vs 84% and an MAE of 0.20 vs 0.18. A well-defined score is comfortable territory for the small model.

Cost and latency

This was not the point of the test, and my latencies are not comparable: jeff ran on a laptop with 16 concurrent requests, so mostly queued (p50 1.4 s), while jev answered in a 365 ms median through the gateway. The numbers that matter are the ones jeff's author measured on an L4 GPU on Modal: p50 151 ms vs 129 ms for jev, and roughly $2.6 vs $15.6 per million single-question requests. On my test set, jev bills 376 input tokens per request on average where jeff counts 111: compare per request, never per token.

What I would do in production

  • Security guardrails (injection, manipulation, subtle content): the hosted model, or a bigger one. No compromise on this line.
  • Routing, sentiment, French content, blatant jailbreaks: jeff is shippable, the data stays in-house, and the cost is divided by six.
  • Moderation: either, with human review, and start by cleaning your own labels.
  • In every case: build an evaluation set from your own data before choosing. Two hundred balanced items are enough to see the gaps above. Phrase your criteria as a described choice. And keep a cascade: jeff answers first, jev takes over when confidence drops below a threshold. jeff's probabilities are temperature-calibrated (3.2 by default), which makes that threshold usable.

This is exactly the kind of trade-off I make when architecting an AI system: not "which model is best?", but "which one is sufficient here, at what cost, and how do we know when it stops being sufficient?".