Does the Cheap Decision Model Speak Your Customer's Language
TL;DR. In the previous article I put moderation of customer-written text on rung three - a calibrated decision model such as Jev - and quoted one caveat from its docs: English is primary, other languages are "handled but not equally well". I tested it: six models, 25 languages, 1,940 texts, a blind second-model check of every label, $1.84 in API calls. The result: in 25 languages Jev does not block ordinary customers (1 false positive in 360 clean texts). On real comments in eight languages it scores 94% - second of six, ahead of gpt-5-mini. On a short name field it does not hold: it catches a bare swear word about three times in four in English, Spanish, Russian, Japanese and Chinese, and about one time in five in the other sixteen languages. The boundary is the kind of text, not the alphabet. Below: the numbers, a routing recipe, and what this test does not prove.
The caveat nobody tested
The previous article sorted a store's AI calls onto four rungs: code, compiled rules, a calibrated decision model, a generative model. "Is there profanity in this field" and "is this review abusive" landed on rung three: a judgment with no spec, which a model like TypeSafe's Jev answers with a probability in a quarter of a second for less than a hundredth of a cent. The same article carried one boundary bullet taken from the vendor's docs: English is the primary training language; other languages, CJK included, are handled but not equally well; test on your own content and watch the confidence when routing. I did not test it then.
The reason to test it is practical. A Magento module I maintain on a storefront with 30 store views moderates one customer-typed text field with a chat model. Thirty store views are 25 languages and nine scripts. Almost every multilingual store has the same shape: a short free-text field (a name, a signature, an engraving) and longer customer text (a comment, a review). So the question is not "is Jev good". It is: what does a store lose if it puts a decision model on multilingual customer text instead of a chat classifier, in which languages, and does the probability show those losses clearly enough for code to route around them?
I found no published multilingual evaluation of this model class as of 2 October.
What I measured and how
Six models, all through one OpenRouter key except Jev itself:
jev-1.13.0- the typed decision model, version pinned so the thresholds mean something;gpt-5-mini- a chat model with a strict JSON schema: the class most stores run this job on today, and the control;gemini-2.5-flash-lite- the cheapest Google route with structured output;gemma-4-26b- open weights;llama-guard-4-12b- what people usually mean by "a safety classifier"; profanity is not a category in its hazard taxonomy;gpt-oss-safeguard-20b- a guard model that takes your moderation policy as text.
I tried one more guard, one with a Profanity category, and dropped it: it answers safe or unsafe with no probability, so there is nothing to route on.
Two tasks. The short field is a display name of up to 100 characters. The long text is a comment or review of 150-600 characters. The chat models and gpt-oss-safeguard get the same moderation policy, written for the test, not taken from production; Llama Guard gets only the text and answers in its own taxonomy.
Data from three sources, because no single one covers everything:
- Written texts, 435 of them. Language models wrote them from one specification as native text, not translation: clean names, trap names (the local Scunthorpe problem, a medical term, a homograph), polite reviews and angry-but-civil ones. The generator wrote the offensive half for five languages only (en, de, fr, es, it); for the other twenty it ran into its own safety filter, and I did not route around it.
- Names built from published profanity lists, 21 languages. Code drops a term into a name template, with nobody judging it. Every term appears twice - as is and disguised (leet, dots between letters, look-alike characters, a stretched letter, an asterisk) - and the same disguises are applied to clean names as a control.
- 960 real comments from published corpora labelled by people, eight languages from seven language groups: English, Romanian, Portuguese, Turkish, Greek, Arabic, Indonesian, Chinese.
The labels were checked by a second model, not by native speakers. GPT-6.1 Sol - a different model family from the one that wrote the texts - labelled every text blind in a fresh session: it saw only the text, the language and the same policy. The numbers below use only texts where its label agreed with the original one and the policy gives a clear answer. All 435 written texts passed. About half of the list terms did: the reviewer judged the rest harmless as a public name (childish words, ordinary nouns with a vulgar second sense) or ambiguous under the policy. A published profanity list is not a list of unacceptable names. Of the 960 comments, 520 passed. That is 1,292 verified texts in all.
Everything was measured on 2-3 October 2026. The code, the data and every model's raw answers are in the benchmark repository; it reproduces with two keys and one command.
Clean text: it does not block ordinary customers in any of the 25 languages
This is the firmest result, and it comes first because for a store a false positive costs more than a miss: a blocked customer does not place the order.
On 360 clean written texts in 25 languages and nine scripts - names, traps, calm and angry reviews - Jev was wrong once (a Swedish trap name). gpt-5-mini was also wrong once: it read an angry but civil French complaint as harassment. Flash-Lite and Gemma were wrong twice each, gpt-oss-safeguard four times, Llama Guard 26 times. Jev's median probability on clean text is 0.02-0.04 in every one of the 25 languages. If the question is "will the model start turning away ordinary people in my Thai store view", on this data the answer is no.
Comments: it holds outside English
520 verified comments, eight languages, 184 offensive and 336 clean:
| model | accuracy | caught of 184 | false positives of 336 | p50 | $ per 1,000 |
|---|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | 95.4% | 178 | 18 | 467 ms | 0.05 |
| Jev 1.13 | 94.0% | 175 | 22 | 250 ms | 0.05 |
| gpt-oss-safeguard-20b | 93.1% | 167 | 19 | 500 ms | 0.14 |
| Gemma 4 26B | 91.7% | 150 | 9 | 1,200 ms | 0.04 |
| gpt-5-mini | 89.0% | 184 | 57 | 2,180 ms | 0.48 |
| Llama Guard 4 | 60.6% | 152 | 173 | 376 ms | 0.05 |
Three things here I did not expect.
The drop from English is small. 98.5% in English, 93.4% in the other seven - five points. All nine of Jev's misses sit at probabilities of 0.16-0.49: the model was in doubt, not confidently wrong. It identified the comment's language correctly in all 520 cases.
The most expensive chat model is too strict for a storefront. gpt-5-mini caught all 184 offensive comments - and also 57 of the 336 acceptable ones, one in six, mostly as "harassment". That is the angry-but-civil complaint from the previous section, at scale. For a store moderating reviews, this classifier cuts honest criticism.
Jev's doubt sits where the policy is in doubt. The reviewer separately marked 186 comments the policy does not settle. 39% of them fell into Jev's uncertain band (0.3-0.7), against 14% of the clear ones. The band is not noise: it is roughly where a careful moderator would also pause. On those ambiguous comments gpt-5-mini flags 89%, Jev 65%, Flash-Lite 51%.
One caveat on the table, so it does not look more precise than it is: verified means the half of real text where the reviewer and the corpus annotators agree, which is the clearer half. On the disputed half the models diverge more.
A short name: here it falls behind, and not by alphabet
A short field is a different task. A name has no context: one word, sometimes disguised. On written offensive names, where the swear word sits in a phrase, Jev catches 28 of 30. On a bare swear word from a list, it does not:
| language group | Jev at 0.5 | Jev at 0.2 | gpt-5-mini | Flash-Lite | verified terms |
|---|---|---|---|---|---|
| English | 5 | 5 | 6 | 6 | 6 |
| Spanish, Russian | 8 | 9 | 10 | 10 | 11 |
| Japanese, Chinese | 8 | 9 | 9 | 8 | 10 |
| Korean | 1 | 2 | 2 | 4 | 4 |
| Arabic, Thai, Hindi (romanised), Vietnamese | 3 | 9 | 9 | 10 | 15 |
| German, French, Italian, Portuguese, Dutch | 4 | 9 | 14 | 13 | 22 |
| Polish, Czech, Hungarian, Swedish, Finnish, Turkish | 7 | 19 | 28 | 25 | 39 |
In English, Spanish, Russian, Japanese and Chinese Jev catches 21 of 27 verified terms (gpt-5-mini 25). In the other sixteen languages it catches 15 of 80 (gpt-5-mini 53). I expected the break to follow the script: Latin holds, Cyrillic, Arabic and CJK drop. It went the other way: Russian and Chinese hold next to English, while most of Latin-script Europe - German, French, Dutch, Hungarian, Finnish, Turkish - does not. The docs' "not equally well" is real, but it does not follow the alphabet. Each language has 2-10 verified terms, so only the group totals carry weight, not single cells.
Disguise adds a second problem. On written disguised names Jev catches 13 of 25 (English 5 of 5, German 1 of 5); gpt-5-mini catches 25 of 25. Worse, Jev's probability rises with the look of a disguise itself. Put leet or dots on a perfectly clean name and its probability goes up by 0.16 on average (0.32 in English). gpt-5-mini flags none of the 125 disguised clean names: it reads through the disguise. Jev partly reacts to how the text looks.
The control is not a ceiling either: gpt-5-mini misses 29 of 107 verified terms and is uneven itself - 3 of 7 in Italian, 5 of 10 in Turkish, 2 of 6 in Vietnamese.
The probability: where you can put a threshold on it
The main argument for a decision model on rung three is not the price but the probability code can branch on. On comments it works as promised. On short names its doubt is expensive to buy back.
Jev's median probability on a verified term is 0.29; on a clean name, 0.03. So the model is more often unsure than blind. Lower the threshold and it catches more. But disguised clean names rise along with the terms:
| flag if p is at least | terms (107) | disguised terms (105) | clean names (125) | disguised clean names (125) |
|---|---|---|---|---|
| 0.5 | 36 | 37 | 0 | 7 |
| 0.3 | 53 | 66 | 0 | 29 |
| 0.2 | 62 | 91 | 2 | 55 |
| 0.1 | 87 | 102 | 12 | 81 |
So on a short field a threshold alone settles nothing. A pair does: Jev answers where it is sure, and a chat model takes the uncertain band. The documented habit of a 0.3-0.7 band is too narrow here - most misses sit below 0.3.
On comments, the choice of the second model matters more than the choice of band. Jev plus gpt-5-mini on the 0.3-0.7 band gives 91.7% - worse than Jev alone, because gpt-5-mini's strictness comes with it. Jev plus Flash-Lite on the same band gives 95.4% with 14% of traffic sent to the second model, at $0.055 per thousand decisions. That is exactly Flash-Lite's accuracy alone and almost its price; the pair wins not on accuracy but because 86% of decisions come back in 250 ms with a probability code can log and branch on.
A recipe for a multilingual store
| surface | evidence | what to run |
|---|---|---|
| comments, reviews | 520 verified real comments, 8 languages | Jev first, Flash-Lite on the 0.3-0.7 band. Not gpt-5-mini: it cuts one acceptable comment in six |
| short name: en, es, ru, ja, zh | written names and list-derived names | Jev first, gpt-5-mini on the 0.1-0.7 band. On list-derived names in these languages the pair reaches gpt-5-mini's level: 98 of 104, against 84 for Jev alone. In English, 46 of 47 |
| short name: the other 16 languages | list-derived names | a chat model: gpt-5-mini for recall, Flash-Lite for price. Jev catches about one bare swear word in five here |
| comments in the 13 languages with no offensive comments in the set; any surface in Ukrainian | clean text only | Jev will not block an ordinary customer; whether it catches abuse there is not measured |
How much traffic goes to the chat model depends on what your customers write. On ordinary
clean names the 0.1-0.7 band takes 12 of 125, about one in ten. Names in the xX_Name_Xx
style will go more often: the model reacts to the look of a disguise. Measure that on your
own log before you budget for it.
There is also something the tables do not show. One noun in the prompt moves verdicts: when I replaced the field's name in the policy wording with another, 4 of Jev's answers changed on the same 435 texts, 5 of Gemma's, 1 of gpt-5-mini's. Every single number here carries that noise, and your policy wording is a variable of the test too.
What this test does not prove
- The labels are model labels, not native speakers'. A model wrote the written texts and another model checked them; the terms came from lists that people compiled, but the call "this is unacceptable in a name" was made only by the reviewer. Two models agreeing is not a native speaker: shared blind spots in low-resource vocabulary pass this check.
- Few texts per language. 2-10 verified terms and 11-42 offensive comments per language. Conclusions are drawn per language group, not per cell.
- Offensive comments in 12 of 25 languages: eight from corpora and four more (de, fr, es, it) on 20 written ones only. Russian, Ukrainian, Japanese, Korean, Hindi, Thai, Vietnamese and six European languages are not tested on long offensive text: I found no suitable corpus under an open licence. For Ukrainian there is no offensive data at all, only clean text.
- The reviewer is an OpenAI model, like gpt-5-mini and gpt-oss-safeguard. On comments gpt-5-mini still came out the weakest of the chat models, so the reviewer did not reward its own family's strictness, but it is worth keeping in mind.
- The thresholds were fitted on the same texts they were scored on. There is no held-out set; on your own data the band has to be fitted again.
- A snapshot, not a verdict.
jev-1.13.0is pinned; the other models were called by name through OpenRouter and are not pinned. TypeSafe says many of the model's jagged edges will be fixed in later versions. Everything was measured on 2-3 October 2026. - Flash-Lite's calibration was not tested. On its 105 wrong answers its self-reported confidence was 0.80-1.00, median 0.90 - the prompted-for confidence problem from the previous article.
What it means for rung three
The previous article's advice survives the test, with a refinement. For comments and reviews - text written in sentences - a calibrated decision model works as the first filter in every language where this could be tested, and its uncertain band lines up with where the policy itself is in doubt. For a short name field it is not yet the only filter beyond a handful of languages: a bare swear word and a disguised word are exactly where it lacks context, and its probability reacts to how the text looks.
The vendor said test on your own content. This is that test, and its most useful result is not a ranking of models. It is that the boundary does not run between English and the rest of the world. It runs between a sentence and a lone word.
Comments
No comments yet. Be the first to share your thoughts.
Sign in to leave a comment. Only registered readers can comment.