Our free translator ran a small machine-translation model directly in the browser (a Bergamot / Opus-MT style model compiled to WebAssembly). The appeal is privacy: the text never leaves the device. The downside is quality. I finally stopped guessing and measured it on the live site with 30 phrases.
The test
30 phrases in both directions between English and Russian, grouped by type. I marked each result as correct, correct with small flaws, or wrong in meaning.
Result: 43% correct, 26% with small flaws, 30% with a changed meaning. Speed was fine: the first translation takes about 11 s (model download and load), after that about 1.7 s per phrase.
What worked and what did not:
Connected prose and business text: good, in both directions.
Technical text: it invented words ("driver" became "водитель" (a vehicle driver) instead of a software driver).
Idioms and casual speech: 5 out of 5 failed. "I'll hit you up" came out as "I will hit you", which reads as a threat.
Ambiguity: "the bank refused the loan" was inverted.
A legal sentence, "shall not disclose ... to any third party", came out with the direction of the prohibition reversed. For a translator that advertises use on contracts, that was the one that worried me most.
The pattern is simple: it fails exactly on what a random visitor tries first, an idiom or a casual phrase.
Trying LLMs as a replacement
I ran the same 30 phrases through 8 models over an API (240 requests). The whole benchmark cost about $0.013. The hard set was 17 phrases where the in-browser model had failed.
Model
Hard phrases right (of 17)
Median latency
Gemma 26B (instruction-tuned)
17
1.46 s
GPT-5.4 mini
15
1.06 s
Llama 3.3 70B
13
0.92 s
DeepSeek V4 Flash
13 (+2 empty)
2.85 s
Qwen3 30B A3B
11
1.96 s
Two things surprised me:
The cheapest model was the worst. Qwen3 30B had the lowest price and the lowest score. It translated "Klöße" as cutlets and "Vorspeisen" as first courses. "Newer and bigger" is not the same as "better for this task"; I had bet on Qwen before measuring and was wrong.
Reasoning models are a trap for translation. One of them returned empty text for all 30 phrases, because it spent the whole token budget on thinking and never reached the answer. Raising the limit just means paying for the most expensive tokens.
I shipped the winner. The cost came out to roughly $0.16 per million source characters, and it fixed the threatening "hit you" case and the inverted legal sentence.
The trade-off I had to accept
Moving translation to a server model means the text now leaves the device, which was the original selling point. The client falls back to the local model if the server call fails. If privacy is your main constraint, the small local model is still the right tool. If quality on everyday text matters more, it is not good enough.
You can compare for yourself: the browser translator and the English to Russian page. What phrases break your favourite translator?
Our free translator ran a small machine-translation model directly in the browser (a Bergamot / Opus-MT style model compiled to WebAssembly). The appeal is privacy: the text never leaves the device. The downside is quality. I finally stopped guessing and measured it on the live site with 30 phrases.
The test30 phrases in both directions between English and Russian, grouped by type. I marked each result as correct, correct with small flaws, or wrong in meaning.
Result: 43% correct, 26% with small flaws, 30% with a changed meaning. Speed was fine: the first translation takes about 11 s (model download and load), after that about 1.7 s per phrase.
What worked and what did not:
The pattern is simple: it fails exactly on what a random visitor tries first, an idiom or a casual phrase.
Trying LLMs as a replacementI ran the same 30 phrases through 8 models over an API (240 requests). The whole benchmark cost about $0.013. The hard set was 17 phrases where the in-browser model had failed.
| Model | Hard phrases right (of 17) | Median latency |
|---|---|---|
| Gemma 26B (instruction-tuned) | 17 | 1.46 s |
| GPT-5.4 mini | 15 | 1.06 s |
| Llama 3.3 70B | 13 | 0.92 s |
| DeepSeek V4 Flash | 13 (+2 empty) | 2.85 s |
| Qwen3 30B A3B | 11 | 1.96 s |
Two things surprised me:
I shipped the winner. The cost came out to roughly $0.16 per million source characters, and it fixed the threatening "hit you" case and the inverted legal sentence.
The trade-off I had to acceptMoving translation to a server model means the text now leaves the device, which was the original selling point. The client falls back to the local model if the server call fails. If privacy is your main constraint, the small local model is still the right tool. If quality on everyday text matters more, it is not good enough.
You can compare for yourself: the browser translator and the English to Russian page. What phrases break your favourite translator?
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | OpenAI shipped Dots on Tuesday. By Wednesday night my terminal had its own, running 100% locally | 0 | 10 | 30-09-2026 |
| 2 | Новости науки свайпами: serverless на AWS, LLM за цент на статью и 12 тестировщиков для Google Play | 0 | 8.77 | 27-09-2026 |
| 3 | I updated our drug profile to 563-drug harm-reduction library in 9 languages on Cloudflare's free tier — here's the stack | 0 | 11.79 | 30-09-2026 |
| 4 | Daily Hacker News for 2026-09-22 | 0 | 10.96 | 23-09-2026 |
| 5 | Google Releases Gemma 4 While Gemini 4 Argon Signals Build | 0 | 10.87 | 30-09-2026 |
| 6 | LLM уверенно называет шахматные ходы, которых на доске нет. Как я проверяю каждый ее ответ кодом | 0 | 5.87 | 26-09-2026 |
| 7 | Daily Hacker News for 2026-09-20 | 0 | 9.38 | 21-09-2026 |
| 8 | Daily Hacker News for 2026-09-29 | 0 | 11.19 | 30-09-2026 |
| 9 | The AI Inference Revolution Is Here | 0 | 15.2 | 15-09-2026 |
| 10 | TensorFlow.js in the browser: why one new tensor shape cost 8-17 seconds, and how I cut a 40 s freeze | 0 | 8.59 | 30-09-2026 |