ESTIMATIONS.md

May 19, 2025 ยท View on GitHub

Average results for translation quality

BLEU (bilingual evaluation understudy) is an automatic algorithm for evaluating the quality of text which has been machine-translated from one natural language to another.

COMET is also an automatic algorithm for evaluating the quality of translated text, based on neuronets.

From author: BLEU was introduced in 2002, and COMET in 2020. As for now, I recommend COMET scores more than BLEU.

Use this results just for reference.

BLEU scores

Higher is better, no_translate can be used as baseline. Average on 100 examples from FLORES, offset = 150:

fra->engeng->frarus->engeng->rus
no_translate3.983.90.570.56
libre_translate47.6649.6232.4330.99
fb_nllb_translate nllb-200-distilled-600M51.9252.7341.3831.41
fb_nllb_translate nllb-200-distilled-1.3B56.815546.0333.98
fb_nllb_ctranslate2 JustFrederik/nllb-200-3.3B-ct2-float1654.8756.7348.4536.85
fb_nllb_ctranslate2 JustFr-ik/nllb-200-distilled-1.3B-ct2-int856.1256.4546.0734.56
google_translate58.0859.9947.737.98
deepl_translate57.6759.9350.0938.91
yandex_dev----------46.0940.23
openai_chat gpt-3.5-turbo (aka ChatGPT)----------41.4930.9
koboldapi_translate (alpaca7B-4bit)43.5130.543214.19
koboldapi_translate (alpaca30B-4bit)---------------24.0
fb_mbart50 facebook/mbart-large-50-one-to-many-mmt-----48.79-----28.55
fb_mbart50 facebook/mbart-large-50-many-to-many-mmt50.2648.9342.4728.56
vsegpt_chat openai/gpt-3.5-turbo----------41.9331.12
vsegpt_chat openai/gpt-4----------44.1634.88
vsegpt_chat anthropic/claude-instant-v1----------41.8829.67
vsegpt_chat anthropic/claude-254.9156.0946.3834.13
opus_mt Helsinki-NLP/opus-mt-en-ru---------------30.41

LLMs with errors:

  • vsegpt_chat tiiuae/falcon-40b-instruct - a lot of fails
  • vsegpt_chat google/palm-2-chat-bison - "I'm not able to help with that, as I'm only a language model."
  • 'koboldapi_translate' on 'eng->rus' pair average BLEU score: 7.00: 80/100 on IlyaGusev-saiga_7b_lora_llamacpp-ggml-model-q4_1.bin, may be adjusting for input prompt needed

COMET scores

SOTA for opensource realization:

  • multi_sources vsegpt_chat:lizpreciatior/lzlv-70b-fp16-hf,fb_nllb_ctranslate2 - this comparable to DeepL and Google Translate (realistic)
  • vsegpt_chat nvidia/nemotron-4-340b-instruct - this require a LOT OF computational resources

Higher is better, no_translate2 can be used as baseline. Average on 100 examples from FLORES, offset = 150:

fra->engeng->frarus->engeng->rus
no_translate231.6632.0633.0325.58
no_translate79.270.1969.344.82
opus_mt Helsinki-NLP/opus-mt-en-ru---------------82.22
libre_translate86.6682.3680.3683.34
lingvanex87.9286.9984.7586.3
bloomz bigscience/bloomz-1b787.8684.1----------
vsegpt_chat recursal/eagle-7b87.1883.6784.5675.94
koboldapi_translate NikolayKozloff/ALMA-13B-GGUF----------84.6487.92
t5_mt utrobinmv/t5_translate_en_ru_zh_large_1024----------86.0586.53
fb_nllb_translate nllb-200-distilled-1.3B89.0187.9586.9188.57
fb_nllb_ctranslate2 JustFrederik/nllb-200-3.3B-ct2-float1688.7488.3287.2588.83
vsegpt_chat mistralai/mixtral-8x7b-instruct88.4587.286.9487.85
vsegpt_chat mistralai/mistral-small-24b-instruct-250189.188.8187.2888.83
vsegpt_chat lizpreciatior/lzlv-70b-fp16-hf88.6987.1786.9188.15
multi_sources vsegpt_chat:lizpreciatior/lzlv-70b-fp16-hf,fb_nllb_ctranslate289.1488.2287.2289.87
vsegpt_chat OMF-R-Vikhr-Nemo-12B-Instruct-R-21-09-24---------------87.93
google_translate89.6988.987.5389.63
deepl89.3989.2787.9389.82
vsegpt_chat google/gemini-2.5-flash-pre---------------88.41
vsegpt_chat google/gemma-3-4b-it---------------88.5
vsegpt_chat google/gemma-3-12b-it---------------89.64
vsegpt_chat google/gemma-3-27b-it---------------89.51
vsegpt_chat google/gemma-2-27b-it---------------89.07
vsegpt_chat google/gemma-2-9b-it---------------88.32
vsegpt_chat meta-llama/llama-3.1-405b-instruct---------------89.72
vsegpt_chat meta-llama/llama-3.1-70b-instruct---------------89.93
vsegpt_chat meta-llama/llama-3.1-8b-instruct---------------86.98
vsegpt_chat meta-llama/llama-3-70b-instruct---------------88.84
vsegpt_chat meta-llama/llama-3.3-70b-instruct---------------89.69
vsegpt_chat meta-llama/llama-4-scout---------------89.17
vsegpt_chat meta-llama/llama-4-maverick---------------89.91
vsegpt_chat deepseek/deepseek-chat v3---------------90.22
vsegpt_chat deepseek/deepseek-chat-0324-alt-fast---------------90.04
multi_sources google_translate,vsegpt_chat:meta-llama/llama-3.1-70b-instruct---------------90.35
vsegpt_chat openai/gpt-4o-mini89.3888.4587.3189.55
vsegpt_chat cot_openai/gpt-4o-mini---------------88.93
vsegpt_chat anthropic/claude-instant-v1----------85.7388.13
vsegpt_chat openai/gpt-3.5-turbo----------86.8788.76
vsegpt_chat openai/gpt-3.5-turbo-instruct----------85.2387.46
vsegpt_chat openai/gpt-4----------87.0289.54
vsegpt_chat openai/gpt-4-1106-preview---------------89.85
vsegpt_chat openai/gpt-4-turbo---------------89.76
vsegpt_chat openai/gpt-4o---------------90.06
vsegpt_chat openai/gpt-4.1-nano---------------89.57
vsegpt_chat openai/gpt-4.1-mini---------------89.89
vsegpt_chat openai/gpt-4.1---------------89.97
vsegpt_chat cohere/command-r-plus---------------89.45
vsegpt_chat qwen/qwen-2-72b-instruct---------------89.38
vsegpt_chat qwen/qwen-2.5-72b-instruct---------------87.85
vsegpt_chat qwen/qwen3-32b (no_think)---------------88.62
vsegpt_chat qwen/qwen3-235b (no_think)---------------89.25
vsegpt_chat nvidia/nemotron-4-340b-instruct---------------90.07
vsegpt_chat anthropic/claude-289.2789.1787.4789.85
vsegpt_chat anthropic/claude-3-haiku---------------89.5
vsegpt_chat anthropic/claude-3-5-haiku---------------89.64
vsegpt_chat anthropic/claude-3-sonnet---------------89.49
vsegpt_chat anthropic/claude-3-opus---------------90.75
vsegpt_chat anthropic/claude-3.5-sonnet-20240620---------------90.78
vsegpt_chat anthropic/claude-3.5-sonnet-20241022---------------90.62
vsegpt_chat cot_anthropic/claude-3.5-sonnet---------------89.94
vsegpt_chat google/gemini-flash-1.5-8b---------------88.34
vsegpt_chat google/gemini-flash-1.5---------------89.27
vsegpt_chat google/gemini-pro---------------89.69
vsegpt_chat google/gemini-flash-1.5 (002)---------------89.57
vsegpt_chat google/gemini-pro-1.5 (002)---------------89.55
multi_sources google_translate,deepl89.6689.8587.890.42
multi_sources google_translate,deepl,vsegpt_chat*89.6689.8587.7690.67
yandex_dev----------87.3490.27
multi_sources google_translate,yandex_dev----------87.6490.39
multi_sources deepl,yandex_dev----------87.6490.62
multi_sources google_translate,deepl,yandex_dev----------87.7490.63
multi_sources google_translate,deepl,yandex_dev,vsegpt_chat*----------87.7190.66
multi_sources deepl,yandex_dev,vsegpt_chat*----------87.6790.77
multi_sources deepl,yandex_dev,vsegpt_chat**---------------91.02
multi_sources deepl,yandex_dev,vsegpt_chat***---------------91.05
multi_sources deepl,yandex_dev,vsegpt_chat****---------------91.06

(yandex_dev represents Yandex.Translate service)

* vsegpt_chat with anthropic/claude-2 ** vsegpt_chat with anthropic/claude-3-opus *** vsegpt_chat:openai/gpt-4o,vsegpt_chat:anthropic/claude-3-opus **** vsegpt_chat:openai/gpt-4o,vsegpt_chat:nvidia/nemotron-4-340b-instruct,vsegpt_chat:anthropic/claude-3-opus,vsegpt_chat:anthropic/claude-3.5-sonnet

IMPORTANT: You interested how it will work on YOUR language pairs? It's easy, script already included, see "Automatic BLEU measurement" chapter.

Chain translation results (use_mid_lang plugin)

Chain translation allow to translate phrases with mid-language (usually English)

BLEU scores

jpn->rus
no_translate0
google_translate27.63
deepl27.48
use_mid_lang google_translate,deepl28.34
use_mid_lang google_translate,google_translate27.62
use_mid_lang deepl,deepl27.62

COMET scores

jpn->rus
no_translate56.85
fb_nllb_ctranslate2 JustFrederik/nllb-200-3.3B-ct2-float1686.46
google_translate87.93
deepl88.11
use_mid_lang deepl->yandex_dev87.56
use_mid_lang google_translate->deepl88.37
use_mid_lang google_translate->google_translate87.67
use_mid_lang deepl->deepl88.43
multi_sources google_translate,deepl88.88
use_mid_lang multi_sources->multi_sources*88.9
multi_sources google_translate,deepl,use_mid_lang**89.05
  • * multi_sources with "google_translate,deepl"
  • ** use_mid_lang with "google_translate,deepl"

vsegpt_chat with anthropic/claude-2 get a lot of fails ("Can't understand", "Can't translate")

More results on different multi_sources settings:

jpn->rus
multi_sources google_translate,deepl,use_mid_lang:google_translate->deepl,use_mid_lang:google_translate->google_translate89.05
multi_sources google_translate,deepl,use_mid_lang:google_translate->deepl,use_mid_lang:deepl->deepl89.15
multi_sources google_translate,deepl,use_mid_lang:deepl->deepl89.04
multi_sources****89.25

**** model: google_translate,deepl,use_mid_lang:deepl->deepl,use_mid_lang:google_translate->deepl,use_mid_lang:google_translate->google_translate,use_mid_lang:deepl->google_translate

Automatic BLEU and COMET estimation

There are builded package to run BLEU and COMET estimation of plugin translation on different languages - so, you usually can reproduce our results.

There are pretty simple estimation based on FLORES dataset: https://huggingface.co/datasets/gsarti/flores_101/viewer

To estimate:

  1. install requirements-bleu.txt
  2. setup params in run_estimate_bleu.py (at beginning of file)
  3. run run_estimate_bleu.py

RECOMMENDATIONS:

  1. debug separate plugins first!
  2. To debug, use less BLEU_NUM_PHRASES.

Settings params:

# ----------------- key settings params ----------------
BLEU_PAIRS = "fra->eng,eng->fra,rus->eng,eng->rus" # pairs of language in terms of FLORES dataset https://huggingface.co/datasets/gsarti/flores_101/viewer
BLEU_PAIRS_2LETTERS = "fr->en,en->fr,ru->en,en->ru" # pairs of language codes that will be passed to plugin (from_lang, to_lang params)

BLEU_PLUGINS_AR = ["google_translate", "deepl", "multi_sources:google_translate,deepl"] 
    # plugins to estimate, array
    # now you can run them in format "plugin:model", that works only if plugin support "on-the-fly" model change (usually YES for synthetic and online plugins, and NO for offline)

BLEU_NUM_PHRASES = 100 # num of phrases to estimate. Between 1 and 100 for now.
BLEU_START_PHRASE = 150 # offset from FLORES dataset to get NUM phrases

BLEU_METRIC = "bleu" # bleu | comet