Capabilities
Evaluating translation
Before trusting a model with a language, measure it. llamay
evaltranslate has a model translate a set of sentences and scores the
result against human references with chrF++ and BLEU, computed the way
sacreBLEU computes them, so the numbers are comparable with published ones.
Run it
llamay evaltranslate -m qwen2.5:0.5b -src fr -tgt en -n 5
model qwen2.5:0.5b
digest sha256:c5396e06af294bd101b30dce59131a76d2b773e76950acc870eda801d3ab0515
file id 590bc10b059c7189
set bundled:translate.jsonl (sha256:cb9b8b58740ebb7ebddf63b9fa5a5f742406f5c85dc87758b9757b2d70b489b1)
direction fr -> en, 5 sentences, greedy (temperature 0), chat user turn
id chrF++ BLEU hypothesis
1 29.05 9.43 The train for Lyon departs at 7:30 AM.
2 89.16 76.92 Please send me the invoice by the end of this week.
3 34.89 16.59 The pediatric clinic opened a new pediatrics service the previous year.
4 84.33 65.80 My grandmother grows tomatoes and carrots in her garden.
5 45.04 14.76 The meeting was postponed to afternoon.
chrF++ 57.45 chrF2++ nrefs:1|case:mixed|eff:yes|nc:6|nw:2|space:no
BLEU 39.52 BLEU nrefs:1|case:mixed|eff:no|tok:13a|smooth:exp
Sentence 1 is a good example of why the score matters: the reference says half past eight, and a 0.5B model wrote 7:30. The report names the exact model file by digest and the test set by digest, so a number can be reproduced later. The last column of each score line is sacreBLEU's signature: the settings that produced it.
Test sets
- The bundled set has 30 sentences in English, French, German, Spanish, Italian and Portuguese. Any pair of those works with no file.
- Your own, with
-set: a.tsvofsource<TAB>referencelines, or.jsonlwithsrc/refkeys or one key per language code.
Twenty or thirty sentences tell you whether a model is usable in a direction; they do not rank two good models against each other. For that, use a few hundred sentences from your own domain.
Flags
| Flag | Default | Meaning |
|---|---|---|
-chat | true | prompt through the model's chat template; -chat=false completes a raw "French: ...\nEnglish:" prompt, for base models |
-gpu | run the layer stack on the GPU when one is available; -gpu=false forces the CPU | |
-json | write the report as JSON | |
-kv string | "f32" | KV cache precision: f32, vq8 (values quantised, keys exact), or q8 (both) |
-kv-page int | 64 | positions per KV page |
-lexicon string | newline-separated word list for lexicon-constrained decoding | |
-m string | path to a GGUF model (required) | |
-max int | 256 | maximum tokens generated per sentence |
-n int | score only the first N sentences; 0 is all of them | |
-ngl int | -1 | llama.cpp spelling of -gpu: 0 is CPU only, anything at or above the layer count is the whole model on the GPU |
-no-repack | keep weights in their file format instead of rewriting them for a faster kernel (on by default for the CPU and unified memory, off for a discrete GPU) | |
-o string | write the report here as well as to standard output | |
-set string | test set: a .tsv of source<TAB>reference, or .jsonl with src/ref or language-code keys (default: the bundled set) | |
-src string | source language code, e.g. en (required) | |
-t int | worker threads (0 picks a default that avoids efficiency cores) | |
-tgt string | target language code, e.g. fr (required) |
Reading the numbers
- chrF++ compares character and word n-grams. It is the better of the two for short sentences and for languages with rich morphology.
- BLEU compares word n-grams over the whole set. It is included because it is the number most papers report.
- Neither is a judgement of meaning. Read the hypotheses: a high score can hide a wrong number, as sentence 1 shows.