Capabilities

Evaluating translation

Before trusting a model with a language, measure it. llamay evaltranslate has a model translate a set of sentences and scores the result against human references with chrF++ and BLEU, computed the way sacreBLEU computes them, so the numbers are comparable with published ones.

Run it

llamay evaltranslate -m qwen2.5:0.5b -src fr -tgt en -n 5
model          qwen2.5:0.5b
digest         sha256:c5396e06af294bd101b30dce59131a76d2b773e76950acc870eda801d3ab0515
file id        590bc10b059c7189
set            bundled:translate.jsonl (sha256:cb9b8b58740ebb7ebddf63b9fa5a5f742406f5c85dc87758b9757b2d70b489b1)
direction      fr -> en, 5 sentences, greedy (temperature 0), chat user turn

id  chrF++  BLEU   hypothesis
1   29.05   9.43   The train for Lyon departs at 7:30 AM.
2   89.16   76.92  Please send me the invoice by the end of this week.
3   34.89   16.59  The pediatric clinic opened a new pediatrics service the previous year.
4   84.33   65.80  My grandmother grows tomatoes and carrots in her garden.
5   45.04   14.76  The meeting was postponed to afternoon.

chrF++         57.45   chrF2++ nrefs:1|case:mixed|eff:yes|nc:6|nw:2|space:no
BLEU           39.52   BLEU nrefs:1|case:mixed|eff:no|tok:13a|smooth:exp

Sentence 1 is a good example of why the score matters: the reference says half past eight, and a 0.5B model wrote 7:30. The report names the exact model file by digest and the test set by digest, so a number can be reproduced later. The last column of each score line is sacreBLEU's signature: the settings that produced it.

Test sets

Twenty or thirty sentences tell you whether a model is usable in a direction; they do not rank two good models against each other. For that, use a few hundred sentences from your own domain.

Flags

FlagDefaultMeaning
-chattrueprompt through the model's chat template; -chat=false completes a raw "French: ...\nEnglish:" prompt, for base models
-gpurun the layer stack on the GPU when one is available; -gpu=false forces the CPU
-jsonwrite the report as JSON
-kv string"f32"KV cache precision: f32, vq8 (values quantised, keys exact), or q8 (both)
-kv-page int64positions per KV page
-lexicon stringnewline-separated word list for lexicon-constrained decoding
-m stringpath to a GGUF model (required)
-max int256maximum tokens generated per sentence
-n intscore only the first N sentences; 0 is all of them
-ngl int-1llama.cpp spelling of -gpu: 0 is CPU only, anything at or above the layer count is the whole model on the GPU
-no-repackkeep weights in their file format instead of rewriting them for a faster kernel (on by default for the CPU and unified memory, off for a discrete GPU)
-o stringwrite the report here as well as to standard output
-set stringtest set: a .tsv of source<TAB>reference, or .jsonl with src/ref or language-code keys (default: the bundled set)
-src stringsource language code, e.g. en (required)
-t intworker threads (0 picks a default that avoids efficiency cores)
-tgt stringtarget language code, e.g. fr (required)

Reading the numbers