A claim about quality is worthless without a method.
So here is ours, published before the results, so we can’t redesign it once the numbers arrive.
Blind, not branded
Editors receive three renderings of the same chapter — ours, a human literary translation, and a raw machine output — unlabelled and in random order.
Working editors, not linguists
The readers are people who acquire and edit books in that language for a living. They are paid for the read and are not told who commissioned it.
Four axes
Voice fidelity, register accuracy, naturalness in the target language, and error count. Scored 1–5 with a written justification per axis.
Published whole
Every round is published in full, including the rounds where we place third. The moment we start selecting rounds, the number is worth nothing.
Round one is in progress.
The first round covers EN→DE and ES→EN, six chapters each, nine editors. Results and every scoring sheet will be published on this page when it closes — whatever they say. Until then, this section stays empty rather than filled with something reassuring.
Run on every manuscript, every time.
- Glossary adherencePercentage of bible terms rendered as specified across the manuscript.
- Expansion ratioTarget-to-source length per scene, flagged when it drifts outside the band for that pair.
- Untranslated fragmentsSource-language runs surviving into the target text.
- Numeric and date integrityEvery figure, date and measurement matched against the source.
- Back-translation divergenceA sample is translated back and compared for meaning loss.
- Paragraph structureCollapsed, merged or dropped paragraphs against the source structure.
None of these measure whether the book is good. They measure whether it is broken. The editor measures whether it is good.