Providers of language models keep lowering their prices; most recently, a major vendor halved the rates of two models. Independent tests show, however, that the list price says little about the actual costs. A new model with a low rate needed around three times as many output units as a top model for the same test tasks and was therefore more expensive per completed task.

As long as companies use AI through flat-rate subscriptions, this stays invisible. As soon as models are built into processes such as pre-coding incoming invoices, the choice of model determines the running costs.

We recommend comparing models with a sample of real cases from your own process and recording the hit rate alongside the costs. In addition, switching models should be technically simple enough to take advantage of every price cut.