Batch APIs, 2 of 26
Half the price.
7 times the wait.
The same 12 support messages classified twice, on gpt-4.1-mini. Once a call at a time, once through the batch API, and both numbers taken off the clock rather than out of a price list.
Every answer came back the same. The whole trade is time.
Exactly 2 times cheaper and 6.9 times slower, with 12 of 12 answer identical. The documented window is 24h and it finished in 83s, which is the number worth quoting because it is the one that happened.
At the size this is for
Twelve messages is a demonstration. A million of the same message, at the per message cost measured above, is $21 one at a time against $10 in a batch. That is the shape of the decision, and it is a multiplication rather than a measurement, so it is worth saying which is which.
A real million would not run as one batch either. It arrives in chunks, some of them fail, and the job has to know which chunks it already finished, which is the same reconciliation problem as any other load.
Four decisions
The turnaround is measured, not quoted
The documentation says up to twenty four hours, and this one finished in well under two minutes. Both numbers are true and only one of them is a measurement. Telling a client the documented figure and then watching it finish in ninety seconds is how you get asked to do the arithmetic again in front of them.
Identical answers, which is the part worth checking
Half price would be a poor trade if the output drifted. Every one of the twelve came back the same through both routes, and that comparison is in the script rather than in a sentence, because "it should be the same model" is an assumption and this is a check.
The discount is bought with latency, so it only fits work nobody waits for
Classifying a queue overnight, tagging a backlog, filling a column on ten thousand rows. None of those care how long one answer takes. Anything with a person at the other end cannot use this at any price, and knowing which side a job falls on is most of the skill.
The prices come out of the file the cost page uses
Not typed in here. Two pages quoting different numbers for the same model is the failure this whole site is arranged to prevent, and the way to prevent it is to have one place where a number lives.
What both runs said
The cost page holds the prices these are worked out from. The open-weight page is the other way of paying less. The load page is the reconciliation a real batch needs. The whole list is 41 requirements from 114 job posts.