Take a look at the API calls you'd use to build your own chatbot on top of any o...

littlestymaar · 2025-07-29T15:51:56 1753804316

Exactly, but caching doesn't work if you switch between providers in the middle of the conversation, which is my entire point.

majormajor · 2025-07-30T04:30:20 1753849820

If you're selectively faking things you don't care. You may not even be aware because the caching is transparent to you and you send the whole set of messages to the system each time either way. From the perspective of the person faking their model to look better than it is, it requires no special implementation changes.

And if you're faking your model to look better than it is, you probably aren't sending every call out to the paid 3rd party, you're more likely intentionally only using it to guide your model periodically.

littlestymaar · 2025-07-30T06:19:38 1753856378

> because the caching is transparent to you

It isn't when you look at your invoices though.

> aren't sending every call out to the paid 3rd party, you're more likely intentionally only using it to guide your model periodically.

I'd you do that, you're going to have to pay each token multiple times: both as inferred token on your model, and as input tokens on the third party and your model.

If the conversation are long enough (I didn't do the math but I suspect they don't even need to be that long) it's going to be costlier than just using the paid model with caching.