Didn't expect that. Luna pricing is crazy now. I don't think there is anything on the market that competes at this price-performance point.
For our production app, OpenAI clearly is the best provider now. Their API is very reliable and has many nice features. The price-performance of the model lineup is incredible. We used open weights model via Fireworks for a long time (e.g. Kimi K2.5). Fireworks is a great provider but we still ran into issues here and there (Same with Anthropic and Google). OpenAI just works, is fast and in my view has a better price-performance ratio across almost all levels of intelligence.
That’s what I was seeing too, and my cache hit rate was below 25% during that same time, leading to significant burn of my weekly limit (Pro 5x plan) via all the uncached input. Doesn’t prove that the overload caused the cache failure, but it does seem to point to some common infrastructure cause. No such problems (cache hit or system overload) with Terra or Luna.
Was that Codex/subsidised-usage or API? I do get overloaded in Codex-account from time to time, but API is rock solid.
They obviously load shed a bit of Codex-sub during peak times, and for the amount of tokens you get for a sub, I don't mind. I just mean the API where you pay-per-token is rock stable.
For our production app, OpenAI clearly is the best provider now. Their API is very reliable and has many nice features. The price-performance of the model lineup is incredible. We used open weights model via Fireworks for a long time (e.g. Kimi K2.5). Fireworks is a great provider but we still ran into issues here and there (Same with Anthropic and Google). OpenAI just works, is fast and in my view has a better price-performance ratio across almost all levels of intelligence.