Currently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.
Deepseek seems to have gotten too cheap. I have been using it for a long time and it's at a point now where my credits balance barely moves even at max setting.
i'm doing opencode <-> openrouter <-> official deepseek api (i don't get the opencode hate, i like it)
how are you doing it?
am also using Kimi K3 via kimi-code
and also GLM 5.2 via ZCode
happy with all three, they're trailing frontier but i figure if i'm running GNU/Linux then i ought to favour open weights models with my €s -- reduced my usage of claude/gpt to the ~$20 tier just to keep abreast of claude_code/codex developments
When the company I work for was evaluating it, there were multiple rough points. Their terms and conditions allowed training on prompts, the default behavior was to route prompts to their servers for conversation summary/labeling. One of their lead maintainers is also super toxic on many issues.
Sorry this is all baseless with no links, I’m on my phone and locating those issues again isn’t something I have time for.
It’s a good tool I just don’t like the privacy policies nor maintainers attitudes.
1. The privacy policy was a bit misleading, but it has since been updated to reflect the exact state of things. [1]. For example, DeepSeek models have ZDR, although their ZDR contract is renewed monthly. It COULD change. You need to toggle a Setting in your account to use DS.
2. At one point (apparently) summary and title generations were handled by Grok. This has changed, by default it uses your 'small_model' configured in your config. By default, it will use a cheap model provided by your provider. E.g. if you have ChatGPT API connected, it will use the cheapest ChatGPT model. OpenRouter users MAY see it routed to a free model however. [2] [3]
The banner on account settings; and a blurb on the pricing page:
"We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice."
I don't mind these prices but I find the need to check against two different time brackets of unequal length an annoying distraction. I guess I need to make some little background app or plugin.
but openrouter says they don't expect the price to change other than through the deepseek api, other people hosting the same model will keep charging the same price.
Unfortunately cache reads with third party providers are all 10-50x more expensive than with DeepSeek, so they're not even close to as cost efficient for multi-round agent use.
Yeah, Dax from OpenCode said that it appears to just be traffic shaping, nothing to do with the inference economics. He also said that OC have already replicated the inference cost in internal experiments.
You can derive a pretty solid yardstick of how things are going for China by what you can find on the aftermarket, currently there is a glut of nvidia 4080s that have had their memory doubled up to 32GB. I'd have to assume they got a good deal buying up piles of H100s or whatever else was eating rack space, or potentially took a loss because they have hit the constraints of the # of cards they can rack.
On the OEM side of things 9070/XTs are also shooting back up in price now that we have <$100 USB 4 egpu docks.
People like to complain about how expensive things have gotten but I think it's pretty neat that there's so much pressure for throughput that it's even viable to buy 4 docks and 4 $850 GPUs and still save money over a single 48GB card.