The competition is real in pricing. Thanks for the Chinese open models, US big players have to cut their inference pricing. We've done a bunch of evals between the models, and Kimi K3 was the first one that actually could compete or be even better than Opus or Sol in our use cases, with a fraction of the price. All our developers use K3 as their programming model, and it now powers a big part of our systems instead of Opus and GPT. Surprisingly the new Sol pricing is quite similar to K3...
Now DeepSeek v4 Flash 0731 is eating Gemini's lunch, and suddenly we saw a price cut (the "introductory price") for 3.7. DeepSeek is of same quality or sometimes better than Gemini for text, Google knows it and they have to compete. Too bad it's too little and too late, it's still 4-5x more expensive in our evals.
And these models are not going away, nor their prices going up because of competition in the inference providers and due to the fact that you can buy/rent the hardware and run them in your own premises.
It's funny how quickly we went from "the greedy US companies are subsidizing prices to keep competitors out of the market" to "the greedy US companies are overcharging because they are greedy."
I'm not happy about it, but API prices seem actually reasonable if you compare it to free-market pure inference providers. They need to buy the same power, same hardware as the big labs without needing any R&D. This means that the coding plans must be subsidized. I'm not happy about that, because it means we'll be paying more, until hardware and maybe power drops in price, which, if we look at e.g. the housing market, may never happen during my lifetime.
No, almost everyone has been in agreement that they’re subsidizing subscriptions, but there have literally been dozens (hundreds?) of threads on HN in the past 12 months with people vehemently arguing that API prices are subsidized.
I think both can be true - they're losing money on the API and they're still charging too much. Which is really where the music stops for American AI investment. I also use K3 now and its perfectly capable for the development work I'm doing. I don't shed any tears for OpenAI or Anthropic.
API are not subsidized. We know this is true from the existence of independent inference providers (a lot of which are crypto companies that would otherwise just be mining if inference weren't actually profitable).
I don't think the rando providers like Novita, Phala, AkashML, etc., which frequently provide discounts to compete for traffic, are VC-funded. They have to actually make money to stay afloat.
Certainly not everyone. I find it far more likely that both subscriptions and API pricing are profitable in their own right. The fact that some people get great value out of their subscription (just like some people get great value out of their car insurance) doesn't mean it's subsidized.
That’s a fair distinction. I think they are willing to subsidize individual subscriptions for people who maximize usage, but their overall subscription business may be profitable.
Both can be true simultaneously. American companies were comfortable with a high cost structure because it padded their revenue numbers. The government tried to block out competitors through import/export controls, and it ended up backfiring by creating more efficient competitors. Very similar to our competition with Japan in the 80s. We'll see if we give the AI labs the kinds of sweetheart protectionism that US car makers got.
I don't think that kind of protectionism will work with a digital asset. It is much more difficult to erect barriers when there are no physical goods and transportation across geographic borders is instant and free.
It will likely be on the enterprise side if it happens. Hard to control supply, but if demand is concentrated to a few hundred entities, that's easy to do.
As one example: Different access patterns. If you're using the app and have a subscription you get cheap tokens in the hopes you need more at which point you pay API prices which are much more expenseive.
I agree with that. However, there’s much dialogue around subsidized tokens for business use too, that people paying for the tokens are also vastly underpaying vs the “real” cost. I certainly don’t know the answer to that. Maybe the Chinese companies are also doing it. Maybe nobody is doing it.
Looking at reserved capacity cost for PTUs on azure, which I think they’d probably not subsidize but can’t be sure, I’m inclined to not agree with the vast undercharging for tokens hypothesis.
Wait, there's 13 providers for Kimi K3 in OpenRouter. I'm having a hard time believing every single one of them provides them without any profit.
And this one is easy to calculate: take your monthly API spend to K3, then rent a stack of 8xB300 for a month and see how much it costs. I would say you're about to save 10-15k dollars per month if you have enough traffic compared to pay per token pricing.
It's not very complex math, and the hardware of course is cheaper if you bought it last year and if you have extra GPUs waiting in your warehouse (depending on if you can produce enough energy cheaply).
I shifted from DeepSeek v4 Flash 0731 to Gemini 3.7 flash on openrouter and price shoots up almost double with no visible change in outcome. So, today I reverted back.
I'm using hermes and keep close control over open router provider to DSv4 flash 0731
Use only the deepseek provider, you can configure openrouter to do that(but they create some obstacles, go to configurations and allow all providers)
Right now i'm testing deepseek harness, don't wait to test it. The plugin architecture and self awareness of workflows really make you think about "what is a tool vs what is a project".
Is there actually a difference in cache rates between OR and official API? I have a preset set up on OR so that I only send traffic to deepseek. The preset is important otherwise you will send traffic to different providers but if you weren't doing this already then what can I say, water is wet, of course cache rates will be awful. I get about 70% cache hit rate with the preset which is appropriate for what I'm doing. I haven't used the official API though.
I was struck by a video ad that Google released yesterday with testimonials by three developers about using Gemini 3.7 Flash [1]. The point they emphasize most is price, followed by latency. The marketing strategy definitely seems to be shifting.
Yes and no. It competes in the mid tear not in SOTA. It's a very valid model if you need things like computer use or image recognition. Especially with the 3.7 "introductory prices". It's multi-modal and better than GPT 5.6 Terra while only a bit more expensive.
And Fireworks did not yet. They are still under the limit of not feasible to self host... Let's see if other providers follow DeepSeek with their flash pricing.
Long-context agentic tasks and Rust engineering are our use cases where Kimi definitely is better than Sol. We can measure our own systems and the numbers say that Sol has no chance against K3 or Opus, and K3 is so so so much cheaper than Opus right now.
You cannot just look at the price tags for these models, you must eval and see the price per task. In our previous eval rounds Sol was more expensive than Opus (with its original price), took much longer, and provided worse results. Kimi does not have these issues, it's just as good as Opus with a smaller price tag.
China's 50 Cent Party being a real and noticeable thing (and the two biggest things they like to shill is open weight Chinese models and the futility of resisting a Taiwan invasion), I have to take things like this with a healthy dose of skepticism without corroborating data, since independent evals didn't show the price per task lead you're showing.
If there's independent data showing this feel free to share a link, I haven't seen it. DeepSWE has been most closely matching what I see in my own use.
Internal reports from company? Maybe not. I'm just saying you have to eval eval eval if you are working in this industry. There's a ton of victories in price, and price is right now the key thing all the customers are talking about.
It's not always Chinese models. For example GLM 5.2 just did not work for us at all. And Gemini is still the best cheap model for non-text agents.
If you don't have a good eval set and if you don't check the models weekly, you are missing on things. And Opus 4.8 is still the absolute quality king for agentic tasks. Too bad it's so expensive.
And the clearest thing here is that Fable, Opus, and Sol are all too expensive. I'd say a healthy 75% cut to token prices and they are back in competition.
Yep. It's a bit scary also. There's a lot of opportunity in the market now, but the downfall of the big US inference labs is going to hurt here in EU too sadly...
We know the price wars are coming. It's gonna be interesting to watch how it plays out in real time. Anthropic/Openai probably want to race to the IPO before the price wars start to have real impact.
Now DeepSeek v4 Flash 0731 is eating Gemini's lunch, and suddenly we saw a price cut (the "introductory price") for 3.7. DeepSeek is of same quality or sometimes better than Gemini for text, Google knows it and they have to compete. Too bad it's too little and too late, it's still 4-5x more expensive in our evals.
And these models are not going away, nor their prices going up because of competition in the inference providers and due to the fact that you can buy/rent the hardware and run them in your own premises.