Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The competition is real in pricing. Thanks for the Chinese open models, US big players have to cut their inference pricing. We've done a bunch of evals between the models, and Kimi K3 was the first one that actually could compete or be even better than Opus or Sol in our use cases, with a fraction of the price. All our developers use K3 as their programming model, and it now powers a big part of our systems instead of Opus and GPT. Surprisingly the new Sol pricing is quite similar to K3...

Now DeepSeek v4 Flash 0731 is eating Gemini's lunch, and suddenly we saw a price cut (the "introductory price") for 3.7. DeepSeek is of same quality or sometimes better than Gemini for text, Google knows it and they have to compete. Too bad it's too little and too late, it's still 4-5x more expensive in our evals.

And these models are not going away, nor their prices going up because of competition in the inference providers and due to the fact that you can buy/rent the hardware and run them in your own premises.



It's funny how quickly we went from "the greedy US companies are subsidizing prices to keep competitors out of the market" to "the greedy US companies are overcharging because they are greedy."


They are subsidizing the non-API use cases and overcharging on the API use cases.


It's interesting if they need to cut off their subscriptions to be able to compete in API prices. Very interesting...


I'm not happy about it, but API prices seem actually reasonable if you compare it to free-market pure inference providers. They need to buy the same power, same hardware as the big labs without needing any R&D. This means that the coding plans must be subsidized. I'm not happy about that, because it means we'll be paying more, until hardware and maybe power drops in price, which, if we look at e.g. the housing market, may never happen during my lifetime.


I'm not actually sure anymore they are overcharging on the API use cases, because there's so much competition.


No, almost everyone has been in agreement that they’re subsidizing subscriptions, but there have literally been dozens (hundreds?) of threads on HN in the past 12 months with people vehemently arguing that API prices are subsidized.


I think both can be true - they're losing money on the API and they're still charging too much. Which is really where the music stops for American AI investment. I also use K3 now and its perfectly capable for the development work I'm doing. I don't shed any tears for OpenAI or Anthropic.


API are not subsidized. We know this is true from the existence of independent inference providers (a lot of which are crypto companies that would otherwise just be mining if inference weren't actually profitable).


Those providers are VC-funded, they don’t have the leeway to pivot away from AI. I would be interested in some examples though.


I don't think the rando providers like Novita, Phala, AkashML, etc., which frequently provide discounts to compete for traffic, are VC-funded. They have to actually make money to stay afloat.


Any public company providing inference who’s claimed its profitability in a way they risk legal penalties if untrue?

(Makes sense to me, just curious about this additional piece.)


Certainly not everyone. I find it far more likely that both subscriptions and API pricing are profitable in their own right. The fact that some people get great value out of their subscription (just like some people get great value out of their car insurance) doesn't mean it's subsidized.


That’s a fair distinction. I think they are willing to subsidize individual subscriptions for people who maximize usage, but their overall subscription business may be profitable.


Anthropic are definitely profitable on my 20 eur subscription, as I have kids and no time.


And every one of them is wrong.


Both can be true simultaneously. American companies were comfortable with a high cost structure because it padded their revenue numbers. The government tried to block out competitors through import/export controls, and it ended up backfiring by creating more efficient competitors. Very similar to our competition with Japan in the 80s. We'll see if we give the AI labs the kinds of sweetheart protectionism that US car makers got.


I don't think that kind of protectionism will work with a digital asset. It is much more difficult to erect barriers when there are no physical goods and transportation across geographic borders is instant and free.


It will likely be on the enterprise side if it happens. Hard to control supply, but if demand is concentrated to a few hundred entities, that's easy to do.


Can you explain better how OpenAI, for example, would simultaneously subsidize tokens while overcharging for them? Not sure I get it.


As one example: Different access patterns. If you're using the app and have a subscription you get cheap tokens in the hopes you need more at which point you pay API prices which are much more expenseive.


I agree with that. However, there’s much dialogue around subsidized tokens for business use too, that people paying for the tokens are also vastly underpaying vs the “real” cost. I certainly don’t know the answer to that. Maybe the Chinese companies are also doing it. Maybe nobody is doing it.

Looking at reserved capacity cost for PTUs on azure, which I think they’d probably not subsidize but can’t be sure, I’m inclined to not agree with the vast undercharging for tokens hypothesis.


Wait, there's 13 providers for Kimi K3 in OpenRouter. I'm having a hard time believing every single one of them provides them without any profit.

And this one is easy to calculate: take your monthly API spend to K3, then rent a stack of 8xB300 for a month and see how much it costs. I would say you're about to save 10-15k dollars per month if you have enough traffic compared to pay per token pricing.

It's not very complex math, and the hardware of course is cheaper if you bought it last year and if you have extra GPUs waiting in your warehouse (depending on if you can produce enough energy cheaply).


Bulk discounts are definitely a thing. Even established businesses do that.

But the biggest discount people see is subscriptions. You get a few thousand dollars of work from a couple hundred dollars.


It's simple, they aren't subsidizing tokens at all.


True for api pricing across the industry. Not as true for 200$ a month plans assuming you used every bit of it.


I shifted from DeepSeek v4 Flash 0731 to Gemini 3.7 flash on openrouter and price shoots up almost double with no visible change in outcome. So, today I reverted back.


Try DS through their own API if that's feasible, AFAIK they're much cheaper than through OS due to cache hit rates.


I'm using hermes and keep close control over open router provider to DSv4 flash 0731

Use only the deepseek provider, you can configure openrouter to do that(but they create some obstacles, go to configurations and allow all providers)

Right now i'm testing deepseek harness, don't wait to test it. The plugin architecture and self awareness of workflows really make you think about "what is a tool vs what is a project".

It has been an amazing experience


They just raised their prices sadly.


Is there actually a difference in cache rates between OR and official API? I have a preset set up on OR so that I only send traffic to deepseek. The preset is important otherwise you will send traffic to different providers but if you weren't doing this already then what can I say, water is wet, of course cache rates will be awful. I get about 70% cache hit rate with the preset which is appropriate for what I'm doing. I haven't used the official API though.


Zenmux say the cache hit rate is 98% for the deepseek flash API. I don't know why, but performance is definitely worse using openrouter.

https://zenmux.ai/deepseek/deepseek-v4-flash


Could be compression or metadata in their pipeline adding noise and making requests less generic.


they will train / retain your prompts though, so to me it's not a valid option


Openrouter is going to cost you a lot more than the 5% fee, unless you lock the provider.


I was struck by a video ad that Google released yesterday with testimonials by three developers about using Gemini 3.7 Flash [1]. The point they emphasize most is price, followed by latency. The marketing strategy definitely seems to be shifting.

[1] https://youtu.be/kacf2bib-X0


You play to your outs. Gemini is far behind on quality.


Yes and no. It competes in the mid tear not in SOTA. It's a very valid model if you need things like computer use or image recognition. Especially with the 3.7 "introductory prices". It's multi-modal and better than GPT 5.6 Terra while only a bit more expensive.


> And these models are not going away, nor their prices going up [...]

Well, DeepSeek just raised prices.


And Fireworks did not yet. They are still under the limit of not feasible to self host... Let's see if other providers follow DeepSeek with their flash pricing.


They did now. Landing somewhere between Terra and Luna now per task, with the quality of Gemini 3.7 flash.


> Opus or Sol in our use cases, with a fraction of the price.

I assume it's highly use case dependent, though?

Even before the price cut seems like Sol was price competitive with Kimi

https://artificialanalysis.ai/models?models=gpt-5-6-sol-xhig...

And now it should be considerably cheaper


Long-context agentic tasks and Rust engineering are our use cases where Kimi definitely is better than Sol. We can measure our own systems and the numbers say that Sol has no chance against K3 or Opus, and K3 is so so so much cheaper than Opus right now.

You cannot just look at the price tags for these models, you must eval and see the price per task. In our previous eval rounds Sol was more expensive than Opus (with its original price), took much longer, and provided worse results. Kimi does not have these issues, it's just as good as Opus with a smaller price tag.


China's 50 Cent Party being a real and noticeable thing (and the two biggest things they like to shill is open weight Chinese models and the futility of resisting a Taiwan invasion), I have to take things like this with a healthy dose of skepticism without corroborating data, since independent evals didn't show the price per task lead you're showing.

If there's independent data showing this feel free to share a link, I haven't seen it. DeepSWE has been most closely matching what I see in my own use.


Internal reports from company? Maybe not. I'm just saying you have to eval eval eval if you are working in this industry. There's a ton of victories in price, and price is right now the key thing all the customers are talking about.

It's not always Chinese models. For example GLM 5.2 just did not work for us at all. And Gemini is still the best cheap model for non-text agents.

If you don't have a good eval set and if you don't check the models weekly, you are missing on things. And Opus 4.8 is still the absolute quality king for agentic tasks. Too bad it's so expensive.

And the clearest thing here is that Fable, Opus, and Sol are all too expensive. I'd say a healthy 75% cut to token prices and they are back in competition.


> I'd say a healthy 75% cut to token prices and they are back in competition.

Surely you don't want them to be the reason the bubble bursts?


Yep. It's a bit scary also. There's a lot of opportunity in the market now, but the downfall of the big US inference labs is going to hurt here in EU too sadly...


We know the price wars are coming. It's gonna be interesting to watch how it plays out in real time. Anthropic/Openai probably want to race to the IPO before the price wars start to have real impact.


Tell me why price adv not working this same on web pages. Why price od advertisment on portals, social media etc. not fall down?


Not including Grok/Cursor seems a major flaw in your model


Oh we did try to get Grok for evals but they had some weird EU limitations last time we checked. Which the open weight models don't have.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: