> There’s no reason we couldn’t decide that we want to err on the side of employing too many people.
One might bring up the personal consequences bourne by surplus employees who're then laid off during the unavoidable corrective phase - or is that not something society should care about? What are you optimising for?
I am quite sympathetic to your position. Seeing those who manage to evade accountability consistently paying a heavy personal price was immensely satisfying. But at the same time, I don't think it resulted in any structural changes that minimise the proportion of accountability-evaders plauging society.
Ideally of course everyone, irrespective of any immutable traits they may have, gets to enjoy a healthy, satisfying, and stable life with plently of avenues for upward mobility. Short of that ideal, a society which equally burdens the rich and poor with devastating, seemingly random, unavoidable life-chaning events is decidedly better than one which only affects the poor.
So for these reasons I don't advocate for the actions of "the kid" but I don't think the consequences of his actions were in any way "bad" per se.
Perhaps if the rank and file at a company see personal consequence for those in the topmost posistions (salary deductions, demotions, or firing) in response to such glaring fuck ups that might even help mitigate some of these morale issues.
I think the issue is that information was expensive until the owners of capital gatekept knowledge itself. Now that information is cheap to reproduce they'd rather gatekeep the conduit to that knowledge and devalue knowledge itself rather than pay up in accordance with the very precedent that they set all along.
I am mostly economics illiterate but I understand a subsidy to be an economic concession given by the state to an entity which gives said entity a relative advantage compared to its peers.
In that sense (which could very well be bogus), letting a company violate individual IP of basically every human is less of an economic concession and more of unconsented to IP open season.
Even if one were to drop "economic" from "economic concession" and instead view a subsidy through the lens of a more general concession, one could say that the US Govt gave US AI companies a legal concession to sidestep the copyright protections of other US entities. But the US Govt should only get to undermine the copyright protection of other US entities - who gave American companies the right to violate the copyright of non-Americans?
Can you please tell my, as someone who is neither Chinese nor American, "why" I should care if a Chinese company stole from another American company (that in turn stole from everyone) to give me a cheaper service that fits my use case?
> to give me a cheaper service that fits my use case?
Because they aren't giving you a cheaper service that fits your use case.
Best Case scenario, it's a trillion-dollar behemoth stealing from a billion-dollar behemoth so they can add their own explicit restrictions/weights on top to influence the masses.
There is no 'robin hood' here, any perceived value you get is clearly and explicitly tainted. "I don't care if it doesn't show me non-party-line results - It makes me a cheap UI !". Ethics/morals be damned.
> There is no 'robin hood' here, any perceived value you get is clearly and explicitly tainted. "I don't care if it doesn't show me non-party-line results - It makes me a cheap UI !". Ethics/morals be damned.
I can't tell if you are talking about Anthropic or Alibaba here.
In a world which already has the likes of Anthropic and OpenAI, having Chinese labs be a counter balance is decidedly better than the hypothetical where American companies had a global monopoly on LLMs.
If your argument is that all present LLM offerings are unethical then that is something I am sypmathetic to. That said, I am also unable to offer a conceivable roadmap to undoing the opening of the LLM Pandora's box so I tend not ground my arguments in anti-LLM advocacy; that would be very 2023 of me.
I'm curious what's been your experience with Sarvam outside of Indic languages - Indian English (perhaps mixed with romanised indic verbiage) and also documents with complex layouts (figures, tables, etc).
I've been quite curious but hesitant about Indian offerings, particularly because they seem to be priced a little higher than what I would think they should be (I could be wrong and simply be misrembering though).
Sarvam is exceptionally tuned for indic languages we have more than 20 languages and it perform well for all in ocr. Iam yet to test with other languages. No any models come close for indic languages like sarvam. I saw they recently dropped price per page to 0.5 inr which is much cheaper. The only downside is the zip file based delivery.
I'm not sure I understand what you're arguing for? There are massive companies that collectively profiting off of stolen IP and are now gatekeeping even their paid offerings - surely consumers will rail against this? Personally, I feel very bad and can't wait for Chinese models to continue improving as much as they can prior OpenAI's and Anthropic's IPOs.
I’m not arguing for anything, actually. The ‘fair’ ship has sailed, even if the pirates somehow get shut down (which would be suicide by USG, won’t happen, national security issue), open Chinese models are not even hiding the fact that they distill from the frontier US labs, thus benefiting indirectly from the stolen content.
Note I don’t particularly like the ‘stolen’ word here as I don’t like when the music and film companies use it in the same context. Copyright infringement? Sure. Theft? No.
> I don’t particularly like the ‘stolen’ word here
Except that's the standard that we've measured everyone with up until the LLM/generative tech boom. I don't see why the benchmarks should change now. I realise my argument doesn't move reality but that doesn't mean we shouldn't call a spade a spade. Said companies carried out theft (or copyright infringement if you prefer) at industrial scale which is far more reprehensible crime against humanity than anything the individuals we think of as "digital pirates" have committed.
> open Chinese models are not even hiding the fact that they distill from the frontier US labs
The difference is they return to the same system that they feed from (indirectly); people get access to model weights even if the entire model isn't open source. The same can't be said for OpenAI, Anthropic, Google etc (who also benefit from Chinese models and train on them).
Sure, the alternatives aren't a panacea of fairness but I'd much rather advocate for and support the thieves who give me a better deal if my choice is limited to thieves. Especially if thieves aren't hostile to their customers like Anthropic is (which is why I replied to you in the first place).
And the Chinese models rip IP just like everyone else before them. Your argument is moot.
This was a problem for 5+ years ago. Nobody cares or at least the majority voice does not care across the world. Cat is out of the bag and there is no way to put it back in.
EDIT: Worth noting that I have long held the belief that if you put data out on the public sidewalk that you should have low to no expectation that it’s IP. It’s how I think about Google Maps data for example. If they want to reap the benefits by not walking it off the a user login than they can feel the pain if folks use that information. Same applies for media that has been bought, Reddit comments or any other datasets.
> And the Chinese models rip IP just like everyone else before them.
The difference is the Chinese models return to the same system that they feed from (indirectly); people get access to model weights even if the entire model isn't open source. The same can't be said for OpenAI, Anthropic, Google etc (who also benefit from Chinese models and train on them).
Further, Chinese models are significantly cheaper and the comapnies aren't hostile to their customers.
> Worth noting that I have long held the belief that if you put data out on the public sidewalk that you should have low to no expectation that it’s IP.
Except your beliefs aren't the cornerstone of modern jurisprudence. Why are models able to reliably produce replicas of Ghibli movies which go well beyond any example you listed?
One might bring up the personal consequences bourne by surplus employees who're then laid off during the unavoidable corrective phase - or is that not something society should care about? What are you optimising for?