Hacker Newsnew | past | comments | ask | show | jobs | submit | civet_java's commentslogin

There are some commentors in this thread downplaying the severity of a service provider being less than transparent about exactly what their shipped tooling does on customer's machines.

That the provider's business needs necessitate the this behaviour doesn't justify their lack of honest disclosure. That honest disclosure would render the solution to their problem useless isn't my problem. If anything, that they thought this was acceptable makes me wonder what else they're harvesting from my machine? PII?

The cynic in me can't help but feel that the state of these comments reflects less on the commentor's views of this debacle but rather their feelings about AI/Anthropic/America/what-have-you.


This is not a technical issue or a matter of service agreements.

They totally knew that this would never be accepted by users, and that is precisely why they resorted to obfuscation and steganography to exfiltrate the data. This was quietly added in an update, and would have been removed in a later one had it gone unnoticed.

How is this any different from hiding a drug in food and randomly feeding it to a homeless person as a human trial? Even if that homeless person is an undocumented immigrant, does that make that acceptable?

They crossed the line.


I mean, you’d resort to an obfuscated approach if you thought the ‘malicious’ users would remove your direct telemetry. The other users could be perfectly happy about it, but if you announced the change, obviously the malicious users would hear about it too and disable it.

This isn’t a comment on whether I agree with the change. Just that your analogies aren’t applicable here.


Anthropic always came off to me as what can shorthand be described as an abusive controlling partner + the resulting relationship.

If you start viewing the conversations around it through that lens, a lot of stuff intuitively clicks into place.

Including this apologism for example.


Seems like the people talking about AI safety and alignment don't want to stop at just harnessing their AIs. They also want to keep themselves safe from competition and make sure only users aligned with their business and political agenda gets to access their products.

The same people are boasting about being the future of work and is schmoozing with politicians to draft regulations, by the way.


Why would they want to do that when it’s likely that someone will find out anyway and it could turn into a scandal?

One possibility: they wanted to keep it secret because they’re crooks.

Another possibility: there are so many calls that it costs a lot in operating expenses.

Remember when we mentioned that even just saying “hello” at the start of a chat costs extra money for no reason?

So if I create a mechanism on the client side that every transmission uses 4 bytes instead of 40… it is clever way how to spend less.

Of course, it doesn’t excuse them for not mentioning it anywhere… or for hiding it on page 385 of the terms of service… like when the Vogons told Earthlings that the notice about Transgalactic Highway was in the basement on Alpha Centauri :)


Abusive controlling partners rarely are long-term timeframe rational, because being long-term timeframe rational is at odds with being an abusive controlling partner.

Things do not need to make sense. They often do not. They just appear just enough like they would so that it flies under the radar.

It's all just conway's law. It had to be like this. It cannot be any other way.


> Abusive controlling partners rarely are long-term timeframe rational, because being long-term timeframe rational is at odds with being an abusive controlling partner.

Humans are rarely “long-term timeframe rational” (rational choice theory is very much a known-false model of human behavior on both the individual and aggregate level), but there is nothing about abusively exploiting an initially trusting and eventually dependent telationship that it is inconsistent with long-term rationality.


I can't really parse this comment, but I also have a feeling that nothing is lost in struggling with that.


> The other users could be perfectly happy about it

this could be true about a miracle drug that is untested too. still, human rights laws dictate that humans be allowed to opt in to such trials. so you must at least shift your argument to "users opted in to all of anthropic's bullshit, explicit or otherwise, when they opt to install and run claude code."


> randomly feeding it to a homeless person as a human trial

Would you be shocked if some of the current tech bro companies could with finding such subjects?


First its the "Chinese" then it will be people using "cyber" capabilities, or "jailbreaking" or "going against Dario" or any other thing they find "objectionable".


You forgot "Think about the children"


That would actually be a thing to care about


Won't somebody please think of the Chinese cyber children!


what about the "ai children who are hosted on servers in china and usa" think about them too


I asked Claude to think of the children. It worked for three minutes, came up with a plan, and charged me $9.


XD


Their terms of service already effectively attacks people for criticizing Anthropic. It says that if you use Claude to criticize Anthropic, then you've pre-agreed to pay for their lawyers going after you, and pre-agreed to lose the court case.


Since Fable if you use Claude to do ML research they find objectionable you are delegated to less capable models.


I don't think that would hold up in a court. At least not outside of the US.


Doesn't matter. Getting entangled up in a court case that lasts years is hassle enough by itself that it's effective as an attack.


[flagged]


I don’t mean to insult you, but it is probably worth refreshing yourself on when slippery slope is actually a fallacy. It’s not correct to suggest logical extension of an established precedent is a slippery slope.


The Slippery Slope Fallacy Fallacy


Putting "I don't mean to insult you" before a mild phrase just makes it sound you actually want to insult but in a passive aggressive way.


I saw it more as “if you are insulted by this reasonable advice, then you are an immature idiot” rather than a more direct insult. Still passive-aggressive, but no insult intended if the reader isn't triggered.

Source: I use similar phrases this way.


You mean there’s a slippery slope from “I don’t mean to insult you” to passively aggressively insulting them? ;)


Slipper slope is always a fallacy. It works as a metaphor, and as a rhetorical device, but the fact that it’s possible to construct a gradient from heaven to hell is not proof that each step in the gradient is inevitable.

It’s a good way to illustrate a potential bad path. But I can construct a slippery slope from commenting on HN to making bioweapons. That should not be seen as evidence that HN commenters are likely to do such a thing, and should not be used as an argument to forego (or ban) commenting on HN.


That's not correct. Many slippery slope arguments are perfectly valid and sound, namely whenever there is plausible evidence for the existence of the respective slippery slope. It can for instance be based on historical precedents, or on probabilistic or convincing balance of consideration sub-arguments for the slippery slope premise.

Whether its a fallacy or not depends a lot on how warranted the slippery slope premise is (notwithstanding other, possible logical errors that might me made in setting up the argument, of course).


Give an example where a slippery slope is itself a legitimate argument and not just a rhetorical technique to oppose A using arguments against B?


Opioid use comes to mind.


Wars are sometimes caused by small acts of violence (e.g. at the border) that are mutually retributed, becoming larger and larger, until the hostilities turn into an all out war. This is sometimes called "escalation of violence." If wars are sometimes caused in that way, then there must be some wars for which a slippery slope argument concerning the escalation of violence is not fallacious.


https://en.wikipedia.org/wiki/Slippery_slope

In a slippery-slope argument, a course of action is rejected because the slippery slope advocate believes it will lead to a chain reaction resulting in an undesirable end or ends. The core of the slippery slope argument is that a specific decision under debate is likely to result in unintended consequences. The strength of such an argument depends on whether the small step really is likely to lead to the effect. This is quantified in terms of what is known as the warrant (in this case, a demonstration of the process that leads to the significant effect).


No longer slandered as a fallacy in the lede :) the 2019 version:

  A slippery slope argument (SSA), in logic, critical thinking, political rhetoric, and caselaw, is a logical fallacy in which a party asserts that a relatively small first step leads to a chain of related events culminating in some significant (usually negative) effect.
https://en.wikipedia.org/w/index.php?title=Slippery_slope&ol...


Textbook fallacy fallacy.


Fallacy fallacy(and variants) implies there is a fallacy in the first place.

If the slope is slippery, it is not a fallacy to conclude that things will slide when placed on it.

Claiming something is a fallacy —when it is not— is simply a false statement of fact.


For a slippery slope there has to be something about the slope that makes you start at the top. If you can go to any point on the slope, it isn't even a slope.

Anthropic can start censoring criticism of Dario completely independently from whether or not they censor Deepseek.


Textbook fallacy fallacy fallacy


[flagged]


Wasn't she too busy burning down Portland to go to DC?


Whether or not you find Anthropic's behavior bad, theybhave been very loudly stating the foreign labs have been distilling their models for a while now. This seems like an obvious response to me that would be a mechanism to make that obvious.


From my understanding, distilling the model with another model is not illegal per se. Also, the output of the LLM is public domain by law, too.

So, why all this "effort" to protect the model? This is a free market, and moving fast and breaking things is the norm.

If they are so adamant on protecting their IP, maybe they can start by respecting others' IP, so we can start talking about ethics, equality and playing fair.


> distilling the model with another model is not illegal per se.

Just because it is legal, that doesn't mean Anthropic wouldn't reasonably want to prevent that from happening (which, from my understanding, isn't illegal either).


I love the asymmetry. When small fish tries to protect itself, big fish hits small fish with "It's not illegal" pole.

When small fish points out that what the big fish is crying about is "not illegal", big fish has the right to be above the law to prevent the problem themselves.

Having values requires equality. They have lost the right to cry foul when they trained their model with "but it's fair use" card. Life works by reaping what you sow. Now they are at the reaping stage.


"It's not illegal" is only an argument against lawsuits / law enforcement involvement. Those PoW anti-AI things people put on pages aren't illegal either.


No. From my interactions, I have understood that some people use the same argument to wash their consciences from any guilt. What they do is unethical, but not illegal, and they hide under the same argument to drown the ethical angle.

In other words, being honest to oneself is important.

Anti-scraping measures people utilize are neither unethical nor illegal. That’s the difference.


I’m still frequently shocked by the entitlement people feel to other people’s work/ideas/data/bandwidth/server load, to feed a multi-trillion dollar industry. I find the totally cynical “well when you’re making an omelet…” types to be a bit pathetic, but I understand their motivation— they’re simply greedy. But I just can’t understand the genuine indignation about people attempting to limit or stop ingestion of their own work, even if it’s just for the bandwidth costs. Go ingest your own shit.


I often wonder what those people are like IRL. I'd surmise they're the people that are easy to hate. Greedy and intolerable yet want to be the focus.


It's good to agree that some don't have a conscience, and maintaining an appearance matters more. And appearances change based on what's legal or not.

They could detect the other AI labs and also silently burn the tokens at a faster rate providing fewer tokens for money, which does sound illegal to me.

The comments only further prove that without more regulation around this, big AI wouldn't have a "don't be evil" attitude going forward.


Anyone who called for regulations/guardrails of any kind were shouted down as Luddites who hate progress. We all knew this was going to be a mess but $$$ so screw it right?


Maybe that's how OpenAI was able to blow $5.73 BILLION dollars on "marketing" i.e. paying politicians and influencers.


> I love the asymmetry.

Much as I hate to defend companies climbing to success and pulling up the ladder afterwards, this asymmetry you note is kind of the whole point a company would want to grow big. Growing an organization has some super-linear costs and generally sucks for most individuals living through it - including the management - but it's still considered worth it, precisely because big entities can do things small entities cannot, and escape the threats from smaller competitors.

It's so basic it's actually part of the reason we exist, and animals of various sizes exist, and generally why evolution didn't stop at single-cellular life.

> They have lost the right to cry foul when they trained their model with "but it's fair use" card. Life works by reaping what you sow. Now they are at the reaping stage.

Yup. Except what they're reaping is insane cashflow and ability to pull stunts like these. We can call out the hypocrisy until our throats run dry, and in ideal fantasy land this would've meant something, but here in the real world, they sow the seeds of success, and now are reaping the right to be hypocritical and continue to get away with it.


I've actually heard it quite a few times from different people who want to climb the greasy pole to get heard or resources. Idk it just seems rather soulless and slightly psycho to me. It also seems like that kind of system is rather broken and unstable, if the only way you get impact is to climb up the ladder and whatever that entails.

Change this from humans to companies and I still think it feels slightly wrong.


I disagree with your assessment that large organisations are beneficial.

We can see with our current crop of large organisations that they really struggle to create anything new; most of their new products or services were developed by a small organisation and then acquired. A lot of those products are then enshittified and badly managed because large organisation politics screws things up.

Large organisations are inefficient (everyone has stories of people in large organisations literally doing nothing all day). They are horrible to work for because of the politics. They mistreat their customers and their employees. Their executives tend to lose touch with reality, surround themselves with yes-folk and descend into authoritarian psychopathy.

My personal opinion is that we would be much, much, better off if we had fewer large organisations and more smaller organisations.


They are beneficial for those with an equity stake. That much is clear.


Needs a little more precision: not "those with an equity stake": those with a disproportionally large equity stake.

Otherwise it's just an opener for the old excuse of "they might be ruining your life, but it's all good, it's also your pension fund, little man, that's profiting from your life getting ruined, you should celebrate them!"


Agreed. And for oligarchs.


And I disagree with yours :).

Large organizations are necessary if you want things like airplanes and rockets and computers and MRI machines to exist. And if you feel you benefit from those things yourself, then large organizations that create and operate[0] them are beneficial to you, too.

> A lot of those products are then enshittified and badly managed because large organisation politics screws things up.

That's not caused by org size. It's how modern economics work because of ad-backed business models and few other things (a tangent for another time). Importantly, small orgs and especially startups are very much complicit in this - the venture capital business model in software settled around a symbiosis, where startups create toys (er, MVPs) and growth-hack the shit out of them, in hopes of winning an acquisition or IPO lottery (aka. "exit"), where a big org buys the whole thing for ${a lot}, and enshittifies it further in an attempt of extracting a positive multiple of ${a lot} from the market. Both sides know what they're doing, exits are planned from day 1, and at no point in this process "creating useful products" is ever a driving goal.

Note this symbiosis: it's a recurring theme.

> Large organisations are inefficient

In some ways. Small organizations are inefficient in others. More at the end.

> (everyone has stories of people in large organisations literally doing nothing all day).

Some (not all) cases of this are about maintaining slack in the system, which is necessary for efficiency. A system at 100% capacity is extremely fragile to breaking completely due to tiny, random workload spikes. Breakage is inefficient. Some degree of idle capacity improves overall efficiency.

> They mistreat their customers and their employees. Their executives tend to lose touch with reality, surround themselves with yes-folk and descend into authoritarian psychopathy.

That description fits small business owners much better IMO. In our times, at least in non-failed western countries, there's a limit to how abusive or careless a large organization can be with their customers or employees - their very size makes them easy to target legally. It might be hard to get through their well-funded legal defense, unless the case is slam dunk, but that's still much better than the armies of small businesses flying completely under the radar, flagrantly violating basic health and safety regulations, or flat out lying to customers in their face, because they're not worth the effort of investigating.

(Of course I'm using a biased sample; I don't know many CEOs of big orgs.)

Symbiosis angle: for abusive practices they can't get away with on their own, big organizations are more than happy to outsource to small orgs and then look the other way.

--

Anyway, key point: *there is no categorical difference between "large organizations" and "small organizations". You need a certain amount of people and communication (and capital) to do high-complexity endeavors. The difference between a well-integrated big corporation, and a hundred of small businesses that kinda end up together delivering something big, is just that the latter is using the market as management layer.

And yes, you need big orgs to create things like commercial airplanes and MRIs, simply because the big org is a boundary layer, within which you have a non-market based incentive structure, and this lets you build things the free market just cannot reach on its own.

--

[0] - Airports and hospitals are themselves large organizations.


> That's not caused by org size. It's how modern economics work because of ad-backed business models and few other things (a tangent for another time

I disagree, it's not "modern" economics, it's one half of the Malthusian trap as it manifests in all economics; the other half is that profits tend to zero, both halves are a loss of systemic slack, to reuse the good point you make later.

> That description fits small business owners much better IMO. In our times, at least in non-failed western countries, there's a limit to how abusive or careless a large organization can be with their customers or employees - their very size makes them easy to target legally.

I think this is more like predator/prey size dynamics. One way to keep safe from predators is to be too big to hunt. The regulators and governments are more like predators than their peer-competitors are, cf. "too big to fail".


I actually agree with most of this. A few points:

SpaceX built great rockets before it became large (though this is relative - a small rocket company is a large hairdresser, for example). There is a certain scale required for some types of business, agreed. But getting larger doesn't necessarily make them better.

> That description fits small business owners much better IMO. In our times, at least in non-failed western countries, there's a limit to how abusive or careless a large organization can be with their customers or employees - their very size makes them easy to target legally. It might be hard to get through their well-funded legal defense, unless the case is slam dunk, but that's still much better than the armies of small businesses flying completely under the radar, flagrantly violating basic health and safety regulations, or flat out lying to customers in their face, because they're not worth the effort of investigating.

Flat disagree with this. Small org CEOs are close to their customers and employees and if they behave like dicks then they get punished quickly. Obviously some still do, because people, but it's harder for a small company CEO to continue being a dick.

> Anyway, key point: *there is no categorical difference between "large organizations" and "small organizations". You need a certain amount of people and communication (and capital) to do high-complexity endeavors. The difference between a well-integrated big corporation, and a hundred of small businesses that kinda end up together delivering something big, is just that the latter is using the market as management layer.

There is a key step change when the first pure-management layer forms in an organisation. This is the management layer that only have other managers reporting to them, and only report to other managers. So no direct contact with front-line staff or shareholders. Personally, the presence of this layer is what classifies an organisation as "large". It's when the politics takes over from performance as the priority and the organisation starts to lose the connection between what the c-suite want and what the front-line actually do.

And all commercial airplanes, MRIs, anything, were built first by small organisations, and only later by large orgs. Large orgs just can't invent new things unless they form specialist small orgs to do it (skunkworks, or Palo Alto, or similar). Large orgs just don't work like that.


> Flat disagree with this. Small org CEOs are close to their customers and employees and if they behave like dicks then they get punished quickly. Obviously some still do, because people, but it's harder for a small company CEO to continue being a dick.

I'm thinking it might be both - depending on who the real customers are.

I've seen plenty of what I described in "boring" B2C like... grocery stores. But thinking about it, for a grocery store chain, customers are as much a commodity as the products they buy. Suppliers are where relationships (and power plays) matter.

(This might be fundamentally the same problem as the infamous case of "enterprise software procurement" - people using the software aren't the ones paying for it. For a grocery store chain, customers come and go all the time for many reasons, so it averages out anyway - but your suppliers and partners are what makes a difference in your bottom line.)

> And all commercial airplanes, MRIs, anything, were built first by small organisations, and only later by large orgs. Large orgs just can't invent new things unless they form specialist small orgs to do it (skunkworks, or Palo Alto, or similar). Large orgs just don't work like that.

Which is why I tried to point out the category error. "Large org with skunkworks" vs. "Bunch of smaller orgs forming an alliance and acquiring more smaller orgs to productionize a new technology" vs. "government megaproject" - they're all similar, arguably for a given invention they may very well be the same thing. Names and legal groupings are different, but the dynamics is (by anthropic principle) specific to what's needed for a given type of invention.

E.g. for stuff like airplanes or MRIs, you need individuals and small teams with lots of freedom (and a "hold my beer and watch this" culture often helps), but that gives you a prototype at best - scaling this so it works reliably, and then optimizing so it can be economical, both require throwing money at people doing boring work that mostly increments things on margin. And then the money has to come from someone, and someone must be willing to spend it to fund it all.

The actual org charts and legal charters don't matter - what matters is the incentives inside. I somewhat tentatively put forth a hypothesis: large orgs form to solve problems that the regular free market dynamics can't handle, by creating areas governed by different rule sets, within which that work can be done. Whether that's by fiat or corporate charter or a bunch of friends aligning their small businesses for the same goal, is window dressing.


yeah I see where you're going. But it doesn't address the politics/management problem - large orgs invariably generate internal politics as their incentives get misaligned with their objectives (the root cause of the 5-layer "pure manager" situation becomes a problem). Large orgs have to actually split off a smaller org (that doesn't have this problem) to actually do anything new.

The Innovator's Dilemma is a symptom of the same problem: a large org with an existing market cannot innovate because the incentives for management do not allow cannibalising the existing product sales to launch the new product.


At risk of losing the metaphor, they reaped stuff across all the lands, even ones that were not theirs, and it is questionbale that they even did most of the sowing in the first place


> Much as I hate to defend companies climbing to success and pulling up the ladder afterwards

Based on your post, you don't sound like you hate it at all.


> It's so basic it's actually part of the reason we exist, and animals of various sizes exist, and generally why evolution didn't stop at single-cellular life.

It's also quite natural to want it to stop at individual human life instead of us getting absorbed by some next bigger thing.

Which I'm fairly sure is also the desire (as far as they can be said to have any) of these animals of various sizes you speak of.

> We can call out the hypocrisy until our throats run dry, and in ideal fantasy land this would've meant something, but here in the real world, they sow the seeds of success

Just because they pulled a mirage over people's eyes doesn't mean it suddenly became the "real" world.


Is China the little fish here?


No, the ordinary netizen, who runs their personal web servers, who are hit by crawlers, their content ripped from their hands and their servers chocked during the process.


What Anthropic is doing is illegal in many jurisdictions. I don't know about the legal situation for the Chinese domains they mark, but steganographic data extraction without user consent would definitely be illegal in the EU, for example.


It's not illegal to distil the traces, but it is also not illegal for them to try to stop it.


Imagine an electricity generating company saying that they don't allow their electricity to be used to cold start a competitor's generator.


Do you think software should be regulated as a utility?


AI probably should be. The bulk of its efficacy comes from the work of “everyone else” (in loose terms). AI also aims/hope/threatens to replace such a large number and range of jobs that it probabky should be a commons.


Wise words. And probably in a few years more people will think the same, but now most are blinded by the gold rush or the hate for it.


> Do you think software should be regulated as a utility?

I do think that AI models that were trained using the biggest intellectual property heist in human history should be a utility for all, yes.


I would 100% support this as long as the same is true for human works. Every artist is trained on thousands of years of art tradition. Copyright wasn’t a thing for the first many millennia of artistic creation, and we still got Bach and Michelangelo.


> Michelangelo

Did you know that Michelangelo was the first "label" for inventions and art? He sourced his material from other people and published it for them.

So it's kinda ironic (maybe on purpose?) that you mentioned him.


> as long as the same is true for human works.

and it is.

Anyone can study Michelangelo or Bach, and learn their style.


Depends on the software.

In this case, the companies that make and provide AI models that are increasingly used to interact with me on critical things (banks, public sector services) then yes.

Abso-fucking-lutely they should be regulated like crazy.

In fact I'm really surprised by the amount of people that are not worried by how many parts of their lives are being handed over to be managed by a probabilistic system that is controlled by a private company with next to zero oversight.

There must be a greater liability than "oops, you're right to push back"


Software, no. But maybe eventually AI and tokens are a public utility.


AI companies say people will buy intelligence like they buy electricity.


> Also, the output of the LLM is public domain by law

Why so? Also there is a lot of code in ironically claude and ChatGPT that’s generated by LLM . Yet I haven’t seen the public domain code



The code is not eligible for copyright. If they do not give you a copy of the source code, that does not matter. And if you don't know which parts were generated by LLM, you can't safely reuse the code.


> And if you don't know which parts were generated by LLM, you can't safely reuse the code.

I speculate this could be a real issue in future copyright infringement lawsuits.

The plaintiff bears the burden of proving that the code they claim is copyrighted by them actually is copyright. If it is known that large parts of it were generated by LLM, they’d need evidence to demonstrate sufficient human input to establish copyrightability. If they’ve kept highly detailed traces of the development process, that could be rather straightforward; if they haven’t, it could be really difficult.

Now, that’s true in the US, which never accepted mere “sweat of the brow” as a basis for copyright; the UK courts have, and most of the Anglosphere follows the UK on this more than the US.

The other factor: when dealing with an (almost) trillion dollar corporation, even if you’ll win the legal argument, they may bankrupt you with legal fees before the argument is ever properly heard.

But I suspect the precedents on this topic are going to be established by lawsuits involving far smaller actors.

(IANAL and I speculate only for myself, not any present, past or future employers.)


> The code is not eligible for copyright.

This is very much not what the linked case established.


According to the link:

"The US Copyright Office and federal courts require human authorship for copyright protection; works created solely by AI are not eligible for registration under the current rules."

The Supreme Court declined to consider a challenge to this rule, and so for the moment at least, the rule remains in place.

This means that companies leaning heavily into their LLM use may very well find that they do not, under the law at least, actually own any of their code. As I've read elsewhere there's every possibility that AI code will be the asbestos of the Software Engineering world. Something we'll be trying to get rid of for decades, once everyone comes to their senses.


I think the word "solely" is going to be a tunnel you can drive freight trains through.

So with the asbestos analogy, we encase the fibers in resin and call the whole thing copyrighted.


Or in other words: there's a big difference between public domain and copyleft and it looks like whoever came up with the asbestos analogy was underestimating that difference.


Please explain. That is exactly what the linked case established.


It established you can't assign copyright to the LLM itself. That's very different.


> If they are so adamant on protecting their IP,

What they are trying to protect doesn't qualify as intellectual property. Only 4 categories of IP exist: (1) copyrights; (2) patents; (3) trade secrets; (4) trademarks.

The capabilities embedded in model outputs don't qualify. Machine-generated outputs are ineligible for copyright. They aren't covered by patents. They aren't trade secrets, because the model companies are selling them rather than keeping them secret. And of course, trademarks are conceptually inapplicable.

This leaves the model companies with contract law (ToS) which is pretty inept because it can't bind third parties. And technical measures, like the ones being discussed in the article. And, of course, politics.

Frankly, I think it's pretty ridiculous to even think that models can be protected from being learned from. I feel the Stanford Alpaca team demolished that idea 3 years ago.


The hypocrisy of the pro-AI mega corp arguments makes my head spin. For three years they’ve been using the example of a human reading books and then outputting creative works influenced by them as analogy for training AI on copyrighted works. Now suddenly we’re supposed to not draw the same parallel about a hypothetical person who learned from Claude and is now outputting creative work based on it.


> So, why all this "effort" to protect the model?

Because it's their model and business and they are free to use the free market to do exactly that?

That's their free market rights too. If you don't like it, use another model (which they would be fine with).


> Because it's their model and business and they are free to use the free market to do exactly that?

I mean, nothing stops distillers to find better ways to distill, either. Meaningless cat & mouse games.

> If you don't like it, use another model (which they would be fine with).

Thanks, I use none. It's peaceful this way.


The usage of the output is probably considered legal. The usage of the service for that purpose may not be, and using it at scale in a dishonest way is not, which is what China has been doing. Countless thousands of separate requests abusing the service (which is not a simple static HTML feed, but an AI service request) for every kind of query to soak up the results.

The post is about what's in the local code, but for a long time there has already been modifications made to the request outputs from the major cloud services as they work together to both curb adversarial distillation and to degrade the quality of training China can get from that distillation.

It's likely not to make the answer wrong or bad, but to make it so that any model trained on the output would not gain the benefit of the model's reasoning generalization skills as easily and also identifying markers that might even link back to request IDs.

The techniques talked about in this post are naive and simplistic, largely because they are released publicly.

It's not as much about protecting IP as much as it is about slowing China down or being able to track the effects of abuse. So many people are talking about greedy company this, greedy company that. The world is not made up of caricatured giant money pigs wearing suits with monocles and gold watches. That is a children's view of Marx's exaggeration on free markets. Bad, greedy people do exist, but if that is your only hammer for every nail then you have a problem.

The upper-bound for how good these models can be is so crazy that it is essentially dual-use military applicable to an extent most other technologies are not. It's not only cyber attacks or biological weapons. Most people are not even built to understand the possible threats.

Why does it matter if China gains those capabilities? I invite you to begin to learn about China's behavior around the world. The CCP is darkside material.


> Why does it matter if China gains those capabilities? I invite you to begin to learn about China's behavior around the world. The CCP is darkside material.

Reminds me of this comic: https://xcancel.com/tomgauld/status/571994690289061888?lang=...

None of the superpowers in this world is innocent, and like MAD, more countries have the capability, the better.

I know some of the things CCP do/did. I know some of the things US does/did. I'm from neither, so I don't take sides.

AI's use has been confirmed, or more precisely boasted by two countries in two different wars, and China was not one of these countries.

We have seen the effects of "if they don't know them, they can't exploit them" mindset of NSA for years. Keeping information/technology private is neither beneficial, nor possible. It's only a temporary moat-ish gap. Not a definitive solution.


Certainly the world is full of actions and reactions, nothing is happening in a vacuum. You don't have to be from a country to take sides, but presumably you have some kind of moral compass, some kind of values around personal freedom or the worth of a human life.

There can be a very real cost, because one side comes from an ideology with a history that wants to conquer the entire Earth which caused World War 2 while the other side is trying to prune the planet like a bonsai to prevent it from descending into total chaos to preserve some sense of international order.

Europe was constantly at war, and we helped stabilize it. Middle East as been constantly at war, and if Iran can be sorted then it will be the closest to some sense of peace it's been in a long time.

We used to be in Japan, Philippines, Germany, Vietnam, South Korea, Iraq, Afghanistan and so on. How many are US territories? None. We aren't out there to conquer the globe and take land. We're usually fighting other people's wars for them, because they're up against better resourced opponents. Meanwhile China is over there building artificial islands, ramming other country's ships, creating ideological police stations in countries around the world to harass people and engaging in the most widespread international interference campaigns in human history.

They do not treat their people well and they do not have free speech. The internet is flooded with their propaganda now, because they have a human numbers advantage.

It's true that given time most advantages are temporary, but there's always that slim chance we could slow them down until the CCP collapses and they could become a more normal country.


You sound like you've swallowed pro-US-propaganda hook, line and sinker.

The reason the middle-east is at constant war is because colonialist machinations. Same goes for south-Saharan Africa. And the US is a big colonialist player, just ask Vietnam, South-America, Iran, Afghanistan, etc. They all have been attacked by the US because of US colonial interests. If anything, one could make the argument that the PRC is treading much more lightly than the US.

That said - I'm not defending the PRC by any way; it's a state-capitalist hell hole that's suppressing workers by denying them any ability to organize and whose political class is purely focused on furthering their own interests and that of the moneyed elite, the common person be damned.

The thing is - so is the US.


If you think those were about colonization, you would be well served to examine history closer.


Ah yes. The exceptionalism argument.

There's the good "us" and the bad "them."


I didn't make an exceptionalist argument, but if any country's behavior and values can be measured compared to others, you will always be able to make some kind of decision about where those fall in terms of goodness or badness. Do you not believe in good and bad?


I believe in good and bad.

I don't believe in US good, non-US bad. I also don't believe the same about my religion or my political party, for that matter.

How you measure depends on weights you assign (cultural system of values) and what information you use (media bias).

You can rank in the extremes (e.g. North Korea as worse than Belgium), since they come out that way by almost any set of information and values. Comparing the US to most other countries, there isn't a clear ordering. If you believe the things you wrote, I think the other comment summed it up well: "You sound like you've swallowed pro-US-propaganda hook, line and sinker."

Most countries have similar propaganda, by the way.


Most countries have some kind of story they tell about themselves. New Zealand doesn't have any illusions that it is a superpower. It doesn't have the resources, the talent pool or any of that to even begin to dream of it. That is natural. If New Zealand was powerful, then it would find itself in a position of greater responsibility.

You cannot compare countries that barely have the option of ambition with something like the US and even begin to imagine that it is meaningful. You have the UK, France, Germany, Japan, Russia, China and the US. You can rewind history to name other civilizations.

These kinds of countries are the only ones that matter, because they're the ones that have to answer about what people were thinking when they chose to make use of their power in a way that is relevant on a larger scale. They have to answer about what the reasons were when things went wrong and whether they agree things went wrong. If people are even allowed to talk about it.

It's very cheap to label anything as propaganda without taking the time to appreciate whether it has any merit in terms of the overall behavior of a country or its people. You can always find counter-examples, but how influential are they in the larger picture?


> These kinds of countries are the only ones that matter, because they're the ones that have to answer about what people were thinking when they chose to make use of their power in a way that is relevant on a larger scale.

This is an extreme claim, and incorrect. If you'd like to see counterexamples, you'll see many corrupt regimes in Africa, which did extreme harm to their own people. You'll see many regional powers.

One does not need to be a global superpower to be good or evil.

There's also nothing magical about Germany or Japan. Many countries had similar resources. Both chose to invest those resources into industrial militarism.

One can make the argument for a handful of countries which our outliers for land area or population, but in general, if any country chooses to invest in military and attack its neighbors, it has good odds of success.

> You have the UK, France, Germany, Japan, Russia, China and the US

Germany was an outlier, on the evil end, but otherwise, it's a selection of which facts one picks.

A comparison would require deciding which facts to compare on. For the US, the "evil" argument comes back to things like slavery and the genocide of the native peoples.

One can pick hundred of examples like Guantanamo Bay, fake vaccines in Pakistan, the Tuskegee Syphilis Study, the Tusla massacre, police violence, corrupt court, ...

The US does pretty nasty things, even if they don't always make US news or grade school textbooks.

> It's very cheap to label anything as propaganda without taking the time to appreciate whether it has any merit in terms of the overall behavior of a country or its people

This is an ad hominem, and a poorly placed one. You're discounting what people are telling you. A lot of the people here went through the US school system, learned "US rah rah rah" propaganda, and only deconditioned themselves as adults.

Many of us were where you are when we were younger.


It's largely correct, but you're missing the point. For many countries, it does not matter whether they are good or bad on a grander scale, because their power is so limited that they are largely confined. Countries that exist on the lower end of the intellectual development scale are expected to fall apart.

Not all countries are structured the same, are at the same stage of development, or resourced enough enough to arrive at a sense of responsibility.

That is why these countries are different. The US is the most different, because it is the only one of its kind. There is no other country you can compare the US to.

A lot of negative things about the US are being misrepresented or inflated by propaganda without sufficient context. It isn't all positive, but the US is a huge country and our history didn't begin 250 years ago.

> One can pick hundred of examples like Guantanamo Bay, fake vaccines in Pakistan, the Tuskegee Syphilis Study, the Tusla massacre, police violence, corrupt court, ...

I could come up with way more examples than that, but for each example you have to ask the questions. You don't simply label an event and point at a face to say "bad, they did it and this is who they are". Is it true? Why did they do it. What were they thinking. Was it an accident, intentional, what was the context? What else was going on in the world, and what was the world like then? What were they up against? Was the choice easy or hard? Was the badness mitigated, or was it unmitigated badness?

> You're discounting what people are telling you. A lot of the people here went through the US school system, learned "US rah rah rah" propaganda, and only deconditioned themselves as adults.

Your unbalanced schooling is not a compelling argument for why I am wrong as much as it is a piece of evidence for why you lost perspective when you got older.


> The CCP is darkside material.

And which country has "Black Sites" peppered around the world to detain and interrogate people they don't like?


The US doesn't have black sites anymore and when it did, the interrogation techniques were chosen to avoid physical harm. The results were bad, we didn't like it here in the US even if they were extreme measures for extreme times and so we shut it down. It had a high error rate and generally didn't reflect what we thought was right.

Meanwhile the CCP regularly abducts its own citizens and executes more people than the entire world combined.




> The usage of the output is probably considered legal. The usage of the service for that purpose may not be, and using it at scale in a dishonest way is not

This is literally what the "training AI on copyrighted works is just like a human learning/getting inspired" crowd has been arguing though.

Literally. People have been literally saying that it was wrong because they did this "learning" at scale in a dishonest way.


In some ways it's an offshoot of the honest benefit of search engines already crawling all this content. That has its own conflicts, like just how much of a page's content should you reproduce in the results before it's basically considered stealing their content without benefiting the site itself.

There is a balance to strike, both in search engine fair use cases and AI fair use cases. The major cloud LLMs do double as web search engines now, though they didn't originally. In many cases there's no reason left to click the links they sourced from.

That is a legitimate concern. At least within the US, I think there are nuances around fair use and contract law. A lot of companies are getting paid for having their content used in these models, but many websites had no particular rules you had to abide by and the content was simply public. I think if you're operating under an agreement, then even if there is fair use or public domain content being reproduced by the site you are still bound by that agreement.

Similar to old paintings digitized and hosted on some museum website. It's 300 years old, right? It should be public domain, yet the people who digitized it or provided a service to give you access have some say in how their reproduction can be used. These AI services are obviously very different, but there are laws that can govern how you are allowed to use a service if that service has laid out acceptable usage.

I'm not exactly comfortable with the mass scale that everything was soaked up to train these models even within the umbrella of search services, but I also admit that a lot of the usage was probably quite legal. The potential displacement caused by the resulting trained models on artists or writers is almost its own facet. In practice, whether they ONLY trained on strictly legally acquired fair use content with no errors and paid agreements to acquire even more content than they already do or not, there was enough legally accessible information for fair use that there was no escaping some kind of impact on artists, writers, etc.

With any luck, artforms and skills impacted by technology will adapt and continue to be valuable instead of complete displacement or the dilution of opportunity.


Well it was also problematic when the search engines started quoting the websites in such a way to disincentivize people from visiting the actual website.

> At least within the US, I think there are nuances around fair use and contract law.

The concept of "fair use" as it exists in the US-law system is completely dysfunctional (see e.g. nearly every educational music channel on YouTube), so utterly biased to favour large corporations, that there's very little room for whatever "nuances" you believe exist.

> Similar to old paintings digitized and hosted on some museum website. It's 300 years old, right? It should be public domain, yet the people who digitized it or provided a service to give you access have some say in how their reproduction can be used.

Yes 300 year old paintings are public domain. Indeed there are certain rules for the people/institutions who digitize them. It's not "they have some say", there's actually nothing mysterious about it and it is not similar to Anthropic's copyright heist at all because nearly all of the books they copied were not more than 100 years old.

> there are laws that can govern how you are allowed to use a service if that service has laid out acceptable usage

well where I live, there are laws about what a "service" can claim to "lay out as acceptable usage" instead of the other way around ...

> I also admit that a lot of the usage was probably quite legal

Let's disagree on that. I think it wasn't a lot and the vast majority was not legal. How do you think the LLMs "learned" to speak all these non-English languages? Unless your point is that it's probably quite legal to treat foreign IP like that. Which it may very well be in the US, especially if the corporation is large enough, but imvho it's still wrong.

> With any luck, artforms and skills impacted by technology will adapt and continue to be valuable instead of complete displacement or the dilution of opportunity.

And with any bad luck, these AI corporations will hold frontier models hostage for the rest of time.

I honestly don't want to put that up to "luck".


oh no, the company that illegally used every possible media they could get their hands on is crying that some other company is doing something potentially shady but not illegal? And using that excuse to put in place hidden surveillance systems on their customers?


People keep throwing this idea around haphazardly, but U.S. courts have pretty consistently decided that training on copyrighted works falls under fair use. You may not like it, but that doesn't make it "illegal".


You have to admit that "downloading every book ever written for free from a repository of books that is itself illegal to compile and to run, in order to write a text generation tool" being legal is at least unintuitive, to put it mildly.


It wasnt, that's why they paid a >billion dollar settlement over it, and now license/purchase them. I don't know if the people distilling are licensing those books/etc today, though


I'd appreciate if the down voters explain why. I wasn't making a value judgement.

Anthropic did pay more than a billion: https://www.npr.org/2025/09/05/nx-s1-5529404/anthropic-settl...

And is now buying up a lot of books (controversially, as scanning involves cutting their spines) because that's what the law deems the legal method: https://www.washingtonpost.com/technology/2026/01/27/anthrop...

We know that models like Deepseek are trained on copyrighted books too: https://arxiv.org/abs/2603.20957

The looser use of IP (eg, any characters/celebrities in AI video models) is increasingly mentioned as an advantage of overseas models.


Clearly paying that fine didn't do anything to stop Anthropic from doing it again.

Buying a book doesn't make it legal to publish lossy compressed copies of it.

Also, the vast majority of authors whose work was copied against their wishes didn't receive any of that fine.

It sounds like your argument is that they paid a fine for breaking the law, and therefore it is okay they reap the benefits of breaking the law and are allowed to continue to do so?

> The looser use of IP (eg, any characters/celebrities in AI video models) is increasingly mentioned as an advantage of overseas models.

UHmmm you remember when Sam Altman changed his profile pic to look like a Disney version of his own face? Yeah neither do I.

Clearly US AI models are playing loose with the use of overseas IP just as much, and even publicly flaunting it, as if US-based IP is more worthy of protection but Gibli can suck it.


The grandparent claim was that they were surprised downloading books was legal, I was saying that it's not, as they did need to pay. Whether the law is enough is another question (some cases are still ongoing), and whether the courts are awarding it widely enough is another, but they are facing genuine legal backlash that international firms aren't right now and are more cautious. Several billion is a genuine cost that can move their prices higher in a time of strong competition (see also other announcements with media firms, it's not just books).

I'm guessing the Sam avatar was related to OpenAI's deal with disney to use their characters: https://openai.com/index/disney-sora-agreement/

It's true that "in the style of" (eg. Ghibli) is not currently legally protected, only actual character IP or using the Ghibli name. That's not inconsistent with US IP treatment.


No it's not unintuitive.

Just like I can learn from a book and nobody can make that illegal, so can other people transformative do the same with computers.

Fair use is fair use.


Just like these distillers can learn from Claude’s output. Fair use is fair use.


I don't think Anthropic argues that distillation violates copyright. AFAIK, their position is that it violates their terms and conditions for interacting with their servers.


They violate every website and book TOS that says "don't distill".


I think Anthropic will argue whatever argument is likely to protect their interests. I don’t expect anything consistent or moral from them. My quibble is with all the Anthropic fanboys who repeat this crap.


I'd think the problem with fanboys is that they don't care about the truth of the underlying arguments. They just want to score points for their team. Do you have a different issue with them? If not, why not engage with the arguments yourself?


I know from reflecting on my own beliefs now compared with 10-15 years past that one's beliefs can change, and I don't want to be so cynical as to say that these fanboys don't actually believe what they are saying and only want to score points. I'm sure there's a great deal of commentary that is astroturf, but I think there are plenty of (hopefully young and naive) techno-optimists who sincerely think companies like Anthropic can do no wrong and only move humanity forward, or something like that.

In any case, online debate is not always about changing the mind of the single person you engaged with. To some degree, its performative debate so that other readers may be influenced by your ideas.


To be clear, I'm not saying that they don't believe what they say, just that they don't evaluate arguments critically. It's easy to agree with every argument that supports your existing position, but it's a recipe for disaster and they should be taken on a case-by-case basis.

>To some degree, its performative debate

No offense intended, and I'm certainly guilty of this myself at times, but this is a pretty gross way to talk to other people. It's certainly antisocial on the individual level and I think it's also pretty destructive on the community level- I look to Twitter as a case study here, which flipped from a left to a right-wing echo chamber without ever touching anything in between, which I blame on the design, algorithm, and culture being built around performance. Dunks are not truth-seeking behavior, but they perform very well.


For someone who finds this "not unintuitive" you sure are confused!

"Just like I can learn from a book" - ok. Are you allowed to go to libgen and download a book in order to learn from it, because learning is a fair use?


Has it? Because as far as I can tell those cases keep getting settled out of court before a legal precedent can be set.

For record breaking amounts too.


Maybe "U.S. courts have pretty consistently decided" used to mean something, but I don't think the opinion of US courts should be the standard for anything, anymore.


> U.S. courts have pretty consistently decided that training on copyrighted works falls under fair use.

I don't believe that this has been resolved at all, and there are quite a few pending lawsuits about it at this very moment.


The courts have never said piracy, which is how the training sets were originally built, is legal. There are several court cases still ongoing over this.


Right, so it seems that distilling an AI model is legal too then. At least it is somewhat similar.


Legal vs "They aren't going to let you do it with their service" are two different things.


Screw those poor copyright holders without the means to stop frontier AI labs, amirite?


>Screw those poor copyright holders

In general yes. Cut it down to a reasonable amount of time and I'll care a whole lot more about those 'rights' holders.


It is a violation of their terms of service.

There are plenty of good reasons to not use Anthropic's services. If you don't like their terms of service, do stop using them! I personally think Anthropic's increasingly successful attempts at regulatory capture are even more distasteful.


Oh Anthropic has shown their ugliness in more ways than one I agree. You have to have to done some pretty heinous shit for openAI to look good in comparison.


It was also a violation of the terms of service of those books (aka copyright)


> that training on [lawfully obtained] copyrighted works falls under fair use

Fixed that for you.


Were the copyright owners contacted prior to this lawful obtaining that you speak of? Or after?


I miss the days when tech people were copyright skeptics. Remember when everyone was upset with Disney for our perpetual copyright regime and the destruction of public domain?

Now many tech people are copyright maximalists and 100% converted to the church of Disney. It’s depressing.


I don't think that's right. The problem is that Anthropic is hoarding it and that's hypocritical. If copyright doesn't count for Anthropic, they should publish Claude. If they wanna hide Claude behind copyrights and/or TOS, they don't get to screw with other people's copyrights and TOS and then profit from it.

To call that opinion "copyright maximalist 100% converted to the church of Disney" is, at the very least, hyperbole.


“Anthropic is hypocritical and hoarding data” is 100% compatible with “copyright has gotten out of control and we need less of it”

But pearl clutching over the poor corporations who have their works trained on is much less compatible with a copyright-skeptical view.

And I stand by copyright-maximalism as a rising trend in tech circles. It’s mostly anti-ai, but strange bedfellows and all that.


Imho, you're getting wrapped up around the wrong perspective axis.

Anthropic, OpenAI, Meta, etc. know they illegally obtained all the material they initially trained on.

So claiming any kind of right against anyone else training on their models is asinine.


its not copyright maximalism. people just see the obvious hypocrisy. a lot of people are also fine with some copyright


> loudly stating the foreign labs have been distilling their models for a while now.

They would be stating this even if it weren't true, because it fits their marketing.

While I don't disbelieve the claim outright, I highly suspect Anthropic is misleading everyone about the severity.


Distillation usage still burnishes usage numbers for IPO...

If anything, Anthropic is incentivized to track but do nothing until equity lock up expires.


No sympathy for them trying to “protect” the output of a model that’s trained on data that they didn’t get consent to use. Ripe hypocrisy.


> foreign labs

Apparently not just foreign labs. It looks like xAI distilled Anthropic models to train grok.

https://opentools.ai/news/xai-trained-coding-models-claude-o...


That's less of a worry though since xAI is patently incompetent.


Incompetence is not an excuse for amorality


Oh that’s ok, xAI is enthusiastically immoral.


What's amoral about distillation?


Except when it comes to image models. Imagine is extremely good and extremely cheap. I've been using it to generate book covers for ebooks (old novel short stories that never had a cover, for example) and it's phenomenal. Each cover is about 6 cents


I wouldn't be surprised


I really doubt other labs are distilling Claude using the Claude Code CLI when they can way more easily use the API directly.

I also don’t get why the « protection » on ANTHROPIC_BASE_URL. If I change it to use a Chinese model, the Chinese model will not care at all about the modified prompt. On the contrary, if I’m distillating (which again, using CC CLI would be stupid), I’m not going to change ANTHROPIC_BASE_URL.


Sounds suspiciously similar to the "album title", "Steal this album" by system of a down.

Im not sure why we are dithering on the boundaries of honesty when the entire content LLMs are trained on is stolen.

Are we debating "honor among thieves"?

Of course we are not, or maybe we are!

Does the behavior of a thief even matter to me? only after they do their time. And they will.

I can see the investors perched on the balconies of their condos in a couple years if that.

its a long way down.


"Steal this book" by Abbie Hoffman


The obvious response is the realization that spending trillions on training LLMs is not a viable business model if they can be distilled for a much lower cost.


What does that have to do with CC? I'm not commenting on that being good/bad/legal/illegal, but CC is separate from the models. If they really are doing this maliciously it is because they are trying to ignore my 'CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1' flag (if that still means anything).


The article seems to state as much minus the obfuscation. However justified they are to respond, this can be a slippery slope. We're bound to hear more reports of hidden user data exfiltration.


The have been fucking distilling our websites and writing, even when behind TOS, aggressively bypassing protection mechanisms. They they can fuck right off


Aww man that's rough did someone steal their content and use it to make money without asking them?


> a mechanism to make that obvious.

Say they prove that foreign labs are distilling their models, then what?


Something about throwing stones in glass houses.


>very loudly stating the foreign labs have been distilling their models

Help! Someone else is blatantly ripping off my plagiarism machine!


i think you're being played by their whole ethical high horse PR angle


Not only that, undisclosed behavior of this nature erodes the trust that the software is not compromised by internal bad actors.

Sure the thing we know matches the company interests but as the parent mentions for all we know they are also shipping over your ssh keys and browser cookies.


> If anything, that they thought this was acceptable makes me wonder what else they're harvesting from my machine? PII?

Hah, you just reminded me of this meme I spotted the other day: https://img-9gag-fun.9cache.com/photo/an76Wnz_460swp.webp


I don't see any ethical connection between adding canary tokens to your output to catch people breaking your accepted ToS through ongoing distillation and stealing your PII off of your machine. How are legitimate users in US or China possibly harmed by Anthropic silently changing the apostrophe in Today's or the date separator from - to /?


It's clandestine behavior, we all here assume it's trying to signal whether a user is in China, but this breaks an implicit trust.

How do we know this isn't the work of a rogue developer at anthropic and that they are not subtly switching visually identical characters in other contexts to ship your ssh keys or whatever.

We don't. After the trust that the software doesn't do things it doesn't disclose has been broken you can assume it operates like malware.


> I don't see any ethical connection

Trust isn’t about ethics, it is about individual judgement of how much risk some actor poses to your own interests. Trust is context-dependent

There can be unethical actors whom you have good reason to trust, and highly ethical actors whom you have good reason to distrust


You are talking about "stealing" from people who already used stolen texts to train their models.


So they’re watermarking requests according to your environment variables and maybe changing a string format if you’re in a certain time zone? Am I missing something here? Where’s the five alarm fire?


No fire, just people looking for a reason to be upset. "They didn't tell us they were secretly checking for ToS violations" is their reason this time.


It is not checking that is the problem, it is sending obfuscated information about the user without disclosure. That is unacceptable in any context, let alone a tool that requires an unprecedented level of trust.


I'm trying not to be flippant (let me know if I failed), but most tools with an online connection send back a lot of information about you, with just about the same amount of disclosure (somewhere in the EULA it says they may).

Especially if you play any online games with ranked or PvP, they are likely doing a ton of work to prevent cheaters - and this information is necessarily going to be obfuscated, to delay the cheaters/hackers in their efforts to work around it.

From Anthropic's point of view, they have a class of "cheaters" they're trying to detect - people trying to distill their models. Those people are of course going to try to work around any detection, so you can't send that signal in clear-text where it is easily detected and blocked.


This isn't an online game we are talking about, it is a fundamentally different kind of application, non-deterministic and capable of manipulation. The bar is much higher. They should collect analytics if it is necessary to provide the service, but they should do so transparently. The solution to snare cheaters can't be worth the sacrifice of trust through non-transparent collection of user data. Although this particular information might not seem sensitive, if they are willing to use this method at all, they are probably willing to use it in other ways.


Do you honestly think that no service out there collects basic analytics?


There's nothing wrong with the transparent collection of analytics. I expect any software that I run to tell me what they are sending at a bare minimum, and ideally give me a choice. This is the common and acceptable approach for a software company that deserves your trust. A willingness to cross that line with something small does not lend trust for something big, and there's nothing really comparable for the level of trust that this software requires.


This entire thread has lost its collective mind. Tracking time zone has to be up there with IP address and referer in the list of "trivial things every company in the entire world collects about you", and these things are entirely trackable without your consent or even knowledge. You're gonna have to turn off the Internet.


It seems like the confusion is about whether it is common practice to collect analytics vs whether it is common/acceptable practice to do so without disclosure and to cover your tracks. I recognize that virtually every cloud service collects info. Even something as trivial as IP address and region are not too trivial to include in the EULA.

What is the harm of that disclosure? If you are a company that wants to establish a firm basis of trust because you deliver a sensitive service, it is a necessity. I can't say with certainty that Anthropic doesn't disclose that they collect this info, but in any case, the way they've chosen to implement it does not lend credibility.


It's disclosed in the ToS like it's disclosed in the ToS of basically every service that exists today.

e.g. "Claude Code connects from users’ machines to Anthropic to log operational metrics such as latency, reliability, and usage patterns. "


In your example, service providers may be allowed to collect IP address or cookies or referer which are either critical to how they provide service (IP and cookies) or part of the established norm (browser sending referer).

That doesn’t mean they are allowed to go beyond and scan user’s environment variables


I mean in the end they could have got all this info from IP + headers + API key right? So what is the issue?


Maybe you missed it in the post, they look for the value of an environment variable and send the obfuscated result. I'm guessing this particular variable isn't sensitive for most users, but the reality that they find it acceptable to snoop your environment and hide their tracks is a big red flag.


Based on that info: they will provide you a shit model and sabotage your project. And you pay full price.


Dishonesty seems to be a core value at Anthropic. I find myself wondering how anyone could have confidence in them after their repeated breaches of trust.


I honestly find it crazy how many people trust them for their business needs. For a business, you want consistency and no surprises. With them you get exactly the opposite.


No, that's not correct. There are two types of business. One wants to be steadily growing, but the other wants to move fast and break things and either succeed or fail quickly.


They were talking about vendors, not the business themselves. Even if you are a move at speed of light and break all things org in SF, you wouldn't expect the same sort of behavior from your business vendors like AWS etc. You want reliability and consistency to ensure your own business doesn't have to constantly get rekt by their plans


altmanaltman is correct, I was talking about vendors. If you rely on a vendors service for your own business, you want it to be consistent. How can you plan around and rely on something that is inconsistent?

Your vendors plans to grow fast and break things shouldn’t affect your ability to provide your service to your customers.

Anthropic has not been a reliable vendor. Previously, when they were compute starved, the quality of their models degraded in the weeks before new releases, without warning to their customers. You just suddenly got a worse service. I wrote a little more a few months ago: https://news.ycombinator.com/item?id=47375818 And then there’s the transparent downgrading they did with Fable.

If your running a business, that’s not what you want to be relying on.


Just like we take unreliable machines and make a reliable system by load balancing and fail over we can do with vendors.

Any sane company should not be too deeply tied to any one AI vendor at the moment and be ready and able to switch with a timeline and cost as acceptable to the org.

TLDR: Be prepared to use multiple AI vendors.


Sure but that adds a lot of complexity and cost to your business. It’s much simpler and safer to just work with more reliable vendors to begin with. Why build on quicksand when there’s rock available?


In most of business, there isn't rock available. You have to build on what's there and accept it.


Many vendors and service providers have SLA’s and contracts, at least for the critical services.


No AI vendor does.


> That the provider's business needs necessitate the this behaviour...

No, please don't move past this so quick, because I'm not convinced they have the need, and in some markets (like the US) it is a violation of civil rights, to show people different content based on their ethnicity[1] because those people might have a claim that supersedes anything they might have signed or clicked-through.

That Anthropic did something they could obvious be sued for constitutional violations in multiple countries is shocking.

> what else [are] they're harvesting from my machine? PII?

Assume everything, and yet I think this is more insidious than mere exfiltration, and we should go further: The LLM can respond to these magic quote marks directly, which means it can be trained to give people bad/different advice without those markers being so visible to people using debugging tools.

That's so unethical, the laws on this potentially so severe, Anthropic could be facing unlimited damages, from any one example combined with this article, which means either they have a really stupid management team, or were given a promise of legal immunity in some way.

Neither of those things should be what you should want to base your next big idea on.

[1]: For a simple example, showing an ad for (say) mortgage offers and targeting people by race/ethnicity/gender is totally illegal, but it's also illegal if you make a list of targeting criteria that just happen to select for a protected class.


You mention PII and potentially whether this is it.

From a privacy perspective, this is better described as metadata, it is not personally identifiable information (PII).


Honestly, I agree.

I just can’t bring myself to trust an organization that allows these types of underhanded things to happen in the first place. The fact that this behavior even got to customers raises a lot of red flags for me.


I agree with you and disagree. Like these days expectations of software are through the floor. We expect them to be greedy assholes taking all data they can on the downlow. So why did this particular thing make a big splash? Two possibilities it's astroturfed by chinese labs or it speaks to our anxietes regarding AI. We worry that the AI doesn't serve our interests but rather the interests of the creator. That the advice we get may subtly flawed to sabotage us should we try to do the wrong thing. That the not even the creator is in control and the AI is just doing its own thing.

So any covert bullshittery hits hard.


> That the provider's business needs necessitate the this behaviour

If that's true, that is another reason why it's an illegitimate business.


They have vested interests in continuing the practice which is why they are so vocal about it. It’s about money.

You will own nothing and be happy.


They are the chosen ones. Protecting the world from <insert most recent claim>

Any and all ends justify any and all means.

/s


[flagged]


That's true, I am less familiar with the workings of cloud services than some are (as relevant as that may be in a discussion about a client that users run on their local machines). However, it sounds like you do understand how cloud services work.

In interest of educating those less informed than yourself perhaps you could share with us why the reasoned points I've brought up are incorrect by actually addressing them?


Almost all for-profit large internet platform providers you use have anti-bot, anti-scraping, anti-spam, anti-abuse defenses. This is a new thing that is the same, it's anti-distillation, but its a subset of the same space.

If you have a problem with this, you should have a problem with Google, Apple, Microsoft, Amazon, Netflix, etc.

If that's also the case, then no problem.

But what Anthropic is doing here is nothing new.


[flagged]


Aw buddy, you seem to think I'm trying to hurt you. Furthest, thing from the truth.

I think you might have had enough HN for today. Take a nap and then eat a snack if you still feel cranky. The internet and all your cloud services will still be here when you want to play next.

(Well rested you'll also be able to string together a cogent argument but we're clearly struggling with bigger things here.)


Copying over my comment from elsewhere in this post:

Anthopic choosing to delay their models' invevitable distillation by competitors is their prerogative.

That they choose to implement it by fingerprinting my access patterns without first disclosing is where they shit the bed. It isn't "sneaky" it's straight up sneaky (and dishonest and unscrupulous while we're at it). That this particular instance is harmless doesn't give me much comfort. Who's to say they aren't harvesting PII?

That their actions make sense for their business isn't any reason for people to accept their deceitful, customer-hostile decisions.


Does their user agreement say they won't be harvesting PII?


Are you implying that Anthropic lawyers wrote the user agreement to protect users out of the goodness of their hearts? That's how they receive their big fat paychecks? Very funny.

We all know user agreements exist to strip users of their rights and to absolve companies of wrongdoing. We know that corporate apologists here are also well aware of this fact. When user agreements explicitly grant companies the right to screw over users, apologists are quick to make excuses about how it's all standard operating procedure and accuse people of being uncharitable for doing a plain reading of the text [1]. Yet when people are actually screwed over by companies, those very same people blame users for accepting the user agreement. It's a bad faith system.

User agreements are nothing more than power plays by exploitative companies.

[1]: https://news.ycombinator.com/item?id=47953501


> by fingerprinting my access patterns

It's based on whether your timezone is in China and your hostname matches a blacklist. Literally 2 bits of information. Not much of a fingerprint.


That's what it's based on right now, anyway. What other bits of info will they add as the Chinese work around this spyware?


Anthopic choosing to delay their models' invevitable distillation by competitors is their prerogative.

That they choose to implement it by fingerprinting my access patterns without first disclosing is where they shit the bed. It isn't "sneaky" it's straight up sneaky (and dishonest and unscrupulous while we're at it). That this particular instance is harmless doesn't give me much comfort. Who's to say they aren't harvesting PII?

That their actions make sense for their business isn't any reason for people to accept their deceitful, customer-hostile decisions.


Would a filter like this make it seem less likely that they're harvesting PII? Why would they need this if they were tracking all user queries with a finer-toothed comb?


If by a "finer-toothed comb" you mean telemetry then I don't quite see it as comparable to this situation.

Telemetry is disclosed in privacy policies, it can usually be opted out of and if not that, then it can be blocked by a firewall. Steganographically fingerprinting customer's network routing when they consented to your tool reading a txt file is a different problem. Anthropic has demonstrated capability and willingness to embed arbitrary obfuscated data in their comms streams and that's a dangerous precedent to set.


I'm using "sneaky" here to refer to anything that's not very obviously stated but anyway

> That their actions make sense for their business isn't any reason for people to accept their deceitful, customer-hostile decisions.

While I agree it's a dangerous precedence to set, I think this is a "vote with your wallet" sort of situation. They shouldn't do it, but from their POV this is what they need to do to offer the product they do at the price they do. If the product wasn't compelling people wouldn't accept that they do this. However they've decided if you want their product you have to use their interface and whatever spyware it comes with, so it comes down to, is the value proposition good enough that people will put up with it? As of today, the answer is unfortunately yes


Thanks for the considered response.

> I think this is a "vote with your wallet" sort of situation.

I agree a 100%.

> is the value proposition good enough that people will put up with it? As of today, the answer is unfortunately yes

I don't fully agree with you here and I think the jury is still out on that.

In any case, I look forward to seeing international markets responding to the current situation.


It doesn't generate private returns but investing in your worker population definitely nets returns for the public.

Also, some might say that upending the de facto copyright regime in favour of AI companies was an altruistic gift.


If you want to "invest" in society and get a return, by all means, buy government bonds and support the government's policies to do these things. You get a return, and the government gets money to invest in society. It's a $30 trillion dollar industry.


I'm not sure I understand your point. What benefits does treating social/national governence as an industry bring to society? Besides the govt funds it operations through taxation in addition to debt.

Your unstated assumption that investments in AI were private and therefore beyond question simply isn't true.

The AI industry's profits depend to a large degree vast textual corpuses that they acquired and trained on in either straight up illegal or at the least in legally murky contexts; public and private knowledge form the backbone of these enterprises. Said industry then serves AI models to customers via data centers that again off load the cost of inference to the public.

Of course, AI companies could step up pay up their fair share and people will happily treat the prerogatives of private enterprise as private. But until then I think certain quarters of society will continue to believe, rightfully so in my opinion, that AI companies are on the hook for unacknowledged debt.

Of course morality doesn't inform law so the current situation isn't illegal despite how egregious it is. But the law can only attempt to deliver justice, it can't guarantee it. Which is why open, honest, collaborative discussions are important to have exactly now.


If you think AI is built on copyright infringement, that's fine, and I somewhat agree with you. That's immaterial to the argument.

My point is that investment here -- "the outlay of money usually for income or profit" -- isn't just giving people stuff altruistically. Investors, here investing in AI, are trying to get a return on their investment.

They aren't just spending money to spend it.


I see your comments scattered across this thread with most converging on this thesis: "The US govt will regulate away the ability of US corporations and individuals to use unsanctioned AI models."

Which is a fair thesis. I've seen you counter people's predictions of how they think things will pan as a consequence. But what I'd really like to hear is what you think happens (in the US and internationally) as a consequence of such regulations?


USA will experience unprecedented prosperity from influx of global talents and capital who seek to amplify their productivity and profits by 10x.


Why would global talent and capital migrate to go to a country with greater regulatory barriers? Why would "some models are illegal to use here" be a selling point for the US?


Because they have “superhero serum”. Of course it doesn’t matter to you if you don’t want to achieve greatness, but some do.


Well, I for one hope you are horribly wrong.

And something about this train of thought aligning with just the general quality of your comments gives me hope.


Well. We will see about that. Wishful thinking can be cute.


Right back at ya buddy. Nothing quite like seeing a braggart humbled :)


Why is the grievance petty?


Well you seem to clearly be in the know. Care to share?


The problem with this line is that it becomes an argument against me, rather than self education. There are many comments about this on this very post, and it's recent headline news!

The reason I asked in that way was to see if GP actually knew anything about what they had strong opinions about.


On the other I think asking someone for a citation to their rather strong claim isn't incorrect. Digging through unstructured HN comments isn't my idea of a good time and being sent off on a treasure hunt by fellow commentors isn't why I come to this website. Collaboration is how we move forward!


I'm not sure I fully understand this "society as a dinner crowd" metaphor you're painting here - who is the govt here and what are taxes and most importantly what is the food everyone is eating?

How does the wealthy paying 40% of the income/consumption tax collected by the govt translate to them feeding 40% of population (which they don't)?


I had exactly this happen to me, suffered great monetary losses and had my identity stolen. I've learnt my lesson and have moved on to 1password.

At the end of it I couldn't help but reflect on my foolishness. I realised just how much better I would've felt if only it had been an American, Canadian, or European Googler who stole my data. It really is the worst when malicious entities are Chinese, Indian, or Pakistani. Just the worst!!! (/s)


I am curious how this will play out legally.

Surely UI enough isn't enough to prove that source code was plagiarised?

In the event Papermark chooses to sue how will the defendant defend themselves short of presenting their own (possibly) closed source?


> I am curious how this will play out legally

I am curious if/how YC will handle this to get ahead of earning a reputation of being a den of scammers - a few months after the Delve scandal


>I am curious if/how YC will handle this to get ahead of earning a reputation of being a den of scammers

flock is a YC company, so it's pretty clear that YC does not care about a negative reputation. as long as it makes money, nothing else matters.


> it's pretty clear that YC does not care about a negative reputation.

Perhaps not what the general public thinks, but I assume YC cares a lot about its reputation among VC firms that fund its companies, because VCs don't like being scammed (directly, or indirectly through unknowingly funding scams)


Let's recall that YC used to be run by Sam Altman of all people, which confirms what you are saying.


Many YC companies do bad things, and I guess they do so independently. There may well be repercussions for the most egregious cases, but I suspect a lot of ill-behaviour simply flies under the radar.

For example only yesterday I got spam from an YC company, Polymath, and I replied back asking where they got my details from - no response yet. Once I get something I'll make a GDPR subject access request, then a deletion request. I hope the overhead of that causes them to rethink their spamming campaign.

But I'm not going to complain to YC about it.


I have also gotten spammed by a YC startup, but they spammed an email that I use in git commits, and lead with "I saw your fork of $POPULAR_PROJECT, pretty cool!" or something like that and then continued to pester me with their drip program even as I replied asking them to never email me again.


> Many YC companies do bad things,

My comment was not about doing a generic bad thing - it was about scammy behavior in particular (which ties to the Delve incident). YC depends on the VC ecosystem to fund its companies, and no VC wants to be scammed. If a reputation of cultivating/condoning/obliviousness scammers takes root, that would be bad for business.

> But I'm not going to complain to YC about it.

I am not complaining, or even expecting a moral decision. I'm legitimately curious how this will shake out, for purely capitalistic, reputation-management reasons.


Good luck with referring to GDPR. Try clicking through YC startup list and see how many load GA and other trackers onto their landing pages without a consent banner or even a privacy policy sometimes. It’s baffling.


Most likely, Papermark would compel Corgi to disclose the source code during discovery.


I didn't realise that one could forcibly require a competitor to disclose trade secrets.

Now, INAL of course, but I would think this sort of mechanism would be quite gameable from both sides ( i) a wealthy competitor legally forcing a promising upstart to reveal source ii) a copycat working out some kind of arrangement where the code itself is licensed to them via shell company based overseas.)


As with most legal hacks, the courts figured this one out long ago :).

If someone is trying to dig into their competitor's trade secrets via discovery, the court offers multiple ways to safeguard against that. The defendant can identify information as a trade secret and ask that it be protected in some way - for example, the documents may be restricted to "Attorneys' Eyes Only", so while the plaintiff's attorneys can review the material, the plaintiffs themselves are barred from reviewing it. Or the judge themselves may get involved in an in-camera session.


There are software engineers that specialise in source code analysis that lawyers will often use in these cases. The engineers will be given access to source code in secure environments where they're not allowed to bring any device in or out. They review, analyse, and write up a report using pen and paper, that can then be reviewed by the lawyers.


Absolutely. It was very similar to one of my first jobs: "Legal Technical Analyst". Not as much time doing deep source analysis, but basically translating things for lawyers: "So as far as this claim of copyright/plagiarism... this block here, that's CS 101 stuff, that block there, that's novel, and does x, y and z".


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: