It’s still not thinking just because we call the reasoning texts “thoughts”. To imply otherwise in the direction of these models truly “thinking” is just a failure to understand the technology.
Later Wittgenstein said, basically “meaning is use” - words mean their usage, there’s not some root unshiftable core of meaning behind them. If usage shifts because people keep saying LLMs are thinking, then, well they’re thinking.
Yet similarly, Paulg could also come out with “When I put it to sleep, my MacBook really is asleep. We need to better grasp a world where Macs are awake or asleep and perhaps even dreaming”. This points to the issue with assigning the idea of thinking to LLMs, rather than saying they’re calculating or computing or processing when they’re running their digital matrix computations. Something important to us about thinking has nevertheless not been captured in assigning terminology used for the human experience and applying it to an electronic device. And in no world is my thinking process the same thing as an LLM computing an output, even if those outputs lead to the identical sentence “the cat sat on the mat” - functionalists will tell you that this doesn’t matter and functional equivalence is equivalence but it’s hard to believe that without looking at basic empirical facts about what’s behind the outputs.
Ultimately Paul has noticed something about language from his own first principles, not about thinking.
It does not. But it does make the world that little bit more annoying, like whenever someone posts anything about AI these days there’s that select few that love to pose those absurd “okay but we’re just electrical impulses too, how can we be sure LLMs aren’t actually conscious and alive too?” “”””philosophical”””” questions.
Things like these make my blood boil for how unserious and completely off-the-wall stupid they are.
Hell, someone the other day equated a plane to a bird because both can fly… this is the type of thing we ought to expect from children just learning what philosophy is. Not adults.
In German an airship doesn't "fly", it "fährt". One could have reasonably invented new words for 'flying' depending on whether their motion is caused by flapping wings, buoyancy of light gases, rotating propellers, jet engines, or being fired out of a canon's barrel.
This is the thing about language, you can have strong opinions about it but language evolves through use. Historical attempts to stop this sort of natural drift has just resulted in dead languages nobody speaks like Sanskrit and Latin.
1) I'm with you in that I'm critical of people who mindlessly say things like "Claude lied to me", and maybe even "ChatGPT hallucinated".
2) On the other hand I'm also critical of people who assert that "Ackshelly you didn't do any 'work' by putting the crate from this table to that table" or "sweet potatoes are not 'potatoes' at all, so don't call them potatoes".
In physics, words like 'work', 'power', 'heat' have very specific meanings. The 'heat death of the universe' is thought to be near absolute zero Kelvin a.k.a. frigging cold, no heat anywhere. The meaning of 'work' in physics has a certain overlap with what is called 'work' outside of physics (to wit, a number of different things). A typical biologist or taxonomist's view on species is much more nuanced than the ordinary guy-with-some-general-educational-background's view on the subject: the latter will tend to see species as an absolute given which science just has to uncover. Sure scientists can err but ultimately the truth of biological species is often conceived as incontrovertible by the informed layman instead of what it is: a tool for thought, a human measure brought in to help make sense of chaotic data. Darwin already said as much.
Maybe physics should never have chosen terms like 'force', 'work', 'power'. As a German speaker let me remark that especially 'power' is a bit puzzling; German has 'Leistung' instead which I think is a bit clearer. But even when scientists go and create their own terms like 'energy' or 'quantum leap', neither of which existed before being introduced with a specific scientific definition attached to them, the public often goes and appropriates them anyway. So "he's got such a cold energy" ("cold" being associated with lack of thermal energy) and "a massive quantum leap" (i.e. "a sizeable advancement" instead of "the smallest possible change") are colloquially possible.
Look at more squishy terms from squishier sciences like 'language' as used by linguists and things become even more difficult. Sure, "the language of Islamic architecture in early medieval Spain" is a nice and valid book title and no linguists will take exception. But talk about the new research on the "language of birds" in the newspaper and you will ruffle some feathers.
One has to acknowledge that people are challenged with finding words to describe a new phenomenon. Sure, we could stop at 'calculate' or 'compute', as in "ChatGPT computed some falsehoods" instead of saying it 'lied'. But even when I say that "ChatGPT's answer was wrong" doesn't 'answer', alone and by itself, already imply a degree of intentionality that is (or should be, as far we can tell) absent from a process that purely relies on solving one big formula with a pinch of randomness thrown in? Then again, "what is the calculator's answer?" was already used in the 1970s when pocket calculators were all the rage.
Poor choices are still poor choices. Words are very curious things with very infinite meanings, shifting between context. There’s probably something to be said about how we shape our languages and how ambiguity of meaning can prevent those who just don’t know from knowing. You know what I mean?
Sure maybe the English language is choosing to adopt the word “thinking” for what is going on here. Some people are choosing “reasoning”. It’s all the same to me though. The thought outputs aren’t any different that the final message for me, regardless of its shape.
I don’t have a better alternative, but I’m pretty sure this just makes people who aren’t in the industry dumber when they see it.
It doesn’t really matter. There’s no way the training loop necessary is sustainable and these things are a fad. How long? Idk. But language will change and there will be a day where using an LLM requires as much training as reading old religious texts. Got to learn more language.
You can probably div up some factions just based off voting data. Red team made it in with just about 1/3rd of the voting population. About 1/3rd of the voting population didn't vote. Just if we're trying to put some numbers to something, those should be verifiable at least. Doesn't say why 1/3rd didn't vote, so it's just 1/3rd unknown faction.
MongoDB and Grafana seem to be doing fine. Idk that this logic holds up at face value. Like the HashiCorp, Elastic, and MinIO license changes pissed people because it was a rug pull on the community of people supporting and consuming those projects. If they had started as AGPL licensed and took the Mongo and Grafana approach, who knows what things would look like today.
So yeah... Google Gemini escapes too. Good stuff people. Seems nobody is going to jail because the CFAA is weak here and arguing some form of criminal negligence hasn't been done yet (probably weak anyway).
> If you're writing for processes with formal highly structured content like manuals, specifications, form content, procedures, information, that sort of thing,
Yeah... no. The people doing this lack the communications training, see the output has the necessary information, and regurgitate it with no effort or care. This needs to stop.
We spent decades format building to make it easy quick and easy to get through something like a runbook. If your commands are bulleted instead of numbered and code blocked, it's wrong. If you didn't crawl through the playbook, it's immoral to hand that to me, you're wasting my time with untested slop.
This is a hill I will die on or absolutely start slaughtering people on. I just refuse to deal with this crap.
I used an LLM a year or more ago to generate a description of our SDLC for a compliance certification. Perfect application of LLMs IMO. Do you want to die on that hill as well?
A lot of documentation generated in corporations has marginal real value.
> If your commands are bulleted instead of numbered and code blocked, it's wrong.
That’s easy to instruct an agent to do. Put it in a skill.
> If you didn't crawl through the playbook, it's immoral to hand that to me, you're wasting my time with untested slop.
Using an AI doesn’t absolve the user of their responsibilities. This is an easy problem to solve, just make sure teams know what’s expected of them and encourage everyone to push back (professionally) against offenders.
If someone sends me AI slop, they're getting chewed out and told I'm not doing it and I'll even tell my boss "no" and why. If I got fired over something like that, then it tells me everything I need to know about company and the leadership's priorities. I'll die on that hill.
To be completely fair though, I'm against the behaviors in information transfer I've been seeing. I've pass along AI generated runbooks, but they look nothing like the default outputs of these models. It's because I took time to apply all the writing knowledge I like to see in my curation. If people are doing this, I can't even tell it's AI writing. My work is done in minutes instead of deciphering so BS pseudo language they developed in their AI workspace (people really need to turn off those memory features).
---
Edit: Also if I'm the guy receiving a security report and it's AI generated and poorly formatted, I'm failing you short of producing something for a human to parse. Simple as that.
> Also if I'm the guy receiving a security report and it's AI generated and poorly formatted, I'm failing you short of producing something for a human to parse. Simple as that.
It’s clear that you’ve never worked for a compliance company and probably have never been involved in a compliance project. As such, you don’t have the context needed to participate usefully in this discussion.
If a security report is written for any particular human at all, I’d consider it to be a failure as an enterprise policy document. It should be written for The System, not for the boss; and LLMs are perfect for producing ritual boilerplate.
We're probably just going to disagree here. These LLMs are built to serve humans. They either need to make the system transparent for the operator or be limited in use to tasks that can be proven in whole (with code that can't be revised without human approval).
Should you think it is wise to trust the machine that can't differentiate subject matters in a chat styled context, you have fun with that fluster cluck when it blows up.
Like Fable is highly useful, but it's really bad at keeping it's responses straight.
In fact, that "it's not X it is Y" pattern always crops up when it reasoned about the idea of X and I never fed it that. It's literally doing that because it can't predict that I'm a different entity despite it being able to say I am a different entity.
Edit: Clarification by removal of incomplete sentence fragment. Edit2: Clarification on the "proven in whole" thing.
And just say: LLMs only amplify the the knowledge you have, even the best models o use like Fable still suffer from promoting false narrative as a chart progresses. Often it’s stuff that can be ignored like the “not X but Y” crap it dumps because its reasoning had assumptions it invalidated. Sometimes it’s directly in your system architecture, because you never expressed preferences for solved foundational issues, you end up with generic http handler setups or whatever the hot web thing is today.
Takes knowledge and lots of it to really be on top of when these things go down a failure mode path.
We already have practices to deal with AI. It’s called user space. Properly air gap the AI, stop creating routes to open internet, don’t run network wires into the faraday cage.
Why the heck someone would risk having an open bridge to the system is beyond me. Like maybe get used to using remote hardware screens (KVMs? It’s been a while since I’ve done datacenter), we have solutions for this that was absolutely skipped.
reply