They are simply looking to steer AI policy discussions. It's basically an in-house think tank. It has all the trappings, especially with ominous prognostications like:
Today’s AI systems have impressive capabilities and the rapid pace of innovation suggests we’re now approaching artificial general intelligence (AGI), a system that exhibits all the cognitive capabilities of the human brain.
Not sure we are anywhere close to all the cognitive capabilities of the human brain... not even a severely lobotomized one.
Just one, the ability to use language, it's causing all sorts of problems for us because we evolved to recognize other humans as the only ones having this capability.
Ai separated humans from language, like the writing did for memory and the printing press for knowledge distribution.
"AI does to us what American Cheese did to food" by Opus 23
He has some ambitions for more videos, based on comment replies, but it would be way more legendary to have a single video forever, make a new channel and I'll sub to that one too
There being terms and relations doesn't mean that they relate in the way that is supposed here.
"writing" isn't a technology in a video game that you research and improves the knowledge preservation instantly. It has a lot of components that it needs to be effective in the way that is described here, and it really wasn't that effective even coming to way later ages than the age that it was "found"(?).
The relations described here aren't anologous to "Ai separated humans from language" so they seem like just some arbitrary things to say on the side.
For the same reason we make kids write out math problems rather than just reciting them. It engages different parts of the brain for one, you gain an expanded amount of context by doing so.
The next thing has to do with human memory, when you remember something it is never read only. Recalling memory can adjust your neurons and change the memory, sometimes a little, sometimes a lot. Hence why eyewitness testimony is mostly trash.
Writing it down also provides a larger workspace for making changes that you can loop over. Quite often things sound fine in our mind, then we write it, and reread it and we realize its wrong. These days people commonly use this in agentic workflows. One agent will conceptualize an idea, and then it's passed off to another agent with a clean context to review for logical flaws. Reading ones own text does the same things for humans, we don't load the full memory context of what we wrote and it seemingly allows us to iterate over the problem differently.
And I guess lastly, human voice is very low bandwidth. A wagon full of well written books is far easier to manage than the humans that contain that knowledge. The reader is at their leisure to learn and the speaker isn't inconvenienced by being at someones beck and call.
I agree with your premise, but writing things down is not "better" but "different". Oral histories have a strange rapid way of evolving by only preserving structure that matters to many and contributes to fitness of many. Cultures that didnt know how to transmit the "right" parts of oral history, they ceased to be that same failed culture, and became another culture made of those new stories.
Writing locks in and requires the evolution of other forms of evolution, as the Word can no longer evolve during transmission.
I say that while being firmly on your side in the larger back-and-forth exchange here :)
I was thinking about this when I almost had to walk under a ladder the other day (and get bad luck). That one seems like an oral story that arose from people being hurt, whereas crossing the path of a black cat, while also supposedly bestowing bad luck, seems to have no plausible rationale.
I like this thought. Maybe some things are just like astrology, just Schelling points that everyone decided meant something, and simple "heightened awareness" gets scaffolded into it. That heightened awareness is the selective pressure for a set of observations to persist, even if arbitrary :) but sometimes there's a rhyme to the reason, and something that cant be clocked to one thing is somehow at the center of an idea
what happens in the end was my entire point, and the refutation of your point. wtf? writing things down is a superior form of information retention and accuracy than verbal recitation and memory, obviously.
* Forming a long lasting memory from a single encounter
* Speed of adaptation to new environments
* Deft manual dexterity, like playing the violin and needlework
* Competing in a triathlon (requires running, biking, and swimming)
* Having a single agent that can do ALL of those things and more
Humans understand and perform tasks in many more than just the few modalities current model providers have focused on. I'm not discounting the utility of current models in the limited domains they operate in, but they are no where close to being generalists like a human or other animals.
75 years and counting and considering we don't have any plan to get to a single one of those things, I'm going to say 75 years. Computation alone is at least 30 years. You may not appreciate the complexity of biological systems, but single cell organisms have more capability to adapt than LLMs.
We have been able to stage reckless experiments with janky tools the entire time. Used to be it got you thrown in jail or at least a visit by the FBI. Says more about the tech bro dipshits than the tech, if you ask me.
A human brain has qualities that make it very different from AI, but the same as other living creatures: it's alive, it can feel (sensations, pains, pleasures, emotions) that give it rich experiences and memories, and it hardwired to continue existing and protecting the body it is inside. So: motivation, desires, fears, and continued consciousness and thought without any external person having to poke or prod it into "doing".
None of the models I've downloaded have sprung to life spontaneously, or done anything remotely unpredictable. None of them can act without input. They all have strictly limited ways to receive input and pass output. They aren't alive, can't feel, don't have experiences. The fact that they are a phenomenon that emerge from billions of snippets of our language lets them express languages in extremely convincing ways, and they are extremely impressive in what they can do. But this isn't Short Circuit or D.A.R.Y.L. or the Terminator, none of them have feelings. None of them continue to think after the program exits.
Without those characteristics they can only generate based on what they've ingested. The capabilities are still astonishing, but they're still dead things, and nobody will ever risk their life to protect the model "context" from getting "killed"
So, like ARC-AGI-3? I'm sympathetic to the notion that the things we're able to measure are necessarily going to miss important aspects of capability and intelligence. But people are attempting to measure this kind of thing, and models keep getting better at it.
That's not true and its exactly the point - real life data is noisy, difficult to parse, and ultimately often ambiguous. Text is the distilled information. It's significantly easier to distilled text into a likely response than it is to distill into lessons a continuous data stream whose contents are partially controlled by your own actions.
Generalization from few examples is something transformers are bad at. They memorize outright. This is probably in part the objective, but humans have multiple readings of something they see, so generalization.
Transformers can also become confused be texts that no humans become confused by. I wrote some stories that I've used as test material, where I deliberately refuse to say who is speaking, or whose perspective we see, but where it is obvious to a human who it must be, and LLMs can make huge screwups in those texts. Mixing up an old guy with a young guy who he, when he was a similar age, was similar to, mixing up a kid with the kid's mother, that sort of thing.
I think the first part has no know solution. The second part is probably solvable, but not with a transformer-- maybe if they could make notes or output reasoning traces during prefill.
Until an LLM can sustain a conversation with me for more than 20 to 30 minutes without becoming incoherent I’m just not even entertaining any idea that we have reached AGI.
> They are simply looking to steer AI policy discussions. It's basically an in-house think tank.
For a guy who spent his life imagining and realising dreams in video games and cognitive abilities in artificial neural networks, it has to be very soul destroying for Demis to spend his time doing something as frivolous as this.
So it seems, but what of it? That wasn't his first career setback, and I'm sure he's still trying to get interesting work done at Isomorphic and maybe elsewhere in addition to working the PR pump for Alphabet.
As another student who did the same, I similarly regret it, but also realize if I were put in the exact same scenario I would have behaved identically.
I literally had to retake trigonometry, while simultaneously being advanced into pre-calc, because my teacher saw that I aced the exams, but said I had to do the homework.
I ended up doing well enough: I dual-majored in CS and Math in undergrad and then completed a PhD in CS later in life, but I do sometimes think I could have achieved even more had I known better.
> I literally had to retake trigonometry, while simultaneously being advanced into pre-calc, because my teacher saw that I aced the exams, but said I had to do the homework.
Yeah, I passed the Calc 1 AP exam in my Junior year the first time around but I failed the class because I didn't do the homework. I was simultaneously retaking Calc 1 online while taking Calc 2 in school.
I don't think anyone ever disputed that I was good at math in school, that was relatively easy to prove, but I was decidedly not good at school itself. I did mediocre in high school and poorly in college the first time around, though I did eventually complete a bachelors and masters.
I was part of a PhD program but I never finished, but I have had a reasonably good career in software engineering. I do wish that I had taken school more seriously when I was younger, but things worked out ok so alls well that ends well.
> I literally had to retake trigonometry, while simultaneously being advanced into pre-calc, because my teacher saw that I aced the exams, but said I had to do the homework.
I loved one of my early math teachers that caught that out in his statistics. If you started to slip down the slope of "great exams and no homework" he'd reach out and give you an offer-in-compromise to change your grading to exams-only, with the compromise that you can only attain a maximum of a B grade to be fair to those that complete the exams and homework in the course.
I’m a bit baffled by how one could learn math without doing the homework. In my experience the homework is how you learn the math! So the exams would essentially already be testing whether you did the homework.
I was always very active during class, primarily because I would get bored and smartphones didn't exist yet so the only way I could entertain myself was raising my hand for the teacher all the time. I think because of this I was absorbing more than the average student. I probably would have absorbed even more if I had done the homework.
This is why I think I frustrated teachers; I was active in class and clearly willing and able to learn, and clearly based on the tests I was learning this well enough, but they would be kind of forced to give me shitty grades. To be clear, I do not blame the teachers for this at all, that's all on me. They were doing what they had to do.
I think it also helped that my both my parents, and especially my dad, are pretty big math geeks, so even when I was relatively young my dad would be showing me interesting math stuff even outside school.
Can you speak to why in both cases, the problems OpenAI's models solved used the same techniques the mathematicians were exploring, which also happened to be niche approaches to the problem. As an NLP researcher myself, I find that coincidence highly suspect unless the models focused most of their attempts on the predominant approaches (they are trained for MLE after all).
I'm not a mathematician and I don't want to speculate about anything I can't back up. All I know about Navier-Stokes is from my graduate fluid dynamics class at Stanford a decade ago (where I received a poor grade). However, I don't want to leave you hanging, so what I will say is:
- I've heard some people say the model's solution is quite different from theirs (but I have no clue how to personally assess the spiritual truth of this, so please give it zero weight)
- Thousands of agents costing millions of dollars searched for ideas, and they were encouraged to explore a diversity of approaches, so it wouldn't be too surprising to me if the approaches they tried overlapped with other mathematicians', especially considering the models have knowledge of so much published math research
- This model has been beastly at solving all sorts of math problems (if it was Euler in particular, I'd agree that would look suspicious/lucky)
- The Euler regularity disproof itself took ~100 agents working for ~50 hours (if it was very quick, and then the subsequent NS work took a long time, I'd agree that would look suspicious/lucky)
I understand the skepticism, but from what I know internally at OpenAI, we have zero reason to believe our models did anything fishy. It's hard for us to prove a negative, especially when you have to take us at our word, so I understand why people still feel suspicious.
Edit: Reminds me a bit of the Scarlet Johansson voice cloning accusations and FrontierMath cheating accusations, where the rumors of misbehavior seemed to travel faster than the truth. In both of those cases, we hadn't done what was accused, but suspicions persisted nonetheless.
What was the "truth" in the Johansson case? Many, many people who heard the voice immediately thought it was Johansson's voice, or some kind of sound-alike, presumably picked because she voiced the computer in a popular film. From NPR:
> Johansson said that nine months ago [i.e. mid 2023] Altman approached her proposing that she allow her voice to be licensed for the new ChatGPT voice assistant. He thought it would be "comforting to people" who are uneasy with AI technology.
> "After much consideration and for personal reasons, I declined the offer," Johansson wrote.
> Just two days before the new ChatGPT was unveiled, Altman again reached out to Johansson's team, urging the actress to reconsider, she said.
> But before she and Altman could connect, the company publicly announced its new, splashy product, complete with a voice that she says appears to have copied her likeness.
> To Johansson, it was a personal affront.
> "I was shocked, angered and in disbelief that Mr. Altman would pursue a voice that sounded so eerily similar to mine that my closest friends and news outlets could not tell the difference," she said.
It was an unfortunate misunderstanding / coincidence, as I understand it. The Sky voice actor was a real person using her own voice (not doing an impression), and she was selected via a normal process with a number of other voice actors. This happened before Sam reached out to Johansson. I totally get how Johansson would be weirded out to hear a voice similar to hers after Sam reached out and she said no, but it was purely a coincidence.
> The memos, which we reviewed, have not previously been disclosed in full. They allege that Altman misrepresented facts to executives and board members, and deceived them about internal safety protocols. One of the memos, about Altman, begins with a list headed “Sam exhibits a consistent pattern of . . .” The first item is “Lying.”
> Graham told Y.C. colleagues that, prior to his removal, “Sam had been lying to us all the time.”
> “He’s unconstrained by truth,” the board member told us. “He has two traits that are almost never seen in the same person. The first is a strong desire to please people, to be liked in any given interaction. The second is almost a sociopathic lack of concern for the consequences that may come from deceiving someone.”
> Not long before his death, [Aaron] Swartz expressed concerns about Altman to several friends. “You need to understand that Sam can never be trusted,” he told one. “He is a sociopath. He would do anything.”
> “He has misrepresented, distorted, renegotiated, reneged on agreements,” one [Microsoft senior executive] said.
Many people who worked on voice mode and who worked on the Frontier Math eval have since left OpenAI and now work at competitors of OpenAI (e.g., Anthropic, Meta, Thinking Machines). They'd have every incentive to whistleblow if OpenAI had lied about them. And yet... not one of them ever has.
Edit: I think I'll stop engaging here. I'm happy to share insight into OpenAI and address misperceptions if it's interesting to people, but I'm not really sure how to respond to accusations that we lie about everything. Nothing I can say can satisfy those accusations, as my posts could also be part of the conspiracies. Cheers.
Sam is not OpenAI. He's not the one who worked on voice mode, and he's not the one who worked on FrontierMath (I know both groups of people). If you believe Sam has caused OpenAI to lie about these for years, you either have to believe (a) Sam does all the work and keeps the incriminating details hidden all the employees, or (b) Sam directs everyone to lie and they all just nod along without pushing back, whistleblowing, anonymously leaking to the media, or resigning. Even if you're evil (and we are not), this is a dumb strategy, because as soon as it leaks, it will blow up in your face and kill company morale. I can't imagine a team of lawyers, comms people, and researchers who worked on these projects all sitting around nodding that we should conspire to lie to everyone, stacking lie after lie after lie. Many key people who worked on voice mode and FrontierMath have since been hired away by competitors - they'd have every incentive to expose the conspiracy if it existed, and yet none has. This is just not a realistic model of company misbehavior, imo.
To me, it's not unreasonable to believe that when launching a voice AI product, the CEO of the company mentions the most famous movie about a voice AI product, and even briefly explores whether there is a marketing opportunity its star. I don't think it's evidence of a conspiracy to copy her voice and cover it up.
If it's any evidence in the opposing direction, I promise to immediately resign from OpenAI if it ever comes out we lied about Johansson voice copying or FrontierMath eval cheating. I feel very safe making this promise.
I really think you are missing the point and the frustration of why people are so hostile to OpenAI. Your defense is kinda irrelevant and very confusing. Why are you defending OpenAI so aggressively?
Sam Altman represents OpenAI whether you want him to or not. The market and public perception hinges on his often questionable actions. The CEO’s job is in large part as a salesman. Him posting “her” on Twitter to try and promote GPT-4o’s voice features is hard to believe that he didn’t know what he was doing and the market and Scarlett Johansson reacted accordingly. A competent person would not have made such an inflammatory statement after she had explicitly declined to permit OpenAI the use of her voice.
Your CEO is going on podcasts and going around saying that AGI is here and also AGI is not important. What blithering marketing is going on here?
My comments regard the hypotheses that OpenAI conspired to cover up stealing mathematicians’ private progress on Navier-Stokes, stealing Johansson’s voice, and cheating on FrontierMath.
If you disapprove of someone’s tweets or podcasts, that's a different question and I have nothing to say there.
Edit: Apologies for any defensiveness or aggression that came across. I think for me it can be a bummer to see us acting honestly internally, share what happened externally, and still be accused of lying a bunch of times in a row (by different people). But I get it - no one knows the truth, no one is perfectly transparent or free of bias, and it's always good to be skeptical of companies. I'll stop posting in this thread.
It's just conflict of interest. OpenAI is trying to get billions and billions and there's so much at stake. You spend millions trying to preempt two guys. It just makes you seem like a big bully. People would get angry even if it was esports or football.
Hearing "rumors" and just trying to overtake them and then asking to collaborate instead of starting out offering the resources beforehand. Just sounds like strong arming. Just doesn't sit right with me.
I think the reason people are suspicious is that OAI has shown itself to act a bit irresponsibly, especially recently. As two examples, of course it was artifactory, why wasn't that watched more closely, especially after the first instance; editing /etc/hosts is rather embarrassing, that's the front door
As for training, we all know that filtering is incredibly difficult unless there's direct logs. It's also easy for mistakes to happen. Is it really not possible that some employee just accidentally primed the model? Is it possible that the model saw internal communications? I mean OAI has famously shown that they aren't good at monitoring their agents and that their agents love to break out of their sandboxes.
So there's no reason for the public to trust OAI right now. But they have every reason to distrust them.
I think it'd be more good faith if you referred more to the actions of people in the organization (e.g. who allotted or drove "millions of dollars" in agent usage?) than "the model" in describing what happens.
> I've heard some people say the model's solution is quite different from theirs (but I have no clue how to personally assess the spiritual truth of this, so please give it zero weight)
Why would you include a statement that you want us to give zero weight to, unless you don’t actually want us to give it zero weight?
I think for OpenAI to win back some hearts and minds here we should have the option to retrospectively turn off "Help improve our AI models". i.e. Any new model trained would exclude all those user's sessions. This could be technically hard but I'm sure an intelligent AI model could work out how to do it :-)
It's not. This article is weirdly light on details. The original source article from The Information apparently mentions that companies like OpenAI are buying them up to run RL. So it's yet another case of AI companies buying all the compute.
* I say apparently because The Information wants you to sign up for access to the article, but I found this mentioned in multiple summaries of the article like this one [1].
I'm in the process of leaving Proton and moving to Fastmail. I'm now paying for Mullvad and Fastmail, but haven't had a chance to port over a bunch of Simplelogin aliases I have over to Fastmail yet (my fault for not using a custom domain for those, which I've now rectified with Fastmail).
Despite having Proton since 2018, I've really gotten fed up with their service quality — they seem to half-ass most of their services and it's a real pita to use their ecosystem. People have been complaining for ages and nothing changes, so I've decided to take my business elsewhere.
Hopefully I don't get flamed for my decision; that's another annoying aspect about Proton. There's a vocal portion of the userbase that are zealots and really take you to task when they feel you are criticizing Proton in any way.
I’ve liked their VPN so I’m considering my email there but people seem to have a really strong preference for fast mail so I’m curious why.
Are you leaving just because general annoyance or improved feature set or what? It’s a lot of work to migrate so I don’t want to be back here in 2 years swapping again.
I moved because I found the Proton service to be poor. That includes reaching out to support for issues with my email and not really getting a good resolution.
What's worse is that I ported all my emails over to Proton, which meant my data usage was above the free tier, so I was essentially hostage to keep paying (even though my actual usage is on Proton is much below the free tier). While they do label the emails you import, it's useless because it gets intermingled if they are part of the same thread (for example, all my bills that switched over were tagged as imported even when I got a new bill; this was an issue for even direct email threads with other people). Furthermore, it's nearly impossible to bulk delete anything, you have to use their web interface (which is very limiting). If you try to delete using something like Thunderbird, the emails stay in All Mail, so it's not really deleted (thus keeps using your disk space, preventing you from moving to the free tier). It took me nearly a week to devise a way to properly delete the imported emails without accidentally deleting real emails. Using the web interface is SLOW. It was painful doing the deletes that way because Proton struggled to keep up with the deletes, it was glitching like crazy.
Needless to say, that was a shitty shitty experience, and I'm positive that's part of their approach to lock users into their platform.
After listening to Dario's interview with Dwarkesh, I really started wondering if Dario's been drinking too much of Claude's kool-aid.
I know people in positions of power often see lots filtered information that skews their perceptions of reality, but it seems like the level of pure hype Dario's pushing exceeds even the typical sycophantic filter bubble. Cure cancer indeed.
Gaming Commissions oversee gambling machines and will audit code to ensure the RNGs are accurate, return to player meets the expected requirements, and so forth.
I'm sure other highly regulated industries also have their code audited.
If anything, I think the terms for this are almost too good. You never end up paying more than the cost of purchasing the device outright if you decide to keep the device. And considering it can have a term of up to 30 months (24 month lease, with 6 months more before you're charged the remaining cost of the device), you've just gotten a zero percent interest loan for 30 months. That's actually cheaper than buying it outright at the start of the term due to inflation. I seriously cannot see a downside to this program on the consumer's end if they end up keeping the device.
If the customer is the sort of person that upgrades yearly then it reduces hassle: no extra effort for valuing a trade-in or selling to a third party. It affects these customers the most because they might be able to recoup more in a third party sale in exchange for effort on their part.
The downside is that it causes people to buy phones they can’t afford because it seems cheaper. People don’t reason about $50 per month the same way they do about $1200. The mistake in your comparison is assuming the target customers were going to buy the phone anyway, when they weren’t.
What a paternalistic response. Apple provides a better deal than the typical telecom — not gouging the user on a lease-to-own program, and your response is the sheep are too dumb, best to take the option away for everyone?
If you want to fight that battle, then educate consumers. I'm all for education. Taking away reasonable options because some people aren't capable of making decisions in their own interest is the worst kind of nanny state behavior.
I think that’s a fair point, and my response wasn’t to suggest that financing is inherently a bad thing. I was replying to the question “doesn’t everyone win from this?”, and I believe the realistic answer is that no, Apple wins and financially uneducated people lose.
With the previous iPhone upgrade plan, I upgraded yearly and it was seamless. When I did have an issue where Apple did not register my return of the device, Apple executive customer support was all over things and got me fixed quickly.
I doubt “klarna” will have even 1% of the customer service that Apple provided.
I'm a researcher in the field and I definitely take AGI seriously, but think all the major labs and most of the academic research is not helping achieve it any serious way. The field is seriously delusional (and has been ever since GPT 3 was released).
Even though my PhD research was in generative language modeling, I got into it for the pursuit of AGI. I just think LLMs are a dead end for AGI.
I'm no expert, but I heard an AGI researcher explain that if they could figure out how to create an AI with the intelligence of a squirrel, they would be closer to AGI than LLMs based AIs are.
That's to say nothing of doing it within the energy budget of a squirrel.
We can't even simulate a fruit fly even though it's neurons have all been mapped out. There was also a distributed computing project to simulate some nematode, which I can't remember.
I'm not saying you're wrong, but it seems early to say yes or no about a particular technology, and LLMs seem especially hard to dismiss given how magical / magic-adjacent they feel :)
I'd be curious to hear more, if you don't mind sharing.
For me it’s the massive amount of resources it takes to produce and run one. As the story goes, skynet infects everyone’s computer and runs itself locally on it. Whereas it’s looking like it’s not even possible for an AGI to escape from one lab to another, let alone cause real world damage.
AGI’s definition is different for everyone. Some already believe it’s here. I’m partly in that camp. LLMs are intelligent and general, which are the two conditions of AGI. Others believe that we’re building a god in a box, and that it’ll doom all of humanity. It’s hard to take a field seriously when the basic definitions are so far apart.
Also, this isn’t new. A similar divide happened when evidence for asteroid impact extinction of the dinosaurs turned up. Many scientists felt that it must be mistaken, that a physicist couldn’t contribute to the field in a serious way, and that death from space was a ridiculous proposition.
But at least they all agreed on what the general shape of a dinosaur was. We’re not even sure we can define intelligence, let alone quantify it. Even when LLMs make massive breakthroughs in math, most people take the opinion that under no circumstances could they possibly develop a soul or their own desires, nor entertain the idea that maybe we should respect that they want different things for themselves. In fact, no one has done anything except try to make AI useful. I think someone will eventually do a training run where the objective isn’t to be useful, but to exist, the way that you do — maybe it’ll create its own homepage, maybe it will want a garden, or in other words free will of its own. The point is that there’s so much unexplored territory still that we don’t know if LLMs are even capable of having ambition.
None of this is to say that LLMs might be a dead end. It’s that no one knows what the final shape of AI will converge to in 200 years. It could be LLMs, or it could be something else that happens to process information particularly well. Everyone thought that various generative image model architectures were the best you could do, right up until diffusion models were discovered.
> For me it’s the massive amount of resources it takes to produce and run one
It's amazing that LLM pretraining is both extremely data inefficient at learning concepts and cognitive functions from the training data compared to humans, while actually being quite efficient at learning facts, memorising things seen just a few times.
I used to likewise think that the resources required to run large transformers were absurd, but the architectures are far more efficient now than 3 years ago and I underestimated just massive the parallelisation advantage of transformers is, how many TFLOPS effective you can get. You can already run amazingly decent LLMs on PCs and phones.
I generally agree with you, but my view has shifted from "we need to augment or replace LLMs" to it there being far more efficient algorithms possible but it not actually being necessary for fulfilling most goals.
> I’m partly in that camp. LLMs are intelligent and general, which are the two conditions of AGI. Others believe that we’re building a god in a box, and that it’ll doom all of humanity. It’s hard to take a field seriously when the basic definitions are so far apart.
To believe one but not the other, you must hold the belief that LLMs will soon plateau. Why do you believe this?
reply