The implicit arrangement in corporate use and contribution to open source is that everyone benefits. People contribute work and testing that improves libraries and tools that everyone else can use, making all our projects better. This includes corporations, many of which also engage in this mutual benefit by hiring actual developers to back open source projects. I think it’s a brilliant system, huge fan.
Enter LLMs training off that data. It looks similar on the surface, but if you think about the balance of cost and benefit it becomes clear it is no longer balanced. The training algorithms and models are closed source, and come at a hefty usage cost. AI companies saved hundreds of billions of dollars of paying humans to help train their algorithms by taking advantage of an open system that was open on with another kind of balance in mind.
So in short the work was offered for free within a specific kind of relationship. Now a powerful outsider is steamrolling and taking advantage of that existing open dynamic.
Most FOSS is offered free of charge, but not free of conditions. You still have legal requirements spelled out in the licenses, from attribution to actually requiring you to license your work accordingly.
I feel like the standard "Free as in Gratis vs. Free as in Freedom" argument doesn't quite explain the issue people have here.
Open-source code still had terms, and AI training was an unfortunate blind spot in those terms. If a human being learned how to code 100% from open-source code, nobody saw a problem if that person went on to write closed-source code.
But it feels fundamentally different when the open-source code is just being fed into a machine that uses absurd amounts of processing until closed-source code comes out the other side. It feels like theft.
What if it can be Einstein, but can't draw a Pelican, write a solid college-level essay, or fold clothes?
The ability to do a ton of book learning in training, and pull in tons of related context at once, is superhuman in some ways, but lags a lot in others.
> What if it can be Einstein, but can’t draw a Pelican, write a solid college-level essay, or fold clothes?
Then it’s an expert system.
Stephen Hawking wasn’t very good at folding clothes.
The ‘General’ part of the term ‘AGI’ seems like a trap to me, because there will always be new workflows to master. Can Astra one-shot level completion on some yet-to-be-released video game? If no, does that mean it’s not yet ‘Generally’ intelligent?
You won’t get pure ‘general’ intelligence until you find Einstein’s hidden variables and load the state of the entire universe into context.
Meanwhile, building a series of expert systems targeting specific valuable workflows is useful today and seems like it’ll continue to scale to cover huge swathes of economically valuable workflows.
I think that’s the more interesting thing to be measuring. The surface area of useful economic workflows that can be addressed with expert systems built with today’s tech.
Hitting some ‘Artificial Expert Intelligence’ coverage threshold on economically valuable workflows is what will matter for humans well before pure ‘general’ intelligence.
The only important part of 'general' is the ability to learn from experiential data and update your own model. That's what leads to general capability. Humans can't oneshot any task natively, but we can practice for a while until we uncover often novel methods of accomplishing something.
Therefore: the current transformer architecture is fundamentally incapable of AGI because the models have no mutable long-term memory.
You only have weights (large immutable memory), or context (small mutable memory).
Humans have mutable long-term memory: I can learn a new skill, adapt an old skill to new information, or learn new knowledge today that I couldn't perform/didn't know yesterday. I don't have a training cutoff.
Context engineering is an attempt to paper over this limitation. You can get really far with context engineering and huge models, but you will never get to AGI because there are many tasks where humans' mutable long-term memory outperforms.
For example, a human can invent a new musical instrument and then learn how to play the instrument they just invented. That's inference (inventing an instrument) leading to training (neuroplasticity). Humans have the ability to train our NNs with considerably fewer training samples. Everything that you can do with transformers is in one causal direction: training -> inference.
So if we take a huge with enough compute (CPUs, b200s, petabytes of SSDs), we install on it both the Astra, and the toolsuite to incorporate new sensory inputs (threads/sessions), camera, microphone, temp sensors, the lot, into a new version of the model. This model is then swapped for the old model, or traffic slowly brought over, or even adjusting weights in place.
Then my hypothesis is that thing as a whole could achieve AGI.
This feels like a very close approximation on how we humans evolve our brain. By encountering new experiences/sensations, classifying them as negative or positive to us, filling it away in neurons. Or by training motor skills etc. In the end we get more connections between neurons in our brain and we are capable of more.
Bingo, LLM architecture just does not lend itself to becoming AGI. They can get really good, sure, but they will always struggle with novel input and scenarios.
The more training data that is shoved in to them, the more they'll seem to solve novel situations, but in reality it'll be things that exist in the training data.
Aka Star Trek hologram characters aren't sentient, and actually anyone who things droids in Star Wars can think of a weirdo. C3-PO just kept running out of context and trying to revert to it's system prompt.
> Can Astra one-shot level completion on some yet-to-be-released video game? If no, does that mean it’s not yet ‘Generally’ intelligent?
If a model can't learn on their own to play some new game just as well as humans do, it's not AGI.
It's okay if they would take some hours or days of learning (like humans might), but if they can't do it at all during their normal operation, that's not general intelligence
> You won’t get pure ‘general’ intelligence until you find Einstein’s hidden variables and load the state of the entire universe into context.
But humans have general intelligence. AGI is about matching human ability, and we know this is possible in principle because brains exist
And the only reason LLMs can't write essays indistinguishable from human output is because they aren't RLHF'ed to write like humans.
Folding clothes isn't an LLM's job but if you were to insist, they could certainly do it, as any number of videos from robotics labs will attest. That particular future is already here but definitely not evenly-distributed.
Adding sibling comments, I think some people may be overestimating how well the median human can draw a pelican, or create an SVG of a pelican (depending if we’re comparing to an image generation model, or SVG generation).
Most people can't draw a bicycle. There was an artist 10 years ago that asked people to sketch a bike, and then turned these sketches into 3D renders - quite funny.
I can't draw a pelican. Literally my only point of reference would be AI pelican drawings from the test. Otherwise I wouldn't know how to draw one at all.
I would be able to draw an accurate bicycle, but I'm an outlier on that. Most people could not draw one [1].
Can definitely write college level essays and have for a while. The jobs is that when LLMs first started getting popular, but aren’t quite common professors were that some of the worst students in class started writing the best essays. Now everyone complains because they can detect the slop, but most human writing is so bad. But the really good human writing is still much better.
I would maybe argue that Einstein was the most LLM-like of great thinkers.
A lot of his great discoveries were mostly that he was very knowledgeable about the bleeding edge research in a number of disparate areas, and was able to have the aha moment where he could make the connections for how to integrate them.
A lot of other thinkers who created new fields from scratch are probably way harder for an LLM to crack.
That is very aligned with an LLMs ability to have superhuman knowledge in wide areas.
That’s like that scene in the I, Robot movie when Will Smith’s character is asking the robot “can you turn an emtpy canvas into a work of art, or compose a symphony?” and the robot replies “can you?”.
Not the op, but this TS migration started long before AI was able to help. It was done slowly and carefully, as a project supporting millions of users should. And the benefits are very clear.
Bun’s port was a vibe coding fever dream that happened from one day to the next, with much looser motive, and yet to be proven reliable.
Bun's migration to Rust was nothing more than a marketing stunt to sell more Claude subs under the impression it can perform this kind of work at scale, assuming that most who were convinced by it wouldn't look under the hood at what really took place.
It has its merits as a proof of concept that could eventually be cleaned up and released properly later, but I can't see it any other way.
Too many see it as this miraculous one-shot and are using it as a blueprint to justify more layoffs and buzzword salad in their boisterous LinkedIn announcements about how they're "completely overhauling their strategy" in engineering. Hogwash.
The irony is that the blog post actually points out as pain points the reasons many of us assert languages like Zig are out of place in the 21st century.
A nice collection of heap-use-after-free crash, use-after-free crash, crash and out-of-bounds read, memory leak, double-free crash, race condition crash.
Bun is infrastructure. Why would I want my infrastructure to be unstable? (By the way, 10,000 unsafe blocks last I checked, though the number is going down somewhat.)
I've never used bun on production for this very reason. But nevertheless, tens of thousands of people and businesses do.
And I'm not sure how you're responding to my comment. The parent said "this is a marketing stunt" derogatorily, as if it's slop that doesn't work. This is already the canary build, it's more stable than the current stable, and is actively in production products in wide use.
The parent is objectively wrong, whether or not I personally use Bun.
Marketing could of course be one of the main motivations, but it's not a "stunt". More like a marketing achievement, I guess? Stunt implies smoke and mirrors, and bun rewrite is quite real.
Probably cheaper than doing it by hand, however, that's the short term "port x to y" cost, the longer term cost (or benefit) is a lot harder to calculate.
I don't think irresponsible is the right word, but it has drastically reduced Bun's appeal. All the tools we use have a brand to them, and Bun basically changed their brand overnight to "reckless" in my eyes.
bun has never been fit for production, at least not for load bearing business apps. it’s been haunted by segfault bug reports since the early days, and i personally hit at least one a week when im doing lots of bun stuff. im excited for bun with less segfaults
It's unknowable, because the PR is unreviewable. The Bun migration PR is larger than any model ever made can fit into context. You just have to pray that test coverage is sufficient to catch all of the possible errors, which it almost certainly isn't.
The initial state has no bearing on whether the process was responsible or not. That's measuring along a different axis. If the bun rewrite lands and it breaks someone's app, that's bad no matter whether there's more or fewer bugs in the final state. The important metric in a rewrite of software that's used in production is stability.
It's really not even close to being the same. In the best case, a bug means your app crashes on the new version. In the worst case, something more insidious happens like opening a security vulnerability (say, TLS isn't handled correctly or HTTP headers are mishandled in a way that allows SSRF or request smuggling) or a previously linear time operation is accidentally quadratic (leading to DoS).
You can apply the same FUD to the old version. Your argument basically is “all change brings risk” which is true but doesn’t add any useful insight. It’s always easy to complain and warn about change causing problems while ignoring the problems of the status quo. The “everything is fine” meme in action.
Sure, all change brings risk. But this is change where:
1. The change isn't made by a human.
2. The change wasn't fully reviewed by humans or machines. It's not currently possible for a machine to review the whole thing as one.
3. It's a full rewrite. This isn't a ten thousand line change, it's multiple orders of magnitude more than that.
You're literally making the argument that all risk of any size from any change of any size is equivalent, so just don't worry about it. If you relied on this software before, good luck convincing yourself it's fine to rely on this software now: it's literally not the same software anymore.
No I’m not making that argument, you’re the one making that claim as if it’s the only alternative to your position.
My position is more nuanced. A) what does the test coverage look like B) how is the deployment managed.
I suspect B is going to be my biggest issue - normally you’d deploy this slowly over time to monitor problems and whatnot. But ultimately the real test is seeing how it actually performs in the wild and kinds of problems people report. But you can always keep using the zig version if you wanted. So ultimately it’s a lot of consternation over a nothing burger. You can laugh at them if they screw up the release, but it’s a bold attempt at trying something legit. It took Microsoft 2 years of many engineer hours migrating typescript to Go. If it takes significantly less calendar time and human time, you could reasonably even evaluate what a Rust based typescript looks like vs Go if you wanted to for an order of magnitude cheaper.
I don't think they're wildly different purposes. They're the same purpose (to set shell settings) with different scopes (all users, one user, interactive shells only, etc.).
reply