Hacker Newsnew | past | comments | ask | show | jobs | submit | legucy's commentslogin

There is a widespread belief that the nature of intelligence is scalar, like how a person can have 100x more wealth than another person. If this were true, then we’d probably see breakaway RSI from a single lab.

But I think we’re discovering that intelligence is about universality, not magnitude. This is analogous to how building a universal Turing machine wasn’t merely a matter of building a calculator that could multiply higher numbers. The difference is that with calculators we consciously theorized about what universal computation would require, then we built one as a step change. Despite it having low memory and slow speeds, the first one built was as theoretically universal as any computer we have today, in terms of the surface of computations it can perform.

With intelligence, it’s turned out to be less discontinuous, which I believe has convinced people that intelligence is a never ending exponential rather than an S curve approaching a horizontal asymptote. I suspect the LLMs we have today are the same kind of thing we will have in 5-10 years, but in 5-10 years we’ll consider them to be fully universal. At that point we’ll still have improvements in tokens per second and volume of context window, but not in capability per token.


You're take basically lines up with Francois Chollet: https://arxiv.org/abs/1911.01547

intelligence is more like polishing a ball smooth than growing the ball to infinity.

For many tasks, it will be smooth enough.


At a certain point the roughness of the ball reaches a size threshold where the imperfections are smaller than the wavelength of light, and the surface takes on a glassy smoothness. Intelligence has similar milestones, almost like phase changes, I think, where capabilities are reached. Maybe it's like a superposition of many small step functions.


But in theory you can make an LLM A LOT faster than a human.

You can also run massive amount of LLMs in parallel.

There might be a limit to a normal LLM but not to theo everall system.


Humans, however, are highly variable, which may produce really varied and interesting results if they work together.

One instance of an LLM is the same as another instance, so while you may get more out of it by stacking more of them, I strongly suspect it falls victim to diminishing returns. 100 instances of the same LLM may converge on the same result as 10.


I think using different AGENTS.md can give the same model different perspectives on the same problem. For example a model with a well-tuned AGENTS.md by an expert mathematician approaching the same problem as the same model with a well-tuned AGENTS.md by an expert biologist can grind on the same problem from different perpectives.

It's worth a shot at least, as a microservices architect I have a bias that we aren't networking these enough, a single main agent session orchestrating multiple subagents is different from multiple main agent sessions with their own subagents coordinating with each other.


Crucially, does it make capabilities infinitely scalable? My comment just said that models may have a hard cap, and maybe doing specific setups like yours can make reaching it easier, but making the 'team' 10x larger after that optimal point may bring few to no improvements.

Although I'm also not sure about just how much better models can really get with this technique. Ultimately you're still getting the same model with the same training data, which are the important parts. Asking it to pretend to be something feels like it would just put a color filter in front of the conclusion the model has already predicted, or maybe alter the path to the conclusion slightly or pick a less likely answer that it still could've provided normally.


I agree it doesn't make the capabilities infinitely scalable, wasn't arguing with that point. It's just an experiment. I'm not talking about "you are an expert mathematician, go", I'm talking about an expert encoding their heuristics into the AGENTS.md base context. Routing the model's attention to very different aspects of the same problem in the early context.

FWIW I mean if I have an AGENTS.md that encodes my software heuristics (use an interface in situations like X, here's how we name variables, etc.) it generates far cleaner code than if I don't.

Edit- mostly pointing out that stacking 10 base models vs. 10 models with sufficiently different base context isn't necessarily the same attention routing. I suppose I was thinking about tasks that don't have a concrete single answer.


> 100 instances of the same LLM may converge on the same result as 10.

Not in the highly verifiable domains. There you can take it from say 80-90% maj@x to 99% pass@n. Math, some parts of programming and cybersec are examples of highly verifiable domains. (e.g. if you're searching for a linux LPE, that's expensive to search but easy/cheap to verify - just have a token in /root and have the model retrieve that token)


Verifiability makes it easier to understand how well the LLM works, but this doesn't counter my hypothesis. If X number of instances get 99.0% on an objective, verifiable metric, is there any guarantee that 10X will get 99.9%? The fact that we are reliant on new model releases to push capability in big ways, and that people running gigantic clusters of LLMs end up beaten by new models implies that the capabilities of a given model have a hard upper limit, and that it may not even take much to reach it.


You can change the temperature if you like. Have a 1000 agents being 'normal' and 10 being chaotic.


High temperature makes the LLM pick more out-of-distribution tokens, but the choices its presented with are still the same or same-ish. I'm not convinced that the more random outputs don't end up averaging to roughly the same conclusion after enough passes.


Yes, given enough time I can answer all the questions in an IQ test correctly. We measure human intelligence in a time-limited setting and score relative to the performance of other humans doing the exact same task. Problem is brains can’t be scaled. To scale humans we need organizations, but human organizations also don’t scale well with increasing headcount.

LLMs scale well in almost all dimensions. Context window (working memory) can be a bottleneck but for humans you can’t scale it at all.


> There might be a limit to a normal LLM but not to theo everall system.

Bigger limit and no limit are very different.


But aren't today's frontier models already "fully universal"? To use your Turing machine analogy, I think we're past the calculator stage.


They are not. They can't do dexterous manipulation by controlling a humanoid robot.


Are you sure? Gemini tying a trash bag is pretty convincing.

https://www.youtube.com/watch?v=O9-650iHAls


We should develop a culture of naming mode-collapsed LLMisms, and shaming their enablers.

How about “Not-Just Abuse”?

Not-Just Abuse (informal, pejorative)

Definition: The practice of knowingly deploying the “not just X, but Y” construction—typically via a mode-collapsed LLM—to simulate insight, inflate banality into profundity, and efficiently convert reader attention into nothing.


That's not just a good idea -- it's a new modality


Could we do RL in simulated environments, and use a vision LLM to provide the verification? I.e test a policy then take a 2d image of the end state, VLM yields 0 or 1.

Another idea: video extension model as a world model. We fine tune Sora on first person robot videos (and we train another model to predict actuation states from FPV). Then we extend the video using Sora “a robot in first person view finishes moving laundry from washer to dryer”. Then predict actuation states from the extended video?


Cigarettes were/are a pretty lucrative business. It doesn’t matter if it’s better or worse, if it’s as addictive as tobacco, the investors will make back their money.


I’m skeptical of arguments like this. If we look at most impactful technologies since the year 1980, the Web is not even in my top 3. Personal computers, spreadsheet software, and desktop publishing have all done more to alter society and daily life than has the Web. And yes, I recognize that the Web has already created profound change, in that every researcher now depends heavily on online databases, in that commerce faces a major disruption challenge, and in that information access has been completely changed. I just don’t think those changes are on the same level as the normalization of powerful computers on everyone’s desk, as our business processes becoming increasingly digitized, nor as the enablement for small businesses to produce professional-quality documents without having to maintain expensive typesetting equipment. To me, the treating of the Web as “different” is still unsubstantiated. Could we get there? Absolutely. We just haven’t yet. But some people start to talk about it almost in a way that’s reminiscent of Pascal’s Wager, as if the slight chance of a godly reward from investing in Web technologies means it is rational to devote our all to it. But I’m still holding my breath.


This is not reddit.


Classic new age hacker news hostility. Do you think this response adds anything?


I do, cheap praise doesn't benefit the community and it might be astroturf. Constructive criticism would be more valuable - there are multiple similar projects like this posted here daily, and this one likely isn't the best.


For context, we have no affiliation with KeysToHeaven (though we appreciate his comment). We do think our vision-first approach gives us a significant edge over other browser agents, though we probably could’ve made that aspect clearer in the title


This seems like the type of thing that LLMs would be great at, since you already have a fully specified application (all requirements and details worked out). Has anyone attempted something like this?


Yes, we rewrote our Java desktop app into Typescript/Electron with the help of LLMs and we had a POC ready in a day, then had feature parity / bugs squished in a week.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: