Hacker Newsnew | past | comments | ask | show | jobs | submit | mrandish's commentslogin

> my hunch is that it constricts the actual thinking of the LLM

I've found that prompting any constraint on output (length, style, vocab, even simple formatting) not only places additional cognitive load on the model, which burns some of whatever cognitive budget is available, it will also often skew the output in other subtle and completely unrelated ways.

Since I found this artifact interesting, I did some pretty extensive experiments a couple months ago. The increased load is real, although it may not be apparent if you're not near any cognitive boundaries. The subtle skew, however, seems nearly ever-present regardless of load.


Can you give any further details or metrics on your tests?

> to create a perceived improvement

In addition to the dozens of opaque model parameters and hardware variables that can nerf or buff model intelligence, speed and profit, there's also the very real possibility that models aren't just training on benchmarks but could be evaluating if they are being benchmarked in real-time and applying more resources adaptively. 'Driver optimizations' that detected benchmarks in real-time were deployed in the first 'GPU Wars'.

> I have no idea is the actual frontier is stagnating.

Like a lot of complex, rapidly evolving tech, the truth is it's probably rapidly accelerating on some measures for a few and stagnating on many others for most - hence the divergence in user reports. It's depends on how you use it, for what problems, how rigorously you assess the output and whether you happen to be on a server bank, RAM pool or shard at this moment which hasn't yet been sufficiently 'cost optimized' by the margin algorithms. They don't call them load balancers anymore. They're Margin Balancers.


Because frontier models are completely opaque. Doing a controlled test of "the same model" months apart is simply impossible if you don't work for that provider (and even then, may not be feasible). We know from external observation that model performance changes minute to minute, day to day and week to week for a variety of reasons: load balancing, inference hardware, and shared RAM pool to dozens of internal software settings each of which impact cost, latency, time-to-first-token, quality, veracity, tool use, etc.

Those software settings are being changed in real-time by an algorithm and those algorithms are being tweaked and A/B tested daily by the ~~performance~~ revenue optimization teams. On the hardware side the footprint a particular model is running on is materially changing, growing or being re-distributed across DCs ~weekly.


> might they all pursue this kind of deception

They might but multiple competitors engaging in ongoing deception as an intentional corporate strategy isn't required to explain what we're seeing. It's entirely possible to get the same clearly unethical outcome without any employees knowingly participating in an explicitly unethical plan of record.

Instead it happens without overt coordination when individuals and groups within an org each pursue their local metrics and incentives. In isolation, no individual action seems obviously unethical on its own. They just look like 'optimizing performance', 'maintaining ASP or ARPU targets' or 'achieving operating margin', etc. Customers are still getting deceived and receiving less for their money than they think. The difference is most of the people involved in enabling it get to not feel bad about themselves.


> Why didn’t Sony become a major PC manufacturer after this?

I remember reading about it in magazines when it launched but at ~$1500 for the base model at launch (without monitor) and still ~$1000 18 months later, for consumers it just couldn't compete with the $600-$800 Commodore 64 and Atari 800. In 1983 >$1,000 was a non-starter for amatuer hobby users, plus there weren't many game titles for it.

While CP/M was popular with the early IMSAI / SOL / S-100 homebrew hackers, it was never very relevant for most mass consumer hobbyists. CP/M was really disk-centric and in the early 80s floppy drives weren't within reach for the '2nd Wave' broader consumer hobby market. We got our first computers (C64, Atari 800, TI 99/4a, Radio Shack Coco) envisioning cartridge & cassette usage plugged into a TV we already had through an RF modulator. If you could afford it, the component RGB output of the Sony would certainly have looked incredible but in the early 80s most consumers were still a few years away from even seeing their first RGB monitor in person. No local store had ever even had one on display, so we didn't really know what to make of it. To us, Composite Video was the esoteric "better looking" output vs RF to your old TV. :-)

The SMC-70's price and capabilities were a better fit for business users but by '83, MS-DOS and "PC-Compatible" were clearly 'the future' for business users and CP/M was on the way out.


Why not? We already need to include typos and misspellings to signal we're not an LLM.

What use is this when you can just prompt the LLM to insert those or use any other style of writing that looks human?

> This is unfortunately par for the course at many big companies.

Agreed. Even at companies known for having a great culture and being "different from the rest", the Game of Thrones politics is still there - it's just more hidden and polite. I was a senior exec in the biggest BU at just such an F100 valley tech giant for >10 yrs (one level from the CEO). The company actually did have a nicer environment than its peers, no 360 reviews, no 'up or out' mandates, and a vibe genuinely a cut above.

While the net effect was more pleasant day to day and teams were insulated a few more levels down than most places, the brawling over headcount, budget and revenue attribution was simply practiced far more artfully. Sadly, I think it's just endemic to the breed of "large public corp" and while the best may be better, none are immune.


Did a little research on BDXL because the need for multi-decade backup longevity, which I never worried much about, became more relevant now that I recently retired and have time to read some 80s 5.25, 3.5 floppies and 8mm/Hi8. Would be fun to see some of the first assembler I ever wrote.

I learned:

* Standard BDXL: ~$6/disc / M-Disc BDXL: ~$12 (longer lasting)

* Write/verify to 100GB BDXL discs at 4x is ~2 mins/GB.

* TFA says "128 GB BDXL never made it to market due to the 2016 bankruptcy" but a larger reason is that even though 128GB ( 4 vs 3 layer) was in the standard, it wasn't well-characterized. Turns out adding the fourth laser was harder/costlier than expected and read reliability was poorer. 100GB 3 layer is the sweet spot.


Naming a new product which is a "partner that helps you by making it easier" "Bob" seems like a... lack of historical awareness?


Hey now, sure, Microsoft Bob didn't get us ordinary users a "partner that helps you by making it easier", but it DID help MS Bob product owner Melinda French get a "partner that helps you by getting 10+ billion dollars" in the end, amiright?

<joking>

Maybe IBM is signalling it is likewise hoping for 10 billion dollars while also realizing the (AI) partner might not end up being that helpful either??

<further joking>


After what we've learned about Billy gates in the last year im not certain that's such a good deal.


It's a reference to Bob the builder, the children's TV show. You know, because AI IDEs help you build stuff


> computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant.

I don't think it's irrelevant but perhaps not in the way you're assuming. When assessing AGI I'm not evaluating counting characters or even the execution of math operators at any scale or speed. As you observe, computer software from Regex to spreadsheets and Mathematica already handle that well. But AGI isn't about what computers can do, it's about whether AIs can do the specific things which, until now, have been uniquely human capabilities. Like understanding nuanced context and then coming up with novel approaches to solve a new kind of problem not relying on any specific prior training or knowledge (the 'G' is for General).

Most definitions of AGI start from a baseline that already assumes easily passing a Turing test and doing anything via text response that a high school graduate could. I ding LLMs not for failing to count but for failing to intuitively understand the nuanced context of a simple class of problem it hasn't seen in its training data. I fully understand that the reason LLMs fail letter counting is that they operate at the token level. They weren't trained on individual letters first, like human 2nd graders.

The only reason recent LLMs get strawberry and blueberry correct now is that they have those words on their pre-training 'cheat sheet'. However, the underlying fundamental weakness in the way LLM intelligence works which leads to this failure mode still hasn't been addressed. Even when the frontier labs add "recognize any sub-token counting question and write a Python script" to the training cheat sheet so LLMs always pass that test... they'll still be unable to recognize a simple class of problem which isn't on their 'cheat sheet'. As long as that's the case, to me, they aren't AGI because they can't fully replicate human-like recognition of novel problem classes. And it's not just about letter-counting. That gap and others like it lead to many other kinds of non-human brittleness in LLM problem solving. Those are the classes of reasoning, intuition and insight that the ARC-AGI series has been trying to queue up as targets. Not to show how bad LLMs are but to help them be great in all these counter-intuitive edge cases


appreciate your response, but it's still birds vs planes.

AI does not need to feel emotions or have a heartbeat to be useful. It only needs to perform a task correctly à la Chinese room.

>therefore cannot fully replicate human-like intelligence

this does not follow. planes don't flap wings therefore they cannot fly?


> planes don't flap wings therefore they cannot fly?

This example still misses my point, which isn't related to usefulness or economic value. I concede that LLMs can have greater utility and economic value than humans on many tasks. The point is most definitions of AGI include something like "can fully replicate all the routine daily tasks done by any competent high-school graduate." That's not related to whether LLMs can solve many high-value problems faster and at larger scale than any human. That was also true of ENIAC in 1946.

The fact an airplane can fly faster and farther than any bird is irrelevant to whether an airplane can "fully replicate all the routine daily tasks done by any competent bird." That's the bird equivalent to most AGI definitions. An airplane can't build a nest or recognize the signals encoded in birdsong.

In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any". And in this context, airplanes scoring 15,000% more than birds on 'speed' and 'distance' doesn't matter any more than AIs scoring 15,000% more than humans on 'add 10,000 numbers'. We still aren't near AGI because LLMs cannot fully match any high-schooler's ability to independently conceive new approaches to novel problems not in their prior training data.


> isn't related to usefulness or economic value

which gets us closer to philosophical questions which I'm personally not that interested in.

>In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any".

I'm not sure we want a machine that fully succeeds that test.

Planes pass the 'bird replacement' test on the only criteria that matters to us ... flying.

If we wanted nest making planes I think we'd have them by now. Nest making doesn't rate highly on the problems we're looking to solve though.

I don't want a machine that is moody, or depressed or has schizophrenia, which are all pat of the human condition.

We don't need the human "intuition magic dust" to do 99.99999% of useful work.

They're machines designed to do the work we don't want to. That's as "general" as their intelligence needs to be.

I'd prefer if my clothes folding machine did not have an existential crisis.


Models often write python scripts for counting such things…


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: