Hacker Newsnew | past | comments | ask | show | jobs | submit | throwaway314155's commentslogin

That's clever and all - solid setup, good work. But I still think you either overestimate the average house guest or have particularly savvy/young house guests.

Most people think it's cool but use it once.

The regular-ceiling-lights-as-alarm service gets positive feedback. Weird how smarthome companies never market that, seems easy win.


Not sure I totally understand what you mean by light alarms. Do you just flash ceiling lights to notify the user of something?

No they gradually turn on from 0% to 60% to simulate sunrise. Accompanied by the shutters slowly opening, if you want.

Then after 20 min an actual alarm goes off, just in case.


No Man's Sky and Cyberpunk 2077 are actually proof you can have an awful launch and still pivot into a large player base and continued sales via updates.

Reception for both games is largely positive despite their poor initial response.


Cyberpunk 2077 is in a worse state, though. CDPR axed multiplayer from the roadmap after launch, and the updates it got (eg. working police cars, outfit transmog) should have been there at release.

NMS is no crown jewel, but Hello Games did add multiplayer and held pretty close to the original vision that was offered. It's much more "fixed" than current Cyberpunk 2077 is, at least IMO.


> that distinct off-white shading

Let's see Paul Allen's card.


Is it expected to be available to subscribers?


Yes.


> You have good enough hardware to run good models comparable with Gemini and ChatGPT.

That is at best misleading and at worst outright misinformation.


If you have 2TB of VRAM you can’t run one of the big models which are comparable?


The post was replying to someone with 16GB. (And also: no, even the best open weight models are not as good as what you can use on your ChatGPT subscription. They’ve gotten a lot better, but not that much better.)


[flagged]


[flagged]


How many of these sockpuppets are you going to create?


Your emphasis here seems to downplay the end-result of the attack - which is arbitrary code execution from a seemingly innocent URL merely being read by the LLM. The ACE is pulled off without the user knowing, and seemingly without agent or its auto-mode classifier knowing. There are at the very least _elements_ of prompt injection/jailbreaking in here. The LLM reads content and performs actions described failing to stop itself.


Well, the point is it's more than being read: the task involves downloading and interpreting information which may also involve code. I think it is a pretty good demonstration of what filtering at the LLM-interaction boundary can and can't protect against. And also a good demonstration of how agents will take more action than you might naively assume when given a task unless you specifically limit them.

The end result is the same but it's more of a mismatch of expectations from the user and plain old trickery than it is an attack managing to misdirect the goals of the agent, and it's worth being clear about where the issue is and isn't.


Queue the hoards of Python haters apparently.


*cue

*hordes


This website is the worst.


> They're all in Tokyo ;)

Yeah and effectively any store outside urban America.


I don't think GPT-4 was ever used for beating pokemon with or without a harness. Successful attempts include Gemini 2.5 and Opus 4.7 both using relatively advanced custom harnesses that give access to game memory, notes systems, and one-off hacks to get around parts of the game the model gets stuck on. More recently, Fable 5 beat FireRed with a _very_ minimal harness (screenshots and button inputs). That's the only example I know of but that is a very sophisticated and very expensive model compared to GPT-4.

Most of this doesn't discredit your overall point, though.


GPT-o3 was the first OpenAI one to beat Red. The harness used by GPT Plays Pokemon is the most featureful one of the main competitors (GPT, Claude, Gemini), IIRC.

Community maintained spreadsheet of the runs: https://docs.google.com/spreadsheets/d/e/2PACX-1vQDvsy5Dt_-P...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: