Some of it is. He did the classic “book an entire vacation for me“ except he only asked it to plan.
No one is ever going to just tell a phone to book a vacation for them. It won’t know what you want well enough. Planning is much more reasonable because then you can just adjust whatever you don’t like.
But as a demonstration, it certainly shows how far Siri has come. It probably would’ve just offered to do a web search before.
The things I’ve found it really useful for so far is the “find this piece of information that must exist on my phone somewhere, but I don’t have the slightest clue where“ kind of query.
Would love to learn more about some techniques that "everybody" uses to do this well. So far, everything I've seen that meaningfully advances the frontier has been high-touch (involving human experts in one way or another).
It's fairly easy to describe a task that is slightly harder than an existing one.
For example if frontier models are able to one-shot a database query across 20 columns and 10 tables add one additional relationship then test. Keep doing this until the pass-rate drops below acceptable and now you have your new frontier eval.
My thought experiment was along the lines of "Let's say I'm Anthropic and I want to significantly improve my frontier model's performance on, say, theoretical physics research. How do I build a fully autonomous process capable of constructing an eval that's somewhat outside the current capability in some useful direction (decided by the autonomous process itself)?"
This hits. Ive had an annoying issue with Codex VS Code extension for weeks if not months when queuing up a follow up. Often it just disappears, I think it's just a UI glitch is actually queued but I can't click steer if the UI doesn't update and I find myself automatically going CMD+A, CMD+C before sending a follow up message.
The sandbox still allowed the agents to install additional dependencies (from PyPI etc) that they needed. It did this by locking down all network access with the exception of an HTTP proxy that only allowed read access to PyPI and a few other places.
Thank you, I'd assumed they'd restrict egress at layer 3/4 although I guess then it might just have found an exploit on the http server of an endpoint it was able to access.
Sounds like a bad habit for security testing Ai. It's not that hard to build an internal mirror and proxy that, keeping the real internet physically separated if needed, and truly locked down if concerns aren't as great.
It would be great if we could have AI that wasn't trying to emulate a human. When it expresses emotion, we should see that as a bug that needs to be fixed.
reply