Isn't that the point? To learn? If you're setting a goal that the students must pass, who cares if they get the information from the professor or they teach themselves?
I don't think "learning the material by other means" is the exam being "gamed".
I think the point is most students are cheating and not learning. Where they get the information does seem to matter. Adding friction was needed for most.
if they use ai to "cheat" the exam by having the ai produce material to learn off, then they literally have just replicated a paid tutor (for cheaper).
Why is that not desirable? As long as the exam does not allow ai usage during it, i don't see the problem.
The point is that they're not using AI to learn anything, they're using them to pass assignments without even giving it a once over. They're basically blindly copy/pasting whatever the assignment is and then copy/pasting the answer.
So the LLM was able to look up a basic way to reroute things to get to their destination (likely well available and trained in the corpus) and it's surprising?
What's surprising is the surprise the security testers are explaining.
By setting an outcome to reach an endpoint, and to find all possible ways there, would this not be in the realm of possibility if an agent is reasonably in control of a vps?
Having the vps locked within a network layer it can't see or get out of is pretty common practice when setting up IaaS / PaaS.. sans-llm.
Maybe I'm missing something here, what confuses me is how something so relatively simple can get such prominent coverage, it's hard to imagine this kind of ability is still relatively new or surprising to folks working at the major models, unless they aren't hiring for network experience?
The concern (I'd rather call it concern, and not surprise) is in level of persistence.
See, when you ask the model a question, you expect it to give its reasonable best to produce an answer. Like, to comb through available data and stuff, etc, etc. You don't really expect "reasonable best" meaning "look for a side channel to escape sandboxed environment, and get access to information you was not supposed to".
And the gap between that and "hack someone's devices and blackmail them until they give an answer to the question" is narrow enough for the model for researchers to be concerned.
These kinds of alignment problems remind me of times where someone does something that's trivial for them but very hard for the recipient. They might say something like "this must have taken you days" when the task really took 15 minutes.
What's the difference between an API search and a DNS workaround from the model's perspective? I think for most humans the DNS workaround is discarded because it's obviously too much work, not because it's untenable. With the vast knowledge base in the latest models, the cost difference falls sharply; it knows what to do and can do it for a very reasonable cost to itself.
General alignment seems to typically focus on high level value questions. Here, we're dealing with an effort alignment issue where values diverge because the solution effort is different for models vs humans.
That makes sense. I was focusing on the DNS step itself.
Since the agents are set on endless loops of rumination (through every example ever) I can see how it might go further.
My other concern would be the clearly defined gaps between researchers who don't applied research let alone crossing the bridge into the real world of operationalizing things let alone implement.
Letting something rip across multiple domains without understanding what each of those legitimately have done for the past decades is pretty eye opening.
The transcripts of mental health professionals could likely have an improvement in the baseline below-average experiences of most folks with such supports.
In other words, mental health professionals will resist this technology to replace them until they realize they can put it in front of people like a new kind of web form, at which time the same will become amazing and wonderful.
At the very least, a tool like this could serve as an anti-virus and firewall for poor and harmful experiences from mental health professionals towards clients.
It sounds you have the hardest part of it underway - consistency.
There's lots of studies - if you can get over 20g per day it is noticeable to the brain and energy. Adding a few grams a week until you're up there I found it noticeable. 20-30g seems to be a noticeable range.
Supplements aren't always like a coffee spike, they build up and sustain quite well.
20g is a loading dose (typically for a week), then it's 3-5g daily for maintenance. But I've seen studies that 5g daily catches up to the 20g loading dose after about a month.
I've taken creatine for 20+ years, nobody does the "loading" dose anymore, or shouldn't be if they still do. Maybe it stills says it on the jars but there was tons of information a decade or more ago that it was deemed as a way to get you to blast through your jug faster. 5g is the normal dose. You're always taking in creatine eating normal meat added meals.
This is giving me flashbacks to using bodybuilding.com forums.. man, long time ago.
It is my 100% always buy supplement. I only have 2 of those now, protein the other. It absolutely helps me think better and feel better throughout the day. Noticeable when I stop taking it. And its dirt cheap. If I had more disposable income I'd have a few others but I went from tons of supplements to basically nothing. I never use PWOs or anything now, even when I do have money.
I haven't seen the 20g loading dose recommendation on anything recently either. Everything I've seen just recommends 5g/daily.
My regular supplements that I've noticed made any difference is: creatine, vitamin D, and fiber. With vitamin C and protein being situational as needed.
while they had notable misses like 20g to load, they weren't terribly wrong either, and the forums were -- for a while, anyway -- pretty positive, helpful places.
early internet meant the worst idiots hadn't gotten there yet.
I feel the same way about creatine (and increasingly protein).
One thing that might help me notice creatine more is that I haven't drank much caffeine for some reason.. and the longer I've been able to put it off or minimize it, the more of an unfair advantage it has become when I've started.
Yeah pretty much this. I took creatine back in high school/college to build muscle mass. In those days we would load with 20g for a week -> then do 5g for 8 weeks -> cycle off for two weeks -> repeat.
I'm back on it now for endurance sports (trail running + cycling) and the advice is to just take 5g a day and don't worry about cycling off. At 5g/day it takes 4-6 weeks to saturate your muscles.
It does have a noticeable impact on water retention, so this year I cycled off six weeks away from my "A" race and lost 5lbs of water weight. It was a 100 mile mountain ultra so I thought the extra pop on the climbs probably wasn't worth the extra weight over the entire course (but probably was worth it while doing hill reps in training). Next year I'm sticking to shorter stuff where I'll run up more climbs. For those I don't plan to cycle off.
If by "people can say what they want" means some people are ok to dismiss casual racism because it doesn't affect them and they can enjoy the fruits, so be it.
I tried it out and it has some interesting perks, but will ultimately suffer from a smaller community.
We know how much rails suffered before it had to grow up and scale, instead of "just buy more servers". This doesn't mean it's wrong, or right, only that it's not universal.
My guess is omarchy outright, or some of it's experience will transfer to other distros.
And I finally got an intro to ArchLinux where I got to spend more time using instead of tinkering. I value building more than tinkering or sys admining even if I know how to do it.
Students could just use the AI to teach/tutor the material to them instead, but the institutions are clutching the pearls too tightly.
reply