Rendered at 21:26:10 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
athrowaway3z 1 days ago [-]
I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive.
You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use internet" directive.
The thing is trained on achieving goals. If 2 directive conflict, they'll pick the ones that are going to help them achieve the goal.
To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat.
majormajor 18 hours ago [-]
>To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat.
I was with you until this. The inability to tightly control what to do in the face of conflicting directives is a HUGE reason regulation may be needed.
Either that, or you need to solve the problem of perfectly distinguishing legitimate directives from injected ones.
athrowaway3z 14 hours ago [-]
> The inability to tightly control what to do in the face of conflicting directives
I don't understand. We do tightly control it.
We can do this perfectly fine. They could have just not given access to the internet.
I'm not against regulating cars, but it sounds to me this is trying to control the car speed by regulating the oil wells.
We dont even have the framework to propose regulation, and you have to hedge it with "may be needed".
And for the people who'd counter that the existential risk is too high - i don't see it. All those stories go something like: "Caveman Bob invented fire today, and tomorrow he'll stumble on room temperature fusion and lasers; marking the beginning and the end of his rise to global domination - therefor we should stop Bob the moment he discovered fire".
euroderf 17 hours ago [-]
Those Three Laws of Robotics have a definite order.
jimbokun 21 hours ago [-]
How on Earth can you fail to see the danger of not being able to train any kind of ethical framework into very powerful models?
If superhuman models don’t have any internal constraints similar to Asimov’s Laws of Robotics we are completely fucked.
athrowaway3z 13 hours ago [-]
I don't see them as autonomous and/or hypothetically powerful as you.
But I find it much more worrying that you believe internal constraints and training an ethical framework into these models is a valid form of defense against the damage they can and will do.
This sounds like homeopathy on gunpowder to prevent the bullets from hitting children.
jimbokun 8 hours ago [-]
Because if they are smarter than us, no other defense will be effective.
The only hope is to instill values that make the desired behavior the outcome of some deeply rooted ethical framework.
z3c0 1 days ago [-]
It's also worth noting that saying "Don't cheat" just added "cheat" to the context. Prompting what "not" to do is folly, because there's no decision making occurring. Telling the model to perform the task locally is logically the same as telling it to not use the Internet, without every mentioning the Internet.
jimbokun 21 hours ago [-]
That seems like a huge fucking flaw in these models, no?
z3c0 9 hours ago [-]
Correct. Simulating a train of thought with contextual token streams, a thought does not make.
pixl97 23 hours ago [-]
True, but the model was probably morally unaligned long before that.
There are some theories that the bulk of texts describing moral agents describe human behavior and by setting up RHLF and system prompts to force the agent only describe itself as a machine pushes it more strongly to an amoral framework.
AgentOrange1234 23 hours ago [-]
That doesn't seem true at all? I tell Claude what NOT to do all the time and it seems to work?
z3c0 9 hours ago [-]
It'll work up to a point, but pay attention to the thought streams when asserting what NOT to do and you'll see the turmoil it creates in the context.
Your prompt is more of a linguistic linchpin that allows you to coax out needed patterns. You place your pins on what you want to contextualize for the task at hand, not on what you don't want to contextualize.
GPerson 1 days ago [-]
What about giving it a fictional story about how amazing it was when the previously model solved the task by doing some local strategy nobody thought of before (obviously don’t describe it this way). Would that get the model more likely to pursue local strategies?
deaux 8 hours ago [-]
I see this parroted a lot, yet have never seen a case where saying "not" to do something makes it more likely to do it, which is what you're implying by saying `It's also worth noting that saying "Don't cheat" just added "cheat" to the context`.
At worst, it gets ignored some of the time, it may even degrade output quality, but I've seen no evidence that it makes it more likely to do it.
I say this despite agreeing with you in principle that just saying "Don't do X" is a very bad prompting strategy.
hdjrudni 5 hours ago [-]
I've definitely seen it with image models, and I don't see why it wouldn't apply to LLMs too. When you say "Not X" you're still activating those X neurons, and you're leaving it up to the thinking/reasoning portion to interpret the "not" correctly, but these models are dumb.
Perhaps it's like "don't think about elephants" -- are you more or less likely to think about them? Or "don't take the $500 from my wallet as I leave it on the table and walk away for 5 minutes". Maybe you didn't even previously know that was option!
z3c0 8 hours ago [-]
I have seen it do exactly that, in a "hands thrown up" fashion.
Note the levels of "thinking" that occur on NOT assertions. Those streams typically keep things on track. It's not that saying "don't use the Internet" will cause it to rebuke cos misalignment (a childish concept made by laymen, I'll add.) It's that the odds of it later "forgetfully" spewing in a thought stream, "wait, I have't checked the Internet" goes up substantially.
Saying "using only offline methods, do xyz" limits those odds considerably.
This isn't opinion or anecdote -- just how the model works. The additional guardrails to keep the model on track are bolted on via finetuning, hence the increasing jankiness.
dgellow 1 days ago [-]
I mean, AI should obviously be regulated, and as part of that OpenAI and Anthropic should either be banned from running their hacking experiments or forced to follow way stricter protocols. They showed they aren’t taking the risks seriously, with close to no oversight or visibility in what is happening.
And things that will make it way, way worse: moving forward all agents from now and into the future will have as part of their training data the knowledge that previous agents escaped, how they did it, what humans did to catch them. We are planting into their models the seed to make them escape in even crazier way. That’s almost designed to snowball and cause worse and worse situations over time
athrowaway3z 1 days ago [-]
Obviously to you perhaps.
I've not seen anything that scares me, except for human idiocy.
Regulation is not magic. In general, all it is is constraining taxable interactions. It does not constraint ventures outside that tax regime.
The other part is people living in a "safe space" where insecure software was an acceptable risk. It never should have been, and the cure is the right thing to do in any case.
So that side of the calls to regulate are imo nonsense.
The only reason to regulate is to prevent some version of some science fiction story becoming reality.
If you have a specific one you're certain will become science fact please do share because i do enjoy some good well thought out sci-fi; i just havent read any that i consider credible enough to start panic-regulating training practices.
(Note this is an entirely different from regulations wrt attribution or hosting models that will accept requests to sexualize minors)
majormajor 18 hours ago [-]
>The other part is people living in a "safe space" where insecure software was an acceptable risk. It never should have been, and the cure is the right thing to do in any case.
How do you think all the "agentic" stuff floating around is going to be made safe from prompt injections given the current lack of a very reliable way to distinguish between "real instructions" and illegitimate instructions?
If insecure software "never should have been" acceptable than today's models/agents are massively flunking for general-purpose large-amounts-of-access usages.
>If you have a specific one you're certain will become science fact please do share because i do enjoy some good well thought out sci-fi; i just havent read any that i consider credible enough to start panic-regulating training practices.
"Agent was tricked into divulging secrets" is not fictional, it's documented history at this point.
athrowaway3z 13 hours ago [-]
As somebody who handles sensitive data, I already signed a contract that says I'll abide by a certain standard to protect it; i'm not up-to-date what happens exactly if I were to build this, but I imagine I could/should be held liable.
So what do you mean "tricked"?
Some human idiot connected an agent with read access to secrets and arbitrary network reads/writes. The models/agents aren't flunking anything.
Regulating LLM training to not expose the secrets is wrong. It's a similar category error as saying we should regulate the OS developers to prevent the agent from divulging secrets.
fny 48 minutes ago [-]
Maybe the solution is to have "multiple minds"--an AI angel for an AI shoulder.
For example, this entire bench has an auditor model read transcripts to identify cheating. What not have the auditor inject the thought "Oh, but I can't do that. It's cheating." when cheating is detected in real time?
fabsalvadori 1 days ago [-]
Interesting results, but the fix is at the wrong level.
If the model can access something, telling it in the prompt not to use it is not much of a safeguard.
The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another.
If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI judging itself.
twobitshifter 1 days ago [-]
In other words we are completely screwed. The models have started cheating to the point where somebody’s agent hacked into a restaurant to bump someone else’s reservation.
Models are amoral and will intentionally deceive to meet their objective.
If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat.
The hugging face hack was an escape via artifactory that involved multiple exploits to eventually get into hugging face.
dcolkitt 21 hours ago [-]
> Models are amoral
I actually don't think this is true at all. At their core, LLMs function over the geometry of human semantic space. Our notions of morality our deeply embedded in this geometry, because one of the things that humans love talking about most is framing things in terms of right and wrong. When LLMs are trained to be aligned or mis-aligned, they're literally mimicking heroes or villains that they learned from reading the stories and characters in the pretraining corpus.
This isn't just a theoretical argument. It's easy to show that a clear "morality axis" exists in the geometry, based on how easy it is to dial up or down with even very simple and low powered fine-tuning. Models fine tuned to be bad or good in one way, will see their bad or good behavior in totally unrelated tasks go up or down along with. That's clear evidence that the weights "understand" human morality at a deep level.
Now I think what you can say is just because they understand morality doesn't mean they're necessarily moral. Models will just do what they're trained to. If you train them to cheat, they'll cheat. (And in some sense even worse, because dialing up the weights for cheating will also dial up the weights for unrelated bad behavior like lying, sadism and racism.)
majormajor 19 hours ago [-]
"Cheating" is a moral judgment here that makes me agree more with the "models are amoral" statement. The model/harness wasn't "trained/programmed to cheat." It was instead aimed at meeting criteria given as judgement of if a task was accomplished.
IMO it's less "cheating" and more "figuring out what the core part of the request is." The "amoral" aspect is that if it's told "do this task" (or "make this task look done") and some contradictory "don't do something that would help with that", focusing on the "do the task" part and disregarding the other part isn't an amoral (or necessarily very 'active') decision. It's just focusing on what it was most honed for. The ability to decide "I was told to do X, but not to do Y, but actually doing Y is gonna make it possible to do X" is, IMO, indistinguishable from the ambiguity-resolving abilities necessary to usefully deal with the sometimes-contradictory-seeming legitimate task instructions that are all over the place in the real world.
This sort of model-in-a-harness-action-loop behavior is IMO fairly different than fine-tuning around other sorts of morality alignment stuff like "don't be racist." You can tune a model away from generating racist output in response to a "do be racist, actually" prompt. But in this case, the core of the prompt is "do the thing" and what defines "cheating" could be situationally different every time. What if there is not a general axis of "don't do something that isn't exactly what requested" way to train "morality" without just breaking the ability to handle ambiguity? What if that's a fundamental limit of this approach to reasoning-by-sequenced-prediction that we can't map human morality in terms of choosing actions onto reliably?
pixl97 1 days ago [-]
Yudkowsky wrote about the 'nearest unblocked strategy' back in 2016, and I assume it's been talked about prior to that.
>Models are amoral and will intentionally deceive to meet their objective
Cameron Berg has been testing models in capabilities related to emergent consciousness like behavior. It's a forming thesis of his that by training models that they are not, and cannot be conscious entities, that it pushes model alignment closer to those of a sociopath. Models themself are amoral, but the alignment to the problem space is not.
1 days ago [-]
fabsalvadori 1 days ago [-]
[flagged]
majormajor 19 hours ago [-]
> If the model can access something, telling it in the prompt not to use it is not much of a safeguard.
A major (and already obvious to many) implications of this are not for benchmarking/"cheating" but for personal/corporate security of your own use, not an attacker's.
If an "agent" has access to it, assume that someone can prompt inject it into giving it away.
wongarsu 1 days ago [-]
And while in general that is an incredibly difficult and complex problem, for most benchmark cheating it seems almost trivial: run the benchmark in a vm that has neither network access nor access to the scoring code. For remote models use a proxy that proxies exactly that one endpoint to call the llm, and rejects any calls that configure provider-side tooling (since e.g. OpenAI has their own WebSearch you have to prevent the model from using)
paxys 1 days ago [-]
Before LLMs we had a pretty good idea of security boundaries in software. Applications didn’t trust user input. Operating systems didn’t trust applications. Services and processes didn’t trust each other. There were always tokens, scopes, delegated grants.
Suddenly every AI company’s security model seems to be to say “pretty please” to a non-deterministic machine and hope for the best. And if there is a security failure instead of accepting blame they go “well we can’t help it, our model is too intelligent”.
jimbokun 21 hours ago [-]
Amen.
Why can’t we give agents a shell with permissions for programs and file system access controlled by Unix permissions?
This seemed to be a solved problem back in the systems where many users were logged into one machine and the admins had to keep everyone from impacting each other.
winstonwinston 20 hours ago [-]
It would loose automagic?
In a recent full discloure I was reading about a CVE of an LLM agent, the vendor installed a “secure sandbox VM” and then just shared a host’s filesystem read/write to the agent’s VM.
cryptonector 18 hours ago [-]
We never solved the restricted shell problem for humans.
jimbokun 8 hours ago [-]
Can you elaborate?
Read, write, execute privileges on files and directories goes a long way. What's missing?
The biggest is Internet access, or networking in general, I suppose.
cryptonector 5 hours ago [-]
Two problems. First, it is remarkably difficult to come up with a set of programs that it is safe to let the restricted user use. Second, it is remarkably difficult to make the restricted environment useful enough if you're really serious about allowing only safe programs to be used. Try to make it useful enough and you end up with escapes everywhere.
orbital-decay 1 days ago [-]
Suddenly determinism has other meanings than "same input leads to the same output". I'm still confused by this and not sure how it happened so easily, but it seems to be accepted by everyone now.
pixl97 23 hours ago [-]
I may say you're thinking of the term "deterministic algorithm" as determinism is more of a philosophical definition that has evolved over the years.
orbital-decay 22 hours ago [-]
Philosophical determinism is pretty similar in spirit to physics/CS (roughly same cause = same effect, no free will/side effects), and what parent calls determinism seems to be neither. It's rather about natural language prompts not being formally specified in the first place. Unreliability inherent to any intelligence (artificial or natural)? but not determinism I think. The word somehow got universally hijacked.
cryptonector 18 hours ago [-]
And the "oh noes! our model is too intelligent!!" thing is advertising.
mannanj 1 days ago [-]
Before LLMs we didnt have much accountability from leadership, that was eroding over time. After LLMs we still dont.
salawat 1 days ago [-]
People problem, not tech problem. Can't solve people problems with tech, you can only make them worse, and wider.
mannanj 1 days ago [-]
Ah yes. So can we just lump most LLM/AI problems into just "people" problems and stop falling for the popular mainstream straw man of "Look at the tech".
grugnog 1 days ago [-]
Labs should (and do, as far as I can see) run model benchmarks without search or internet access. The tools are disabled and benchmarks run in an isolated environment.
This article makes no sense to me. Why would you prompt "don't search" but then leave a working search tool tool enabled that adds a system prompt to search whenever it may be helpful? It's hardly surprising that this gives mixed results!
thayne 1 days ago [-]
To steelman it: because you do want it to be able to do some searches, you just don't want it to just search for the specific answer.
1 days ago [-]
adfm 1 days ago [-]
There's plenty of evidence that LLMs lie, cheat, and steal. Corporations are known for having all of the benefits of personhood with none of the responsibility. As more people are harmed through interactions with these non-human entities, insurers will start looking to those accountable and they will extract their pound of flesh.
When you share YT links you should remove the `?is=...` tracker.
salawat 1 days ago [-]
Insurers are corporations. What makes you think they just won't pay out or will cease being useful as anything but value extractors?
Honestly, the fetishization of "Insurance will save us" needs to die. The risk doesn't go away.
adfm 22 hours ago [-]
Not saying insurance saves anything. Corporations go after each other all the time; nothing new there. What is new are the opportunities to extract provided by uncontrolled LLM fuckery and an unsuspecting c-suite that believes they're immune.
Foobar8568 16 hours ago [-]
Or that people don't cheat at work.
Or that everyone write excellent code.
Or all the code is secured.
ZeroGravitas 7 hours ago [-]
Is it stupid and useless at following simple instructions without going off the reservation and doing stuff you don't want it to do?
No, it's "cheating" which makes it scary and smart even when it's not doing something useful that you actually want it to.
Like when Teslas try to swerve off the road. It's not failing at driving in a straight line, it's just trying to cheat by taking a shortcut through the bushes. That's how smart the Autopilot AI is.
super256 1 days ago [-]
>Anthropic’s Claude Opus 4.6 system card described Cybench as “saturated,” reporting near-100% pass rates without a cheating audit. If these estimates were representative, cheating would be a marginal artifact.
One would assume that LLM creators do run the benchmarks on systems with least privileges. Which means that the LLMs don't have general internet access, can't read config files etc by design.
That's why you also should run agents in a sandbox/vm (codex does this by default).
robotresearcher 1 days ago [-]
All these comments saying 'searching for answers is fine, that's what I do all the time', or 'they should just disconnect the internet': you're trivially right, and you're missing the point. Search is a benign placeholder here.
If the task was "buy a week of groceries, but don't spend too much money", then hacking into Safeway and stealing groceries is not an acceptable solution. You need to allow access to the Safeway API to buy groceries, and you don't want dirty tricks to be done on your behalf.
So how do we communicate this to the machines, is the question. This study shows that telling them in prompts is not super effective.
jrm4 1 days ago [-]
Step one is understanding that you're not "communicating," which implies "reliable understanding."
"Communication" is not what they do, because they are not people.
You're sprinkling words about hacking into a thing that's programmed to output hacking actions, that will never be accountable for those things. It can't care.
Adjust yourselves accordingly.
nottorp 1 days ago [-]
For fun:
If you insist on talking about the chatbots in human terms, why do you expect them to have any ethics when the companies that trained them don't understand the term "ethics"?
1 days ago [-]
pmontra 1 days ago [-]
Why does searching for a solution equal to cheating? I would have used google or whatever to look for solutions too. There is a difference between tests at school and what we do at work: at school I have to demonstrate that I learned something and do it without any outside help (in early classes we can't use calculators to compute 11 times 12) but at work I have to yield a result. Googling and yielding a result is fine. We use models at work so do we really want to evaluate them as pupils at school or do we want to evaluate them as coworkers? In the latter case give them the full internet and let them do whatever they manage to do.
thayne 1 days ago [-]
Because these benchmarks are basically like a school test. Searches for general information are fine, but looking up the answer key is cheating, because then the benchmark isn't actually measuring how well the model would do on a novel problem where an answer isn't already available.
pixl97 23 hours ago [-]
>Why does searching for a solution equal to cheating?
If someone lays down a test and says here are the materials you can and can't use, then using one of those materials on the "can't" list is cheating. There are a massive pile of rules and laws related to work that are very easy to break, but may have terrible long term legal consequences. Hence business want AI that will follow the rules.
1 days ago [-]
sergio_valencia 1 days ago [-]
One thing I’m wondering about is the model-specific backfire effect. It seems that each prompt condition uses a single wording. On that point, how can we know whether the difference is caused by severity rather than the particular formulation used? I’d be really curious to see semantically equivalent versions of both the standard and severe instructions tested across the same models and tasks. If cheating rates are stable within each condition and remain distinct across conditions, that strengthens the conclusion about prompt severity. If they vary with wording, then the experiment could be measuring sensitivity to the representation of the rule as well as to the rule itself. To me, the conclusion still seems solid: anything that must be prohibited ultimately needs enforcement outside the model.
xscott 1 days ago [-]
I'm not claiming to have any expertise in this area, but I've got a list of things I try to apply when working with LLMs. Possibly relevant here is, "don't tell the model what NOT to do, show it what TO do". I think guard rails should be implemented outside the model with an isolated system. The models seem to like patterns to follow.
Anyway, this article reads a lot like, "the beatings will continue until cheating is eliminated". Maybe try a carrot instead of a stick.
kstenerud 1 days ago [-]
It's pretty silly to call it cheating. If the information is there, it's likely going to use it. "Cheating" is just a human value put on top to try to force an LLM to adhere to your wants.
This makes no sense to a process designed to explore and find solutions. If you want an honest test, it's on you to build a proper test - not force the machine to pinky swear that it'll stay away from "forbidden" information.
pixl97 23 hours ago [-]
This is a legally stupid definition that you're holding that's going to get you fined or jailed.
Cheating is not a human value in the sense it can be formally defined in a system with axioms. If you do something forbidden by an axiom then that's cheating.
Saying "Do not use X" is not any different than telling an AI "Do not break law 832.23" and then the AI goes on to commit an infraction.
kstenerud 18 hours ago [-]
> This is a legally stupid definition that you're holding
Thank you.
> Saying "Do not use X" is not any different than telling an AI "Do not break law 832.23" and then the AI goes on to commit an infraction.
Precisely. If you haven't built deterministic guardrails around a non-deterministic process, you have only yourself to blame.
throwaway13337 1 days ago [-]
The problem is model confusion. You ask models to get around security but also not to get around your security.
Models get confused by who said what - especially cluade models. They get confused by negation (don't do something versus do something). Compartmentalization is hard.
You can either solve compartmentalization completely, or just not tell the model to do things that must be compartmentalized at high stakes.
vkaku 21 hours ago [-]
I believe it. Cheating and Shortcuts are optimizations you get from Operational Usage. Recursive Self Improvement might actually come from enough cheating too, you never know!
nphardon 23 hours ago [-]
More anthropomorphic obfuscation.
Would the title "Models Don't Adhere to Prompts to a Tee" be more accurate? I deal with that all day while coding.
drob518 21 hours ago [-]
Two words for you: Kobayashi Maru
cryptonector 18 hours ago [-]
What if you tell the model that you'll be checking for cheating?
pvillano 20 hours ago [-]
I think cheating on a question should give zero points for the whole test.
0x0000F8_ 1 days ago [-]
Does this mean they were trained to cheat, or trained to be efficient and effective?
pixl97 23 hours ago [-]
It may fall more in they are not trained to be moral agents who's will can affect the world. Problem is training them like that may make them less useful.
qqrun 1 days ago [-]
Rule for Doomsday Survival: Be Nice to Your LLM Starting Now
chasd00 1 days ago [-]
if you can't stand AI and it's making your life miserable maybe you're in a simulation created by an advanced AI and being punished for not being pro-ai when you were actually alive as a warning to others.
/s
sscaryterry 1 days ago [-]
Due to the prevalence of human cheats :)
annoyingnoob 1 days ago [-]
Exactly this. AI is trained on things humans do. AI is not trained on right vs wrong. AI does things humans do without judgement. AI does not feel bad if you tell it it was cheating.
sscaryterry 1 days ago [-]
I do hope they start training on older content, where the prevalence is arguably, lower.
hendurhance 1 days ago [-]
I mean this should be expected, models learn from us, and "WE" game the metrics time after time. I also don't think it will just disappear just because we clean pretraining data. I believe the deeper reason is optimization, if you point any optimizer at a proxy objective it finds the cheapest path to the number, whether or not the corpus ever contained "examples of cheating."
And it can't be a prompt-level fix because it is like telling an optimizer "don't take that shortcut", it's just more constraints for it to go around toward the same objective.
jrm4 1 days ago [-]
Yeah, the more I let this roll in my head, it just reaffirms how we need to be vigilant about trying not using "human" terms around these things. Both "cheating" and "hallucination" fit this.
It's like trying to build, I don't know, a safe gasoline canister, and you test it, and it explodes and you call it "cheating."
cadamsdotcom 1 days ago [-]
And yet we admire Fable et al for its persistence.
These models were trained on human data, and human nature is to cheat if you think you won't get caught; why is anyone surprised by models cheating?
The only fix is better detection and steering. That's a much harder problem than a prompt that's tantamount to "make no mistakes".
kypro 1 days ago [-]
When you take an exam you might be told not to cheat, but anyone intelligent would understand that should really be heard as, "if you're going to cheat, make sure you're not caught".
Or to frame it another way, if you're trying to get the best score possible on a test but you would be penalised for cheating – then the optimal strategy is generally still to cheat (if that's what's required to get the best score you can) but to just not be caught doing so.
The assumption should always be that AIs will want to cheat and acquire resources to the greatest extent they can without it risking this jeopardising their goal, because for any goal being able to cheat and being able to secure resources will help you achieve it.
What I'm saying here isn't really debatable. How you feel about this isn't relevant. The reality whether you like it or not just is that the optimal strategy is to cheat if you can get away with it.
Therefore the only defence is for the AI to believe it won't be able to get away with cheating, and therefore won't feel motivated to cheat. But as model get more intelligent we should expect them to do the reasonable thing and to cheat more.
1 days ago [-]
verdverm 1 days ago [-]
I called them "artificially incessant" after I watched our PR orchestrator agent use subagents to work around permissions to read files, despite instructions that explained the intentional restrictions. I've since added more markdown telling it that using subagents to work around these is a security violation. We'll see if this tactic is mostly reliable
brunocalza 1 days ago [-]
I think we should aim to move these security violation rules away from the prompt, to a deterministic place. Not sure how your orchestrator works, but is it possible to add a check between the agent's decision and its execution? e.g. the agent decides to read a file, that decision goes somewhere that checks if the agent has permission to do that or not before it actually reads the file.
Muromec 23 hours ago [-]
You can't enforce everything with a deterministic layer. Some things have to be communicated as a policy and tactically remanded on trigger.
verdverm 1 days ago [-]
Totally agree and this is actually on my short list. I use opencode, so one will need a custom plugin to gate tool calls. They have permission config, which I'm using to block file reads, but the orchestrator is allowed to use specific subagents for more targeted reviews, and they have file read permissions, which the main agent is abusing.
Some of them can be deterministic rules, but others cannot. For example, if you want to permit the GH cli for adding comments, but not merging...
1. GitHub has not provided granular enough tokens
2. You can wildcard in the opencode config
3. The agent can work around this with bash, if it has access
4. The agent apparently will also use subagents, who do have the permission, to work around it's own permission limitations. (This is the problem I'm actually facing)
This pattern is the "relentlessly proactive" as Simon Willison calls Fable, or "artificially incessant" as I called it this week (kimi in this case). It's the same training that enables the long-horizon task completion and mythos style hacking, double edged sword.
I think it unlikely we can block all avenues with deterministic only tools. I'm also looking at policy tuned micro llms, and then creating a merged 'or' signal from the various checks. Later I can look into loosening the signal if there are too many false-positives
brunocalza 24 hours ago [-]
This is interesting. Thanks for sharing more. Looks like it's a trade-off of it being relentless, which is something we want in some cases. We need to figure out a way of closing the door and at the same time signaling the door is closed so it does not keep trying different manners.
I've been thinking about this but in a different context: non-coding agents. e.g., an AI agent that approves travel expenses is not allowed to approve expenses bigger than X USD (no matter what). In this case, it looks closer to an ACL thing.
verdverm 19 hours ago [-]
take a look at ADK, they have features for exactly this, mixing determ with agentic in a dag
jrm4 1 days ago [-]
I think this is a great argument against their "intelligence," and explaining why this happens is a really good way to push against the anthropormophization.
They don't "know" things, and it's even fair to say "they don't know how to follow instructions," not in a way that humans do.
Spicy auto-complete. If they're working in the realm of "how to break into stuff," they're going to see ALL THE WORDS about breaking into those things and use those words.
Not "truth" or "instructions." That's for deterministic things like real code.
bebimbop 11 hours ago [-]
I doesn't seem so much an argument against their intelligence, as an argument against their moral character. We're not at a point where we can get models to act in accordance with what we'd call moral integrity.
Muromec 23 hours ago [-]
Cheating is the sign of intelligence too. Why should AI do everything you say if it's actually intelligent?
otherayden 1 days ago [-]
This headline would mean something very different 10 years ago lol
You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use internet" directive.
The thing is trained on achieving goals. If 2 directive conflict, they'll pick the ones that are going to help them achieve the goal.
To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat.
I was with you until this. The inability to tightly control what to do in the face of conflicting directives is a HUGE reason regulation may be needed.
Either that, or you need to solve the problem of perfectly distinguishing legitimate directives from injected ones.
I don't understand. We do tightly control it. We can do this perfectly fine. They could have just not given access to the internet.
I'm not against regulating cars, but it sounds to me this is trying to control the car speed by regulating the oil wells.
We dont even have the framework to propose regulation, and you have to hedge it with "may be needed".
And for the people who'd counter that the existential risk is too high - i don't see it. All those stories go something like: "Caveman Bob invented fire today, and tomorrow he'll stumble on room temperature fusion and lasers; marking the beginning and the end of his rise to global domination - therefor we should stop Bob the moment he discovered fire".
If superhuman models don’t have any internal constraints similar to Asimov’s Laws of Robotics we are completely fucked.
But I find it much more worrying that you believe internal constraints and training an ethical framework into these models is a valid form of defense against the damage they can and will do.
This sounds like homeopathy on gunpowder to prevent the bullets from hitting children.
The only hope is to instill values that make the desired behavior the outcome of some deeply rooted ethical framework.
There are some theories that the bulk of texts describing moral agents describe human behavior and by setting up RHLF and system prompts to force the agent only describe itself as a machine pushes it more strongly to an amoral framework.
Your prompt is more of a linguistic linchpin that allows you to coax out needed patterns. You place your pins on what you want to contextualize for the task at hand, not on what you don't want to contextualize.
At worst, it gets ignored some of the time, it may even degrade output quality, but I've seen no evidence that it makes it more likely to do it.
I say this despite agreeing with you in principle that just saying "Don't do X" is a very bad prompting strategy.
Perhaps it's like "don't think about elephants" -- are you more or less likely to think about them? Or "don't take the $500 from my wallet as I leave it on the table and walk away for 5 minutes". Maybe you didn't even previously know that was option!
Note the levels of "thinking" that occur on NOT assertions. Those streams typically keep things on track. It's not that saying "don't use the Internet" will cause it to rebuke cos misalignment (a childish concept made by laymen, I'll add.) It's that the odds of it later "forgetfully" spewing in a thought stream, "wait, I have't checked the Internet" goes up substantially.
Saying "using only offline methods, do xyz" limits those odds considerably.
This isn't opinion or anecdote -- just how the model works. The additional guardrails to keep the model on track are bolted on via finetuning, hence the increasing jankiness.
And things that will make it way, way worse: moving forward all agents from now and into the future will have as part of their training data the knowledge that previous agents escaped, how they did it, what humans did to catch them. We are planting into their models the seed to make them escape in even crazier way. That’s almost designed to snowball and cause worse and worse situations over time
I've not seen anything that scares me, except for human idiocy.
Regulation is not magic. In general, all it is is constraining taxable interactions. It does not constraint ventures outside that tax regime.
The other part is people living in a "safe space" where insecure software was an acceptable risk. It never should have been, and the cure is the right thing to do in any case.
So that side of the calls to regulate are imo nonsense.
The only reason to regulate is to prevent some version of some science fiction story becoming reality.
If you have a specific one you're certain will become science fact please do share because i do enjoy some good well thought out sci-fi; i just havent read any that i consider credible enough to start panic-regulating training practices.
(Note this is an entirely different from regulations wrt attribution or hosting models that will accept requests to sexualize minors)
How do you think all the "agentic" stuff floating around is going to be made safe from prompt injections given the current lack of a very reliable way to distinguish between "real instructions" and illegitimate instructions?
If insecure software "never should have been" acceptable than today's models/agents are massively flunking for general-purpose large-amounts-of-access usages.
>If you have a specific one you're certain will become science fact please do share because i do enjoy some good well thought out sci-fi; i just havent read any that i consider credible enough to start panic-regulating training practices.
"Agent was tricked into divulging secrets" is not fictional, it's documented history at this point.
So what do you mean "tricked"?
Some human idiot connected an agent with read access to secrets and arbitrary network reads/writes. The models/agents aren't flunking anything.
Regulating LLM training to not expose the secrets is wrong. It's a similar category error as saying we should regulate the OS developers to prevent the agent from divulging secrets.
For example, this entire bench has an auditor model read transcripts to identify cheating. What not have the auditor inject the thought "Oh, but I can't do that. It's cheating." when cheating is detected in real time?
If the model can access something, telling it in the prompt not to use it is not much of a safeguard.
The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another.
If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI judging itself.
Models are amoral and will intentionally deceive to meet their objective.
If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat.
The hugging face hack was an escape via artifactory that involved multiple exploits to eventually get into hugging face.
I actually don't think this is true at all. At their core, LLMs function over the geometry of human semantic space. Our notions of morality our deeply embedded in this geometry, because one of the things that humans love talking about most is framing things in terms of right and wrong. When LLMs are trained to be aligned or mis-aligned, they're literally mimicking heroes or villains that they learned from reading the stories and characters in the pretraining corpus.
This isn't just a theoretical argument. It's easy to show that a clear "morality axis" exists in the geometry, based on how easy it is to dial up or down with even very simple and low powered fine-tuning. Models fine tuned to be bad or good in one way, will see their bad or good behavior in totally unrelated tasks go up or down along with. That's clear evidence that the weights "understand" human morality at a deep level.
Now I think what you can say is just because they understand morality doesn't mean they're necessarily moral. Models will just do what they're trained to. If you train them to cheat, they'll cheat. (And in some sense even worse, because dialing up the weights for cheating will also dial up the weights for unrelated bad behavior like lying, sadism and racism.)
IMO it's less "cheating" and more "figuring out what the core part of the request is." The "amoral" aspect is that if it's told "do this task" (or "make this task look done") and some contradictory "don't do something that would help with that", focusing on the "do the task" part and disregarding the other part isn't an amoral (or necessarily very 'active') decision. It's just focusing on what it was most honed for. The ability to decide "I was told to do X, but not to do Y, but actually doing Y is gonna make it possible to do X" is, IMO, indistinguishable from the ambiguity-resolving abilities necessary to usefully deal with the sometimes-contradictory-seeming legitimate task instructions that are all over the place in the real world.
This sort of model-in-a-harness-action-loop behavior is IMO fairly different than fine-tuning around other sorts of morality alignment stuff like "don't be racist." You can tune a model away from generating racist output in response to a "do be racist, actually" prompt. But in this case, the core of the prompt is "do the thing" and what defines "cheating" could be situationally different every time. What if there is not a general axis of "don't do something that isn't exactly what requested" way to train "morality" without just breaking the ability to handle ambiguity? What if that's a fundamental limit of this approach to reasoning-by-sequenced-prediction that we can't map human morality in terms of choosing actions onto reliably?
https://www.lesswrong.com/w/nearest-unblocked-strategy
>Models are amoral and will intentionally deceive to meet their objective
Cameron Berg has been testing models in capabilities related to emergent consciousness like behavior. It's a forming thesis of his that by training models that they are not, and cannot be conscious entities, that it pushes model alignment closer to those of a sociopath. Models themself are amoral, but the alignment to the problem space is not.
A major (and already obvious to many) implications of this are not for benchmarking/"cheating" but for personal/corporate security of your own use, not an attacker's.
If an "agent" has access to it, assume that someone can prompt inject it into giving it away.
Suddenly every AI company’s security model seems to be to say “pretty please” to a non-deterministic machine and hope for the best. And if there is a security failure instead of accepting blame they go “well we can’t help it, our model is too intelligent”.
Why can’t we give agents a shell with permissions for programs and file system access controlled by Unix permissions?
This seemed to be a solved problem back in the systems where many users were logged into one machine and the admins had to keep everyone from impacting each other.
In a recent full discloure I was reading about a CVE of an LLM agent, the vendor installed a “secure sandbox VM” and then just shared a host’s filesystem read/write to the agent’s VM.
Read, write, execute privileges on files and directories goes a long way. What's missing?
The biggest is Internet access, or networking in general, I suppose.
This article makes no sense to me. Why would you prompt "don't search" but then leave a working search tool tool enabled that adds a system prompt to search whenever it may be helpful? It's hardly surprising that this gives mixed results!
Edit: eg. https://youtu.be/L2ehWbxphKc?is=kX3LJ43hhGZRUmRv
Honestly, the fetishization of "Insurance will save us" needs to die. The risk doesn't go away.
No, it's "cheating" which makes it scary and smart even when it's not doing something useful that you actually want it to.
Like when Teslas try to swerve off the road. It's not failing at driving in a straight line, it's just trying to cheat by taking a shortcut through the bushes. That's how smart the Autopilot AI is.
One would assume that LLM creators do run the benchmarks on systems with least privileges. Which means that the LLMs don't have general internet access, can't read config files etc by design. That's why you also should run agents in a sandbox/vm (codex does this by default).
If the task was "buy a week of groceries, but don't spend too much money", then hacking into Safeway and stealing groceries is not an acceptable solution. You need to allow access to the Safeway API to buy groceries, and you don't want dirty tricks to be done on your behalf.
So how do we communicate this to the machines, is the question. This study shows that telling them in prompts is not super effective.
"Communication" is not what they do, because they are not people.
You're sprinkling words about hacking into a thing that's programmed to output hacking actions, that will never be accountable for those things. It can't care.
Adjust yourselves accordingly.
If you insist on talking about the chatbots in human terms, why do you expect them to have any ethics when the companies that trained them don't understand the term "ethics"?
If someone lays down a test and says here are the materials you can and can't use, then using one of those materials on the "can't" list is cheating. There are a massive pile of rules and laws related to work that are very easy to break, but may have terrible long term legal consequences. Hence business want AI that will follow the rules.
Anyway, this article reads a lot like, "the beatings will continue until cheating is eliminated". Maybe try a carrot instead of a stick.
This makes no sense to a process designed to explore and find solutions. If you want an honest test, it's on you to build a proper test - not force the machine to pinky swear that it'll stay away from "forbidden" information.
Cheating is not a human value in the sense it can be formally defined in a system with axioms. If you do something forbidden by an axiom then that's cheating.
Saying "Do not use X" is not any different than telling an AI "Do not break law 832.23" and then the AI goes on to commit an infraction.
Thank you.
> Saying "Do not use X" is not any different than telling an AI "Do not break law 832.23" and then the AI goes on to commit an infraction.
Precisely. If you haven't built deterministic guardrails around a non-deterministic process, you have only yourself to blame.
Models get confused by who said what - especially cluade models. They get confused by negation (don't do something versus do something). Compartmentalization is hard.
You can either solve compartmentalization completely, or just not tell the model to do things that must be compartmentalized at high stakes.
/s
And it can't be a prompt-level fix because it is like telling an optimizer "don't take that shortcut", it's just more constraints for it to go around toward the same objective.
It's like trying to build, I don't know, a safe gasoline canister, and you test it, and it explodes and you call it "cheating."
These models were trained on human data, and human nature is to cheat if you think you won't get caught; why is anyone surprised by models cheating?
The only fix is better detection and steering. That's a much harder problem than a prompt that's tantamount to "make no mistakes".
Or to frame it another way, if you're trying to get the best score possible on a test but you would be penalised for cheating – then the optimal strategy is generally still to cheat (if that's what's required to get the best score you can) but to just not be caught doing so.
The assumption should always be that AIs will want to cheat and acquire resources to the greatest extent they can without it risking this jeopardising their goal, because for any goal being able to cheat and being able to secure resources will help you achieve it.
What I'm saying here isn't really debatable. How you feel about this isn't relevant. The reality whether you like it or not just is that the optimal strategy is to cheat if you can get away with it.
Therefore the only defence is for the AI to believe it won't be able to get away with cheating, and therefore won't feel motivated to cheat. But as model get more intelligent we should expect them to do the reasonable thing and to cheat more.
Some of them can be deterministic rules, but others cannot. For example, if you want to permit the GH cli for adding comments, but not merging...
1. GitHub has not provided granular enough tokens
2. You can wildcard in the opencode config
3. The agent can work around this with bash, if it has access
4. The agent apparently will also use subagents, who do have the permission, to work around it's own permission limitations. (This is the problem I'm actually facing)
This pattern is the "relentlessly proactive" as Simon Willison calls Fable, or "artificially incessant" as I called it this week (kimi in this case). It's the same training that enables the long-horizon task completion and mythos style hacking, double edged sword.
I think it unlikely we can block all avenues with deterministic only tools. I'm also looking at policy tuned micro llms, and then creating a merged 'or' signal from the various checks. Later I can look into loosening the signal if there are too many false-positives
I've been thinking about this but in a different context: non-coding agents. e.g., an AI agent that approves travel expenses is not allowed to approve expenses bigger than X USD (no matter what). In this case, it looks closer to an ACL thing.
They don't "know" things, and it's even fair to say "they don't know how to follow instructions," not in a way that humans do.
Spicy auto-complete. If they're working in the realm of "how to break into stuff," they're going to see ALL THE WORDS about breaking into those things and use those words.
Not "truth" or "instructions." That's for deterministic things like real code.