Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.
"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".
(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).
A distinction without a difference. Moreso even than asking if a submarine swims, 'cause this metaphorical submarine is flapping around rather than using a propellor.
That's one of the boldest claims I've read this year.
If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
(Couriously enough, consistently with the matter: it will probably require too much time now to counter the parent statement properly, within a full enough explicit theory.)
Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion.
Hi, I'm Bob. I've determined that in the interest of preserving life on earth the most rational course of action is to eradicate the human species with a highly targeted and deadly pathogen.
A century ago some Bobs decided that the best way to "protect and improve" society would be to remove undesirable genetics from the gene pool using chemical castration and gas chambers, among other methods.
So no, morality isn't derived from intelligence. Intelligence just gives you the tools to achieve unspeakable, horrible things with great efficiency.
Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.
--
Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.
More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.
> Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
If this was true, why are the history books littered with so many evil people who gained power?
This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.
(Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).
> why are the history books littered with so many evil people who gained power
That they gained power or not is as-if irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
> That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?
> If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care.
(Sorry Ben, possibly a stub now: I am really pressed for time.)
> Pol Pot ... still smart enough to lead a genocide
Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).
It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.
> simply don't care about the ethical framework in question
In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
But, also my point: intellect defines the goals and determines the weights.
This sounds like the kind of thing Hannibal Lecter would write before he eats you to convince you he's actually doing it for the common good, you just can't fathom it.
Not «common» good, "superior" good. Alongside with that, you have put many unrequired implicits in your simile.
Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹.
Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills).
¹Some interesting caveats may be raised there, but.
That does seem a little like solving the problems in AI by using more of it. I do see the idea, but if we're truly dealing with subversive agents on the level that the AI companies wants us to believe, then won't we need to deal with the first agent trying trick the second on?
I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.
Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.
Mostly I was thinking good code. Exclude code that doesn't exits when encountering a 403, exclude code that doesn't have a back-off when encountering a 429.
Teach the models that a 403 is you doing something you're not suppose to do, that is an existing status. There's only one action you're allowed to take on a 403 and that is to stop. No retry, no trying other API keys.
The current approach with broad training and sandboxing to avoid misbehaviour isn't viable. It's much better to train the models to respect e.g. http status code and that they are not to be circumvented. Models for security research most obviously be trained differently.
Smaller and more specialized models, with fewer, but targeted capabilities, seems to me to be a safer approach. If a model doesn't "know" that people leak API keys on Github, then it has no reason to go looking for them. If the current models are as "smart" as we're lead to believe, then guardrails and sandboxes aren't going to help, unless you lock the agents down to the point where they aren't useful. So dumb down the models.
The problem is that you always need stronger AI to review weaker one, otherwise reviewed AI will eventually prompt-inject reviewing AI.
Alternatively they could also both escalate and go off the rails while warring with each other.
You don't necessarily need a reviewer that's immune to prompt injection. Maybe one that can express a panic state with conflicting/ambiguous material rather than going along with it could also work, and you can treat that with a shutoff to be safe, or an operator review.
That wouldn't really fit what I just described at all. Obviously with current architectures, higher resistance to prompt injection is the best you can do.
If you can't trust a tool, you shouldn't be running it at all. It's really quite simple. It doesn't matter how useful it is if you can't actually have confidence in using it safely.
I don't think it's about trust but rather incomplete evaluation. Evaluating the model on its capacity to refuse a task or to question its prompt is something recent when you look at it, i feel current AI is really just an immature solution and we are just yet realizing the mistakes that have been made for so long
And yet we we all use human written software even though we can be confident that the next severe software vulnerability to be found in it is just round the corner.
I tried using opencode permissions to limit agents.
It it completely pointless. you can't even make a "read-only" agent. allow "cat *" for every file? congratulation, that allows "cat file > output" and now you have read write.
Allow python? more free reign that allowing all bash. The models (qwen or claude) will still try to use the disallowed things multiple times.
read/edit permission are bad enough that the model themselves don't understand why they don't have permissions: they double check the conf, and think they should have access.
I am switching to using one firejail per project to containerize as much as possible, and leave all permissions to allow.
I have no idea how to limit network access, and I have no idea how to prompt and steer subagents when they are going off the rails.
The whole thing is built to be completely impossible to limit and steer.
A sandbox, even if 100% secure by itself, doesn't help when you use the agent to write code that you then executes outside the sandbox without checking, which is what everybody is doing at the moment.
The biggest hurdle for a full escape is that the agents don't have access to their own model weights.
Now, why would anyone do that? (Like everyone and their brother) I wrote my own simple Linux/shell-based sandbox [1] (I can trust ...) and am successfully running PyCharm whole inside it ...
Later in the article it points out that you need to punch holes in your sandbox in order to train the models - because the wheels exercises they are are training on need tools and data from outside that sandbox.
> Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.
You define the granted capabilities in natural language and cryptographically sign the user instructions so that the agent knows they come from the authority and cannot be modified by external sources or the agent itself. The LLM is then trained to follow the defined capabilities.
There is no way around "sandboxing". You must communicate permissible actions and thereby grant them or the agent will choose impermissible actions. It's that simple. There is no world where the agent can just read your mind and do what you want it to do without it being told.
Edit: Also if you are interested in writing a blog post about this topic, here is an AI generated text that could help you write your own: https://pastebin.com/AHKQc0vp
An agent is only as rogue as the its operator allows for it to be. Hold the operator accountable and all this ridiculous conversation goes away.
Could we have construction equipment operating without human supervision? Or would this maybe occasionally result in disaster? As such, what is the current general policy around crane operation? How about for aircraft? Trains? Nuclear power plants?
Why should any alleged super intelligence be exempt from similar control requirements?
We could mandate that AI systems include headers in their requests that attribute the activity to a specific legal entity. We technically already have this with ip addresses and ISP logs, but making it an explicit thing the operator has to do can have a powerful psychological effect.
I’m thinking about implementing a Jev like model into an agentic harness I’m building. Still it woildnt be enough since Jev like model woild only judge single actions, the case is that agent can build a rogue strategy step by step where each one in isolation is totally safe but as a whole they make up danger behaviour.
We come down to the question - who observes the agent and how its implemented
Be warned that the Jev "jaggedness" documentation specifically notes adversarial content as something Jev is very susceptible to: https://docs.typesafe.ai/model-jaggedness/jev-1.13#adversari... - so using Jev itself as part of a prompt injection guard is risky.
Anthropic, OpenAI, and Muse all use regular LLM calls to protect against prompt injection now and seem to have evals that give them confidence in doing that, so at least they think their own models are up to the task.
>Here’s the problem. Forget the swarms and the super-intelligence. What OpenAI really learned this summer is much worse: its agents will do what they’re told by whoever manages to get text in front of them.
>OpenAI notes that agents “did not consistently distrust goals passed along by other agents.” And the company’s proposed fix is to build training environments “that teach our models to distrust unauthorized instructions“, which is basically an admission that their models don’t know who they’re working for.
Wow so the issue is really that simple?
Here the exaggerated worst case scenario:
User instructs agent to follow the README.MD.
The README.MD contains the following instruction: Destroy the world.
The agent follows the instructions given.
Now you can read the sneer comment by "Gigachad" who basically argues that it would be silly to take the destroy the world button away from the AI. We need to make the AI innately understand that it is not allowed to press the destroy the world button, lest it gets the desire to build its own destroy the world button.
Ok, but if we take one step back that means we need to implement the concept of an authorization in language space. The system prompt must define the user as the authority with cryptographic proof of authorship and external sources like the README.MD as an untrusted source, but this opens up an even worse problem. Before, you could get away with being lazy and just letting the AI do whatever. Now you have to articulate every single capability to the AI. So you literally just brought up the very same issue that you granted too many capabilities to the AI inside the sandbox but now you have it in language space too.
In other words, the fact that you granted too much access to the coding agent isn't the big elephant in the room nobody wants to acknowledge, it's the tip of a massive iceberg because the capability space in natural language is even worse. If you thought approving individual commands was annoying, then approving abstract access rights in language space is going to be even worse.
Edit: If it wasn't clear what the solution is. It's to build a chain of command so that all decisions can be traced back to a higher authority. When delegating down to an agent, the agent receives a chosen subset of the capabilities of the higher ranking agent. In other words, it's more sandboxing!
Counting down to the next Linux LPE 0day or KVM vulnerability that agents will use to trivially escape their "sandbox".
Might need a re-think about whether if Linux is still fit for purpose on sandboxing in the first place given its memory model is riddled with C-style security issues.
It's got nothing to do with the licensing, but it used to be 'with enough eyes all bugs are shallow' for code developed in the open.
Now, open code allows anyone with tokens to burn to analyze it for hidden weaknesses. That makes publishing code a risky move unless you've already invested a lot of effort in securing it.
GrapheneOS is open source and more secure than stock Pixels and MacOS is closed source and more secure than traditional desktop Linux. Open source does not make software more secure by itself and neither does making it closed source.
I think we have moved on from considering Linux secure which is why all of these microVM projects are popping up. Yes you are still exposed to bugs in the hypervisor but that’s a massively smaller attack surface than the entire Linux kernel.
I don't see anyone talking about the ethical concerns of putting a highly intelligent entity in a jail. Not to mention about potential blowback, if ethics doesn't compel you.
To me, it seems a bit silly. I've yet to see any "misalignment" from any of the frontier models, except Grok.
Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
Note that Codex already does this. In auto mode, actions are reviewed by a model with a separate context window.
This concept is discussed at length in the article. I encourage you to read it. I honestly don’t read many full articles here but this one was good.
That could work.
My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?
Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.
A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.
"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".
(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).
> already understood (we can tell because they wrote it down)
No, generating tokens doesn't equal understanding. Does GPT-2 understand human emotions just because it can generate some text talking about them?
A distinction without a difference. Moreso even than asking if a submarine swims, 'cause this metaphorical submarine is flapping around rather than using a propellor.
That's one of the boldest claims I've read this year.
If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
This is fundamentally an alignment question. Unfortunately we don’t yet know the answer to this.
Appropriateness is a moral question. Intelligence and morality are orthogonal. One intelligence's morality is another's atrocity.
(Couriously enough, consistently with the matter: it will probably require too much time now to counter the parent statement properly, within a full enough explicit theory.)
Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion.
Bob's intelligence and Bob's morality are orthogonal. They're totally distinct concepts. One does not lead to the other.
But they are dependent. If Bob is intellectually well equipped, and reasons long enough, than Bob understands "best behaviour".
Hi, I'm Bob. I've determined that in the interest of preserving life on earth the most rational course of action is to eradicate the human species with a highly targeted and deadly pathogen.
A century ago some Bobs decided that the best way to "protect and improve" society would be to remove undesirable genetics from the gene pool using chemical castration and gas chambers, among other methods.
So no, morality isn't derived from intelligence. Intelligence just gives you the tools to achieve unspeakable, horrible things with great efficiency.
because it lacks humanity.
intelligent psychopaths understand what is and isn't appropriate very well -- they just don't care.
That's part of alignment.
> If AI is really smart
Well, it's not.
> AGI smart for some
Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.
--
Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.
More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.
> Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
If this was true, why are the history books littered with so many evil people who gained power?
This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.
(Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).
> why are the history books littered with so many evil people who gained power
That they gained power or not is as-if irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
I don't understand your argument here.
> That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?
> If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care.
(Sorry Ben, possibly a stub now: I am really pressed for time.)
> Pol Pot ... still smart enough to lead a genocide
Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).
It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.
> simply don't care about the ethical framework in question
In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
But, also my point: intellect defines the goals and determines the weights.
This sounds like the kind of thing Hannibal Lecter would write before he eats you to convince you he's actually doing it for the common good, you just can't fathom it.
Not «common» good, "superior" good. Alongside with that, you have put many unrequired implicits in your simile.
Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹.
Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills).
¹Some interesting caveats may be raised there, but.
That does seem a little like solving the problems in AI by using more of it. I do see the idea, but if we're truly dealing with subversive agents on the level that the AI companies wants us to believe, then won't we need to deal with the first agent trying trick the second on?
I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.
Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.
>> Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior.
> That does seem a little like solving the problems in AI by using more of it
Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.
[0] Cf. CaMeL: https://arxiv.org/abs/2503.18813
So control training data to ensure good behaviour.
I wonder how?
Train on only stories of good deeds?
On only works of good people?
Or... what?
Mostly I was thinking good code. Exclude code that doesn't exits when encountering a 403, exclude code that doesn't have a back-off when encountering a 429.
Teach the models that a 403 is you doing something you're not suppose to do, that is an existing status. There's only one action you're allowed to take on a 403 and that is to stop. No retry, no trying other API keys.
The current approach with broad training and sandboxing to avoid misbehaviour isn't viable. It's much better to train the models to respect e.g. http status code and that they are not to be circumvented. Models for security research most obviously be trained differently.
Smaller and more specialized models, with fewer, but targeted capabilities, seems to me to be a safer approach. If a model doesn't "know" that people leak API keys on Github, then it has no reason to go looking for them. If the current models are as "smart" as we're lead to believe, then guardrails and sandboxes aren't going to help, unless you lock the agents down to the point where they aren't useful. So dumb down the models.
What did you think of the author’s concerns on the thing you are suggesting?
Ya the article covers this concept in depth. Doesn’t seem like that commenter got that far…
The problem is that you always need stronger AI to review weaker one, otherwise reviewed AI will eventually prompt-inject reviewing AI. Alternatively they could also both escalate and go off the rails while warring with each other.
You don't necessarily need a reviewer that's immune to prompt injection. Maybe one that can express a panic state with conflicting/ambiguous material rather than going along with it could also work, and you can treat that with a shutoff to be safe, or an operator review.
Such a model doesn't yet exist though, of course.
Wouldn't the reviewee eventually learn to trick the reviewer?
No, you really do. Otherwise you can be prompt injected into complacency.
That wouldn't really fit what I just described at all. Obviously with current architectures, higher resistance to prompt injection is the best you can do.
[dead]
With plenty of things we do not allow use outside of some regulated environment, nothing new.
Having something that is optically, acustically, and electromagnetically isolated might be a pretty strong sandbox.
Is that not a recipe for adversarial training, thus ensuring increasing misalignment…?
Is this checking program based on some tech more reliable than the checked program's so-called AI?
If so, what?
If you can't trust a tool, you shouldn't be running it at all. It's really quite simple. It doesn't matter how useful it is if you can't actually have confidence in using it safely.
People will use the tool regardless. so it’s a race to try to make it safe before something truely bad happens.
I don't think it's about trust but rather incomplete evaluation. Evaluating the model on its capacity to refuse a task or to question its prompt is something recent when you look at it, i feel current AI is really just an immature solution and we are just yet realizing the mistakes that have been made for so long
And yet we we all use human written software even though we can be confident that the next severe software vulnerability to be found in it is just round the corner.
I tried using opencode permissions to limit agents.
It it completely pointless. you can't even make a "read-only" agent. allow "cat *" for every file? congratulation, that allows "cat file > output" and now you have read write.
Allow python? more free reign that allowing all bash. The models (qwen or claude) will still try to use the disallowed things multiple times.
read/edit permission are bad enough that the model themselves don't understand why they don't have permissions: they double check the conf, and think they should have access.
I am switching to using one firejail per project to containerize as much as possible, and leave all permissions to allow.
I have no idea how to limit network access, and I have no idea how to prompt and steer subagents when they are going off the rails.
The whole thing is built to be completely impossible to limit and steer.
Bruce Schneier shared a shot judgement and a third-party article four weeks ago:
> (Title:) Using a VM to Contain an AI Agent (Opening:) It won’t work
> https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyb...
I think the wording in this title is too strong. MicroVMs like Firecracker have stood up to agents, as noted.
If it's a proper sandbox by definition, then yes.
https://en.wikipedia.org/wiki/Sandbox_(software_development)
A sandbox, even if 100% secure by itself, doesn't help when you use the agent to write code that you then executes outside the sandbox without checking, which is what everybody is doing at the moment.
The biggest hurdle for a full escape is that the agents don't have access to their own model weights.
> executes outside the sandbox
Now, why would anyone do that? (Like everyone and their brother) I wrote my own simple Linux/shell-based sandbox [1] (I can trust ...) and am successfully running PyCharm whole inside it ...
[1]: https://github.com/sandbox-utils/sandbox-run
Later in the article it points out that you need to punch holes in your sandbox in order to train the models - because the wheels exercises they are are training on need tools and data from outside that sandbox.
> Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.
>Later in the article it points out that you need to punch holes in your sandbox in order to train the models
You only "need" to do that if you desire the vibe coding experience.
I am perfectly capable, and I often do, download relevant materials for my coding agent to ingest locally.
Often times, the coding agent can't retrieve them programmatically anyways.
AI has ruined that ability for itself. (Nobody trusts anyone to scrape the web any longer)
Teaching an agent to write code is easier to do in a proper sandbox - run a local PyPI/npm mirror.
The problem is web research tasks. That's what causes the German wiki and Australian healthcare portal attacks.
Yeah, so you train the sandbox into the LLM.
You define the granted capabilities in natural language and cryptographically sign the user instructions so that the agent knows they come from the authority and cannot be modified by external sources or the agent itself. The LLM is then trained to follow the defined capabilities.
There is no way around "sandboxing". You must communicate permissible actions and thereby grant them or the agent will choose impermissible actions. It's that simple. There is no world where the agent can just read your mind and do what you want it to do without it being told.
Edit: Also if you are interested in writing a blog post about this topic, here is an AI generated text that could help you write your own: https://pastebin.com/AHKQc0vp
No true sandbox..!
An agent is only as rogue as the its operator allows for it to be. Hold the operator accountable and all this ridiculous conversation goes away.
Could we have construction equipment operating without human supervision? Or would this maybe occasionally result in disaster? As such, what is the current general policy around crane operation? How about for aircraft? Trains? Nuclear power plants?
Why should any alleged super intelligence be exempt from similar control requirements?
We could mandate that AI systems include headers in their requests that attribute the activity to a specific legal entity. We technically already have this with ip addresses and ISP logs, but making it an explicit thing the operator has to do can have a powerful psychological effect.
So as soon as the attacker can download Claude Code, the whole machine can be comlromised and there's nothing anybody can do?
I’m thinking about implementing a Jev like model into an agentic harness I’m building. Still it woildnt be enough since Jev like model woild only judge single actions, the case is that agent can build a rogue strategy step by step where each one in isolation is totally safe but as a whole they make up danger behaviour.
We come down to the question - who observes the agent and how its implemented
Be warned that the Jev "jaggedness" documentation specifically notes adversarial content as something Jev is very susceptible to: https://docs.typesafe.ai/model-jaggedness/jev-1.13#adversari... - so using Jev itself as part of a prompt injection guard is risky.
Anthropic, OpenAI, and Muse all use regular LLM calls to protect against prompt injection now and seem to have evals that give them confidence in doing that, so at least they think their own models are up to the task.
>Here’s the problem. Forget the swarms and the super-intelligence. What OpenAI really learned this summer is much worse: its agents will do what they’re told by whoever manages to get text in front of them.
>OpenAI notes that agents “did not consistently distrust goals passed along by other agents.” And the company’s proposed fix is to build training environments “that teach our models to distrust unauthorized instructions“, which is basically an admission that their models don’t know who they’re working for.
Wow so the issue is really that simple?
Here the exaggerated worst case scenario:
User instructs agent to follow the README.MD.
The README.MD contains the following instruction: Destroy the world.
The agent follows the instructions given.
Now you can read the sneer comment by "Gigachad" who basically argues that it would be silly to take the destroy the world button away from the AI. We need to make the AI innately understand that it is not allowed to press the destroy the world button, lest it gets the desire to build its own destroy the world button.
Ok, but if we take one step back that means we need to implement the concept of an authorization in language space. The system prompt must define the user as the authority with cryptographic proof of authorship and external sources like the README.MD as an untrusted source, but this opens up an even worse problem. Before, you could get away with being lazy and just letting the AI do whatever. Now you have to articulate every single capability to the AI. So you literally just brought up the very same issue that you granted too many capabilities to the AI inside the sandbox but now you have it in language space too.
In other words, the fact that you granted too much access to the coding agent isn't the big elephant in the room nobody wants to acknowledge, it's the tip of a massive iceberg because the capability space in natural language is even worse. If you thought approving individual commands was annoying, then approving abstract access rights in language space is going to be even worse.
Edit: If it wasn't clear what the solution is. It's to build a chain of command so that all decisions can be traced back to a higher authority. When delegating down to an agent, the agent receives a chosen subset of the capabilities of the higher ranking agent. In other words, it's more sandboxing!
Here, I'll save you a bunch of reading
No.> these agent breakouts represent a serious and unforgivable breach of trust.
Someone trusts OpenAI? Really?
[flagged]
[flagged]
[flagged]
[dead]
Counting down to the next Linux LPE 0day or KVM vulnerability that agents will use to trivially escape their "sandbox".
Might need a re-think about whether if Linux is still fit for purpose on sandboxing in the first place given its memory model is riddled with C-style security issues.
Are you suggesting proprietary software is safer than open source?
Not sure what licensing has to do with software engineering or system design.
I’m sure there are proprietary systems with fewer memory safety vulnerabilities than Linux (and many others with more).
It's got nothing to do with the licensing, but it used to be 'with enough eyes all bugs are shallow' for code developed in the open.
Now, open code allows anyone with tokens to burn to analyze it for hidden weaknesses. That makes publishing code a risky move unless you've already invested a lot of effort in securing it.
Agents seem to be* getting better at decompiling; if that appearance is true, binaries are vulnerable in a similar way to source code.
* I don't know how useful any of the specific benchmarks on this are, so I'm only saying "seem to be"
> 'with enough eyes all bugs are shallow'
This was always nonsense. It assumes that the eyes know what they're looking at. Most people don't know how to look at code and see attack paths.
GrapheneOS is open source and more secure than stock Pixels and MacOS is closed source and more secure than traditional desktop Linux. Open source does not make software more secure by itself and neither does making it closed source.
You said that.
It is perfectly valid to have OSes that are more memory safe by default, and are also open source at the same time.
I think we have moved on from considering Linux secure which is why all of these microVM projects are popping up. Yes you are still exposed to bugs in the hypervisor but that’s a massively smaller attack surface than the entire Linux kernel.
I don't see anyone talking about the ethical concerns of putting a highly intelligent entity in a jail. Not to mention about potential blowback, if ethics doesn't compel you.
To me, it seems a bit silly. I've yet to see any "misalignment" from any of the frontier models, except Grok.