This is one compaction away from being truncated to "allowed examples: `git add`...".
System prompts aren't safeguards.
A step in the right direction is auto-review, available in claude-code, codex, and Cursor products. This is not foolproof either.
This is why remote calls should be gated through an MCP or other API gateway. The MCP can restrict calls even when the provider lacks scoped privileges for their integration keys.
Does this consistently work for you? I have something like this plus some commands that are explicitly in a deny list in the harness. Roughly twice a week, the model manages to run the deny listed commands, that I need afterwards to manually revert.
Almost completely consistently. I can vaguely remember one slip up in over 6 months of daily usage. Good enough that I'm not inclined to use anything more heavy to guard against this. My agents.md is small and fokused. I only use Sol in pi.dev and Fable in Claude Code.
I have this in agents.md now: