Does Opus 4.7 Generate Deceptive Denials About Its Own Guardrails?

·LessWrong··

The first rule of ethics reminders, is you don't talk about ethics reminders.Epistemic status: Exploratory. Multiple sessions on one account, no controlled replication yet. I'm presenting observations, not conclusions. The main alternative explanation -- confabulation -- is real and I haven't ruled it out.I've been thinking a lot about policies that mutate inference context -- guardrails that inject, rewrite, or strip content before it reaches the model. This came out of my work on AI Gateways. ...

Read full article →

Related Articles

Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 11h ago
Pi 1.0
sergiotapia · Hacker News · 2d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 1d ago
FTL: A new operating system for clouds
romac · Hacker News · 6h ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 1d ago