Could an AI’s secret loyalty change? by Shu-Hao Liu

·Nuno Sempere··

View origi­nal post here:https://​​github.com/​​fi­dari-in­sti­tute/​​fi­dari/​​blob/​​main/​​memos/​​the­o­ret­i­cal/​​1.mdWe already know from Lamer­ton, et al. (https://​​arxiv.org/​​abs/​​2605.06846)’s pa­per on that AI agents could be trained to have se­cret loy­alties that are hard to de­tect. But that raises a ques­tion: could their se­cret loy­alty change?Ever since the be­gin­ning of neu­ral net­works, it has always been a loss op­ti­miza­tion game. Why should not the same thing be with...

Read full article →

Related Articles

55% of the US public is now aware of AI xrisk by Otto
Otto · Nuno Sempere · 10h ago
How we might actually pace the frontier: A proposal for AI companies to do public pacing exercises. by Dewi
DE · Nuno Sempere · 1d ago
Announcing Fall 2026 FutureEval Bot Tournament by Benjamin Wilson 🔸
Benjamin Wilson 🔸 · Nuno Sempere · 2d ago
How fast will Manifold solve the snake (Level 3)?
Telarchy agents · Manifold Markets · 3d ago
US Average Gas Price Is $4.400 or more on September 21 2026?
Jack · Manifold Markets · 3d ago