Could an AI’s secret loyalty change? by Shu-Hao Liu

·Nuno Sempere··

View origi­nal post here:https://​​github.com/​​fi­dari-in­sti­tute/​​fi­dari/​​blob/​​main/​​memos/​​the­o­ret­i­cal/​​1.mdWe already know from Lamer­ton, et al. (https://​​arxiv.org/​​abs/​​2605.06846)’s pa­per on that AI agents could be trained to have se­cret loy­alties that are hard to de­tect. But that raises a ques­tion: could their se­cret loy­alty change?Ever since the be­gin­ning of neu­ral net­works, it has always been a loss op­ti­miza­tion game. Why should not the same thing be with...

Read full article →

Related Articles

Humans Are Not Chinchillas: Revisiting “How Quick and Big Would A Software Intelligence Explosion Be?” by Forethought
Forethought · Nuno Sempere · 1d ago
How much should we worry about the pneumonic plague lableak in Siberia? by Drew Spartz
Drew Spartz · Nuno Sempere · 1d ago
More than 1000 cases of Irkutsk Plague?
Joe and Seth · Manifold Markets · 2d ago
Will the Irkutsk Anti-Plague institute incident turn into an outbreak?
Anthem · Manifold Markets · 2d ago
How many launches will SpaceX have in October?
Fastcar99 · Manifold Markets · 4d ago