Could an AI’s secret loyalty change? by Shu-Hao Liu
View original post here:https://github.com/fidari-institute/fidari/blob/main/memos/theoretical/1.mdWe already know from Lamerton, et al. (https://arxiv.org/abs/2605.06846)’s paper on that AI agents could be trained to have secret loyalties that are hard to detect. But that raises a question: could their secret loyalty change?Ever since the beginning of neural networks, it has always been a loss optimization game. Why should not the same thing be with...
Read full article →