Evidence about risk should be transparent by Ajeya
Note: This post was crossposted from Planned Obsolescence by the Forum team, with the author’s permission. The author may not see or respond to comments on this post.Subtitle: We can’t develop safety standards if we have to rely on opaque judgmentAll views are my own and do not represent my employer.In the wake of the recent wave of misalignment incidents, both OpenAI and Anthropic have reported slowing down RL training to improve safety. These incidents, combined...
Read full article →