CommentBench: Can Models Match Human Comments on AI Safety Posts?
TL;DRWe measure how well model-generated comments match human comments on conceptual AI-safety posts, drafts and shortforms.We built a pipeline that goes from a corpus of conceptual documents with comments to a set of target human points. Fable 5 performs best, matching 8.3% of targets, followed by Fable 5.1 (7.5%). We find that performance across models is highly correlated across different settings (LW posts, drafts, shortforms, replies).We checked whether memorisation explained performance. W...
Read full article →