Where Did D Go? A Gap Between ARC's Motivation and Its Formalism
TL;DR: ARC's post does excellent work motivating a p(doom) estimator equal or better than random sampling; however, they evaluate p(doom) over a naive distribution of inputs, leaving them open to test-deploy asymmetry attacks. Trojan theory and cybersecurity practice suggest a lens and compare mitigation options.Context: I really admire ARC's focus here: If there will always be more deployment samples than testing samples, successful testing must compete with sampling in order to prevent deploye...
Read full article →