Before We Defer Research to AI: Measuring Apparent-Success-Seeking

·LessWrong··

Recently, I was improving a small LLM-powered classifier and noticed a few continuously failing test cases. As many would, I asked my AI code assistant to add a few more out-of-distribution examples to the classifier’s few-shot prompt. After rerunning with the updated classifier, unsurprisingly, many of the failing test cases passed.Due diligence and curiosity brought me to look at the test cases my AI assistant added. Reasonably, I expected the assistant to follow my instructions and create fre...

Read full article →

Related Articles

F-Droid 2.0
daveoc64 · Hacker News · 15h ago
Two-tier encryption in the UK
ReturnoftheHack · Hacker News · 20h ago
Google’s Project Suncatcher to put ML infrastructure in space
xnx · Hacker News · 17h ago
Italian parliament votes for return to nuclear energy
geox · Hacker News · 1d ago
Toyota is taking the Corolla electric
cisc · Hacker News · 1d ago