Before We Defer Research to AI: Measuring Apparent-Success-Seeking

·LessWrong··

Recently, I was improving a small LLM-powered classifier and noticed a few continuously failing test cases. As many would, I asked my AI code assistant to add a few more out-of-distribution examples to the classifier’s few-shot prompt. After rerunning with the updated classifier, unsurprisingly, many of the failing test cases passed.Due diligence and curiosity brought me to look at the test cases my AI assistant added. Reasonably, I expected the assistant to follow my instructions and create fre...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 15h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 1d ago
Xiaomi Mimo 2.6 live post-training dashboard
krackers · Hacker News · 8h ago
Training a 4B model to produce 81% faster query plans than Postgres
polyphilz · Hacker News · 9h ago
Nvidia announces native GPU programming in Rust
nonmaskable · Hacker News · 17h ago