Before We Defer Research to AI: Measuring Apparent-Success-Seeking

·LessWrong··

Recently, I was improving a small LLM-powered classifier and noticed a few continuously failing test cases. As many would, I asked my AI code assistant to add a few more out-of-distribution examples to the classifier’s few-shot prompt. After rerunning with the updated classifier, unsurprisingly, many of the failing test cases passed.Due diligence and curiosity brought me to look at the test cases my AI assistant added. Reasonably, I expected the assistant to follow my instructions and create fre...

Read full article →

Related Articles

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
riordan · Hacker News · 18h ago
Mistral Patent for “Code implemented tool calls”
theanonymousone · Hacker News · 14h ago
Kinney Drugs pulls back AI phone assistant after hundreds of customer complaints
kotaKat · Hacker News · 13h ago
Study links GLP-1 drugs to bigger jump in women's employment than a degree
metadat · Hacker News · 12h ago
Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
HenryNdubuaku · Hacker News · 10h ago