Can Large Language Models Identify Novel Threats? Part 1: Mirror Life and the Classification Gap

·LessWrong··

[Cross-posted from On Failure States. This is Part 1 of an independent AI safety research series examining LLM safety behavior on unclassified emerging threats.]Can an LLM refuse a harmful uplift request when the topic in question hasn’t been identified as dangerous yet? In 2022, mirror RNA polymerase was actually created, a key step towards the creation of mirror life, and in 2024 the scientific community warned against any further research on it.[1][2] Having said that, mirror life is not curr...

Read full article →

Related Articles

Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 18h ago
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 1d ago
A 40ms Go garbage collector pause caused by swap
shellpipe · Hacker News · 13h ago
Car is a smartphone on wheels. Here's who's listening
longhaul · Hacker News · 22h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 1d ago