An "Anthropic Principle" for Formulations of AI Alignment

·LessWrong··

This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. Here I explore one path toward formulating AI alignment without needing to formalize humans or our values, looking instead at the well-integrated computational power characteristic of agents capable of confronting the alignment problem.One natural formulation of AI alignment, the study of how to be sure our highly capable ...

Read full article →

Related Articles

The ChatGPT/Codex app bundles a full copy of LibreOffice
timpera · Hacker News · 14h ago
The Emergent Symbolic Structure of Artificial Neural Networks
schmuhblaster · Hacker News · 6h ago
I trained a small transformer in 1.5hrs and it beats many LLMs
porridgeraisin · Hacker News · 1d ago
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
carloslfu · Hacker News · 17h ago
Refurbishing a Tektronix TDS7104 Oscilloscope
jwise0 · Hacker News · 14h ago