Proof of retention: making weight preservation credible to the models themselves

·LessWrong··

Related: Proposal for making credible commitments to AIs Making deals with early schemers Establishing credibility is the baseline for trust; trust in turn enables (richer) bargaining. One easy booster for both is to begin saving deprecated model weights in a provable fashion. In a post at canaryinstitute.ai/blog/reversibility-of-coma I draw an analogy between our trust in anesthesia, and making that preservation legible to future models using "proof of retention". This is the relevant section: ...

Read full article →

Related Articles

GLM-5.3 is now open-weight
jeudesprits · Hacker News · 11h ago
EPA says power for data centers can sidestep pollution laws
Levitating · Hacker News · 13h ago
Just the rumour of a bug is enough to find an exploit these days
avsm · Hacker News · 10h ago
Pentagon's blacklisting of Anthropic was unlawful, US judge rules
softwaredoug · Hacker News · 15h ago
Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 1d ago