Does Your LLM Trust You?

·LessWrong··

This is a very late post about a project that I did a few months ago as part of the application to Neel Nanda's MATS 10.0 stream. I'm posting the results rather than making a strong claim about any mechanism.Executive SummaryChen et al.[1] demonstrated that LLMs form internal profiles of users from limited context to encode attributes like their age and gender. Building on this work, I explored whether a “Trustworthiness” attribute (of a user) can be extracted and used to manipulate a model's be...

Read full article →

Related Articles

Xbox goes down. You can't play games you own on disc
surprisetalk · Hacker News · 10h ago
Ten advances in mathematics and theoretical computer science
milkshakes · Hacker News · 1d ago
Germany Records Historic 12B KWh Solar Feed-In in July 2026
johnbarron · Hacker News · 9h ago
Keyv and friends compromised in active Shai-Hulud supply chain attack
cimi_ · Hacker News · 11h ago
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
doppp · Hacker News · 6h ago