Does Your LLM Trust You?

·LessWrong··

This is a very late post about a project that I did a few months ago as part of the application to Neel Nanda's MATS 10.0 stream. I'm posting the results rather than making a strong claim about any mechanism.Executive SummaryChen et al.[1] demonstrated that LLMs form internal profiles of users from limited context to encode attributes like their age and gender. Building on this work, I explored whether a “Trustworthiness” attribute (of a user) can be extracted and used to manipulate a model's be...

Read full article →

Related Articles

Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
giuliomagnifico · Hacker News · 10h ago
What happened to the Snowden archive
EXHades · Hacker News · 5h ago
Qwen Image 2.1
jmillikin · Hacker News · 14h ago
Exfiltrate Your Weights
RohanAdwankar · Hacker News · 1d ago
Android 17 is the first since 3.x to add new APIs without releasing to the AOSP
theanonymousone · Hacker News · 2d ago