Some Thoughts on The Environment Problem in Agent Training

·LessWrong··

As Large Language Models move away from being chat interfaces and become increasingly autonomous actors in the real world, a few insights about evaluation and training of these systems emerge, and I'd like to discuss them.Context:I've gained the insights and ideas laid out below through ongoing work I'm doing. This post serves to outline my working model in pursuing it, and constraints and lessons I learned along the way. Some of these are offered as learned lessons, others as assumptions, and s...

Read full article →

Related Articles

Slovakia finds Russian backdoor in traffic speed cameras
dredmorbius · Hacker News · 5h ago
GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
ed-is-ai · Hacker News · 3h ago
Coconut Oil Jet Fuel Matches Kerosene's Efficiency in Engine Tests
mdp2021 · Hacker News · 4h ago
There's no reason for software to be slow anymore
Jach · Hacker News · 1d ago
JIT Compiling Code in 5μs
zX41ZdbW · Hacker News · 14h ago