RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

Apple ML Research··

Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout gr...

Read full article →

Related Articles

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
donsupreme · Hacker News · 5mo ago
OpenAI just dropped 700 preprints of mathematical proofs and counterexamples
tootie · Hacker News · 4h ago
Accelerating Gemma 4: faster inference with multi-token prediction drafters
amrrs · Hacker News · 5mo ago
A couple million lines of Haskell: Production engineering at Mercury
unignorant · Hacker News · 5mo ago
Using “underdrawings” for accurate text and numbers
samcollins · Hacker News · 5mo ago