Scaling Laws for Mixture Pretraining Under Data Constraints

Apple ML Research··

As language models scale, the amount of data they require grows – yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this scarce but valuable target data with abundant generic data, which presents a fundamental trade-off: too little target data in the mixture underexposes the model to the target domain, while too much target data repeats the same examples excessively, yielding diminishing returns and ev...

Read full article →

Related Articles

Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces
nunodonato · Hacker News · 1d ago
OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
donsupreme · Hacker News · 3mo ago
Accelerating Gemma 4: faster inference with multi-token prediction drafters
amrrs · Hacker News · 3mo ago
A couple million lines of Haskell: Production engineering at Mercury
unignorant · Hacker News · 3mo ago
Using “underdrawings” for accurate text and numbers
samcollins · Hacker News · 3mo ago