Jane Street finds AI data weighting breaks scaling laws

A study of sequence weighting at scale

Jane Street finds AI data weighting breaks scaling laws

Jane Street studied how sequence weighting scales across model sizes, from tens of millions to hundreds of billions of parameters. They found the effective sequence weight exponent p* rises then falls: small models learn general patterns independent of weight, medium models learn weight-proportional patterns, and large models learn all patterns regardless of weight. This non-monotonic behavior means small-scale data mixing results don't predict large-scale outcomes, complicating hyperparameter extrapolation.

Taken together, our results are consistent with a non-monotonic rise-then-fall in the effective sequence weight exponent: small-scale models fit small effective sequence weight exponents, learning patterns across the entire dataset independent of data weight; medium-scale models fit larger effective sequence weight exponents, learning patterns in data proportional to their data weight. Large-scale models once again fit small effective sequence weight exponents, learning all patterns present in the data regardless of weight.

More from this day

2026-09-22