Netflix's GenRec: An LLM-Native Ranker Beats Production with 40x Less Data

GenRec: Towards LLM-Native Recommendation at Netflix

Netflix's GenRec: An LLM-Native Ranker Beats Production with 40x Less Data

Netflix introduces GenRec, an LLM-based recommendation ranker that post-trains an internal foundation model on Netflix-specific data and objectives. By verbalizing user histories and item metadata into natural-language prompts, GenRec reduces reliance on hand-engineered features, shifting focus from feature engineering to context engineering. In a large-scale A/B test, GenRec achieved statistically significant improvements in both short-term and long-term metrics over a mature production ranker, despite using roughly 40x fewer labeled examples and input signals. The system runs in prefill-only mode on Netflix's LLM serving stack for cost efficiency.

In a large-scale A/B test against a well-tuned production ranker, GenRec achieves statistically significant improvements in both short-term and long-term online metrics, while using only a small fraction of the Phase-2 labeled data and input signals.

More from this day

2026-08-15