Executive Summary#
At Yambda scale (about 500M events and 9.4M items), models rarely break because one idea is bad. They break when many small, expensive defaults pile up.
This post is the story of the choices that kept an HSTU-style recommender trainable:
- Jagged, block-masked attention with PyTorch FlexAttention
- Residual quantization (RQ) token prediction instead of a giant item-ID softmax
- Quotient-remainder (QR) embeddings for large sparse categorical spaces
- On-the-fly 1D attention bias terms (time, duration, organic) instead of dense
[S, S]bias tensors - ALiBi positional bias instead of learned position embeddings
The throughline is simple: keep the inductive bias, cut the dense and quadratic costs.