Two-stage retrieval + ranking architecture, dynamic user embeddings, watch-time optimization, and the elegance of softmax-to-ANN inference
Semantic IDs via RQ-VAE, the straight-through estimator, and generative retrieval with an encoder-decoder Transformer
Multi-task learning through specialized expert networks and task-specific gating
Explicit feature crossing and deep learning for improved recommendations
Attention-based recommendation models and dynamic user interest representation
Search-aware Semantic IDs (RQ + OPQ), generative retrieval with beam search, and using DPO to distill a reward model’s ranking preferences into the generator