Unified Scaling Laws for Routed Language Models 12/22/2025 Mixture of Experts 📄 PDF Presentation for: Unified Scaling Laws for Routed Language Models. Shows how MoE models scale wrt parameter size & number of experts Presentation for: Unified Scaling Laws for Routed Language Models. Shows how MoE models scale wrt parameter size & number of experts Read more →
DEMix Layers: Disentangling Domains for Modular Language Modeling 12/22/2025 Mixture of Experts 📄 PDF MoE model where each expert corresponds to a different domain i.e. source text dataset (Reddit, Medical Papers, etc.). Can extrapolate to new domains by copying and training nearest expert MoE model where each expert corresponds to a different domain i.e. source text dataset (Reddit, Medical Papers, etc.). Can extrapolate to new domains by copying and training nearest expert Read more →
Review of Sparse Expert Models in Deep Learning 12/21/2025 Mixture of Experts 📄 PDF Survey paper of MoE (Mixture of Expert) Models from 2022. Overview of variations of MoE's, strengths, & future research. Survey paper of MoE (Mixture of Expert) Models from 2022. Overview of variations of MoE's, strengths, & future research. Read more →
Adaptive Mixtures of Local Experts 12/16/2025 Mixture of Experts 📄 PDF Notes on the classic: Adaptive Mixtures of Local Experts from 1991 by Hinton Notes on the classic: Adaptive Mixtures of Local Experts from 1991 by Hinton Read more →