AI Inference 3 MoE Routing During Prefill: Why Experts Receive Unequal Work Aug 13, 2026 From a Regular Language Model to Mixture of Experts Aug 13, 2026 What Porting llama2.c to an OpenCL Device Taught Me About LLM Inference Aug 12, 2026