Meta GEM training efficiency shows why LLM-scale recommenders need workload-specific kernels, precision, parallelism, memory, and profiling.