SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
The problem is that full-parameter post-training of trillion-parameter-scale MoE models on Ascend NPU SuperPOD faces severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. The method is a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution, integrated as SLAI T-Rex. The system achieves 34.22% MFU with a 2.93x improvement over the open-source baseline, and the specialized DeepSeek-V4-Flash model reaches 71.81% zero-shot Pass@1, outperforming GPT-5.4-Mini by 3.98 percentage points. This matters because it demonstrates a full-stack pathway from efficient trillion-parameter post-training on Ascend infra to domain-specialized models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.