DeepSeek's Ascend ports make AI infrastructure compatibility measurable. The public APIs and developer workflow can survive a hardware change, while tensor layouts, compilers, kernels, communication, and performance qualification remain platform-specific. The practical result is a portability budget: record which layers are reused, which must be rebuilt, and which claims require fresh evidence before treating an accelerator as a production substitute.
Reading time: 9 minutes · Evidence date: October 6, 2026
TL;DR
- DeepGEMM-Ascend preserves the
deep_gemmpackage and kernel API surface, yet uses Ascend-specific scaling layouts, alignment rules, compilers, and kernels. - FlashMLA and DeepSelect expose shared Python surfaces while publishing different CUDA and Ascend capability envelopes.
- DeepEP-Ascend aligns with DeepEP's
EPBuffermodel, while several collectives, load-balancing paths, graph capture, and storage modes remain incomplete or experimental. - Vendor microbenchmarks prove that the code can approach the tested hardware's limits under disclosed conditions. Production substitution needs end-to-end workload, reliability, and cost evidence.
- A migration decision should price six layers: interface, data, compiler, kernel, communication, and validation.
Compatibility is a budget, not a label
The usual portability question asks whether a workload is compatible. That boolean hides the engineering cost.
A more useful model is:
compatibility budget =
interface adaptation
+ data and numerical adaptation
+ compiler and runtime integration
+ kernel and communication tuning
+ validation and operations
+ upgrade regression
DeepSeek's September 30 code releases reduce several terms. They provide familiar package names, aligned APIs, public tests, and optimized Ascend implementations. The remaining terms appear directly in repository requirements, source trees, data-layout notes, unsupported features, and benchmark conditions.
This is the infrastructure lesson. A stable interface can preserve application code while the performance-critical system underneath receives a substantial rewrite. Portability comes from separating the stable surface from the platform-specific implementation, then testing both.
Six layers of the DeepSeek Ascend compatibility budget
| Layer | What is reused | What remains platform-specific | Acceptance evidence |
|---|---|---|---|
| 1. Package and API | Package names, public calls, buffer abstractions | Supported parameters and modes | Import tests and API contract tests |
| 2. Data and numerics | Tensor roles and high-level dtypes | Scale packing, strides, index types, KV layouts, accumulation | Cross-backend correctness and tolerance tests |
| 3. Compiler and runtime | JIT workflow and some build conventions | CANN, Bisheng, torch_npu, HCCL/HCOMM, device discovery |
A frozen environment matrix and clean builds |
| 4. Kernels | Operation names and reference behavior | CUDA versus Ascend kernel code, scheduling, pipelining, alignment | Per-shape correctness and performance curves |
| 5. Communication | EPBuffer lifecycle and dispatch/combine concepts |
NVLink/NCCL/RDMA versus UBMEM/URMA/HCCL topology | Multi-rank tests under the target topology |
| 6. Production qualification | Workload definition and service objectives | Failure modes, observability, upgrades, tail latency, TCO | End-to-end replay, failure drills, and cost per accepted workload |
Each lower layer carries more hardware dependence. The upper layers reduce source-code churn. The lower layers determine whether the migrated system meets the original service contract.
Layer 1: API reuse is real and valuable
DeepGEMM-Ascend uses the same deep_gemm package name and points users to the original DeepGEMM interfaces. It supports BF16, FP8, FP4 GEMM, MQA logits, and MegaMoE operations. This preserves call sites and the development model around runtime compilation.
DeepEP-Ascend also keeps the deep_ep package and aligns its public buffer APIs with DeepEP V2.5. Training, prefill, and decoding use the same EPBuffer concept. FlashMLA and DeepSelect place CUDA and Ascend support inside shared repositories and Python interfaces.
This work has economic value. Application teams can preserve orchestration, call structure, and much of their test vocabulary. The port turns a full application rewrite into a bounded backend qualification project.
The boundary appears one layer below.
Layer 2: data layouts spend the next part of the budget
DeepGEMM-Ascend documents a concrete difference: Ascend packs each pair of UE8M0 scaling factors along K into an int16 and stores packed values in MN-major order. It exposes transformation utilities because the efficient representation differs from the NVIDIA path.
DeepSelect provides an even clearer capability boundary. Its Ascend path supports BF16 inputs and int32 indices. FP32 sampling and int64 indices remain CUDA capabilities. Input rows also have stride and contiguity requirements, and TopK is bounded at 4,096.
FlashMLA changed its FP8 and FP4 KV-cache format in the September release. The same release targets DeepSeek V4.1 and removes support for Hopper and earlier model versions from the current branch. Its fused sparse kernel and dense MHA paths remain CUDA-only, while sparse prefill and decode have Ascend implementations.
These are healthy engineering choices. Efficient hardware requires native layouts and schedules. The buyer's task is to record them as part of the compatibility contract. A shared function name guarantees a familiar entry point; the dtype, shape, stride, model revision, and output tolerance define the supported operating envelope.
Layers 3 and 4: the workflow transfers, the toolchain and kernels change
DeepGEMM-Ascend requires Ascend 950, CANN 9.20, Bisheng, torch_npu, Python 3.10 or newer, and a C++20 toolchain. FlashMLA detects CUDA or Ascend during its build and selects separate kernel sources. DeepSelect likewise carries distinct csrc/cuda_kernels and csrc/ascend_kernels implementations.
The reusable asset is the architecture around JIT compilation, Python bindings, reference checks, and test entry points. The machine code path is a platform implementation.
That distinction changes planning. Count the following as infrastructure work:
- Freeze a tested driver, firmware, CANN, PyTorch,
torch_npu, compiler, and repository revision set. - Build immutable images for each supported accelerator path.
- Run correctness tests across every shape, dtype, alignment, and mode used by the workload.
- Keep separate performance baselines because an optimization for one backend may move another backend's bottleneck.
- Requalify the matrix when firmware, compiler, model layout, or kernel revision changes.
The vLLM Ascend installation matrix demonstrates the same systems reality at a higher layer. It validates HDK, CANN, NNAL, PyTorch, TorchNPU, Triton Ascend, vLLM, and vLLM Ascend as one version set. Compatibility belongs to a stack revision, not a chip name.
Layer 5: communication makes topology part of compatibility
Compute kernels receive most of the attention, but large MoE systems also depend on dispatch, combine, collectives, and memory movement.
The NVIDIA DeepEP path uses CUDA, NCCL, NVLink, and RDMA. DeepEP-Ascend preserves the buffer abstraction while implementing communication with HCCL/HCOMM, UBMEM, and URMA.
The Ascend repository publishes its current gaps:
- batched all-gather is available, while Ascend reduce-scatter and all-reduce kernels remain under development;
- load-balancing APIs are exposed, while their Ascend communication kernels are pending;
- PP and Engram interfaces remain experimental;
- hybrid communication, CPU-backed Engram storage, and graph capture remain unsupported in the pinned release.
The public API therefore creates a migration seam, while the capability matrix decides whether a specific workload crosses it. A team using only expanded expert dispatch and BF16 combine may fit today. A team depending on graph capture, CPU-backed memory, or a missing collective carries extra workaround cost.
Layer 6: performance claims need an evidence contract
The repositories disclose impressive numbers. Their value increases when the test cell stays attached.
FlashMLA reports up to 410 TFLOPS for sparse prefill and 360 TFLOPS for sparse decode on Ascend 950, described as 95% and 83% of the theoretical peak. DeepSelect reports a 2 to 20 times speedup over torch.topk, measured by its own test script on supported shapes. DeepEP-Ascend reports high dispatch and combine bandwidth across several expert-parallel sizes.
These results establish three things:
- Working implementations and benchmark code exist.
- The implementations can use the tested hardware efficiently in disclosed microbenchmark cells.
- The repositories expose enough structure for a qualified Ascend team to begin reproduction.
The evidence boundary is equally explicit. DeepEP-Ascend says its bandwidth results used Ascend 950DT, CANN 9.2.0, a PoC HDK supplied to DeepSeek, and additional manual configuration. That configuration was not a public release. FlashMLA and DeepSelect figures are author-run repository benchmarks. A local Ascend 950 reproduction was outside this review.
Use an evidence ladder for procurement:
| Level | Evidence | Supported conclusion |
|---|---|---|
| 1 | Repository and source code | An implementation is available |
| 2 | Public tests and benchmark method | A reproduction path exists |
| 3 | Vendor result with frozen conditions | The tested stack reached the reported result |
| 4 | Independent run on target hardware | The result transfers to the buyer's environment |
| 5 | End-to-end workload and failure testing | The platform meets the production contract |
| 6 | Production observation and TCO | The migration creates durable economic value |
Peak utilization answers whether software can drive one device. It leaves end-to-end model throughput, accuracy, cluster reliability, operator effort, and total cost for later gates.
A migration gate for DeepSeek on Ascend
1. Freeze the workload contract
Record model revision, precision, context distribution, batch and concurrency, training or inference mode, expert-parallel size, topology, quality target, latency objective, and failure budget.
2. Build a capability matrix
Map every required operation and mode to one of four states: supported and tested, supported with constraints, experimental, or absent. Include data layouts and lifecycle behavior, not only Python symbols.
3. Pin the complete environment
Store container digest, HDK and firmware, CANN, PyTorch, torch_npu, compiler, repository commits, runtime flags, and topology. The environment becomes part of the artifact.
4. Validate correctness before performance
Cross-check outputs against trusted reference implementations. Cover boundary shapes, NaNs, padding, alignment, determinism, numerical tolerance, and multi-rank behavior. Run task-level quality tests after any quantization or layout conversion.
5. Measure the real bottleneck
Run kernel curves, then complete-layer and end-to-end workload replays. Track tail latency, throughput, memory, communication, compilation time, failures, recovery, and accepted-task quality. Every optimization can move the bottleneck.
6. Test operations and upgrades
Exercise cold starts, cache invalidation, node loss, rank failure, degraded links, rollback, observability, and a version upgrade. A second hardware backend adds CI cells, images, dashboards, on-call knowledge, and release coordination.
7. Calculate cost per validated workload
Combine hardware, energy, engineering, support, capacity utilization, retries, failure recovery, and upgrade regression. Compare completed and accepted workloads rather than peak arithmetic alone.
This gate turns portability into a controlled experiment. The result may favor Ascend, CUDA, or a mixed fleet. The same evidence contract supports all three outcomes.
What DeepSeek actually changed
DeepSeek moved the decision boundary. Ascend support now starts with public, optimized implementations and aligned interfaces rather than a blank backend. That lowers discovery cost, preserves application structure, and gives kernel teams auditable starting points.
The remaining budget sits in the layers where hardware differences create performance: data representation, compiler behavior, kernel schedules, communication topology, and production validation. DeepSeek and Huawei can pay part of that budget through deep model, kernel, and hardware collaboration. Other adopters receive the resulting infrastructure, then qualify it against their own operating envelope.
This is stronger than a compatibility slogan. It is a map of what can be reused and what still deserves engineering time.
FAQ
Does DeepGEMM-Ascend run existing DeepGEMM calls unchanged?
It preserves the package name and published kernel API surface for supported operations. Efficient inputs still follow Ascend-specific scaling layouts, alignments, toolchain requirements, and supported hardware conditions. Treat source compatibility and operating-envelope compatibility as separate checks.
Does API compatibility imply equal performance?
API compatibility reduces application changes. Performance depends on shape, dtype, layout, compiler, kernel, firmware, topology, and workload. Reproduce both correctness and performance on the target stack.
Can these ports run on older Ascend hardware?
The pinned DeepGEMM-Ascend and FlashMLA documentation targets Ascend 950. DeepEP-Ascend says its published measurements establish no result for other Ascend generations or CANN versions. Older hardware needs its own support and validation matrix.
Do the published benchmarks prove that Ascend replaces CUDA?
They prove efficient implementations in specific test cells. A replacement decision also requires feature coverage, end-to-end workload results, quality, reliability, operator effort, upgrade behavior, and TCO.
What is the minimum useful proof of concept?
Choose one representative workload, freeze its full stack, pass cross-backend correctness, measure end-to-end service objectives, and complete one failure and upgrade drill. That small contract reveals more than a broad feature checklist.
References
- DeepSeek: DeepGEMM
- DeepSeek: DeepGEMM-Ascend
- DeepSeek: FlashMLA
- DeepSeek: DeepSelect
- DeepSeek: DeepEP
- DeepSeek: DeepEP-Ascend
- vLLM Ascend: Release notes
- vLLM Ascend: Validated installation matrix
- Agent Skills Compatibility Across Four Coding Agents
- OpenAI Jalapeño Benchmark Audit
- How to Benchmark Qwen3.8-27B for Real Workloads
Next action: create a one-page compatibility ledger for the target workload. Mark every required API, dtype, shape, collective, compiler version, and failure mode with its evidence level before buying capacity or scheduling a migration.