Ascend 910

✓ In Stock

256 TFLOPS FP16 data center AI processor with 310W power for large-scale AI training and high-throughput inference

Product Overview

Description

The Ascend 910 is HiSilicon's flagship AI training solution, delivering exceptional performance for data center AI workloads. With massive parallel processing capabilities, it accelerates training of large neural networks.

Product Series

Ascend

Primary Application

AI model training

Key Features

  • 256 TFLOPS FP16 training performance
  • 32 Da Vinci AI cores for massive parallelism
  • High-bandwidth HBM2 memory
  • PCIe 4.0 and 100Gbps interconnect
  • Scalable to thousands of cards
  • CANN software stack support

Specifications

AI Performance 256 TFLOPS FP16 / 512 TOPS INT8
Architecture 32-core Da Vinci
Memory 32GB HBM2, 1.2 TB/s
Power 310W typical
Process 7nm
Interconnect RoCE v2, 100Gbps
Interfaces PCIe 4.0 x16
Scaling Cluster up to 4096 cards
Temperature 5°C to +45°C
Form Factor Dual-slot FHFL card

Applications

AI model training

Electronic system design

Large language models

Electronic system design

High-throughput inference

Electronic system design

Scientific computing

Electronic system design

Documents & Resources

FAE Expert Insights

R

"The Ascend 910 is a serious contender in the data center AI training market. In my experience deploying large-scale training clusters, the 256 TFLOPS FP16 performance is competitive with A100, and the HBM2 bandwidth eliminates the memory bottlenecks we often see with GDDR-based solutions. The cluster scaling works well - I've supported deployments up to 1024 cards with good linear scaling efficiency. The CANN stack has improved dramatically; TensorFlow model migration is now straightforward with the conversion tools. For customers building AI infrastructure, the TCO advantage over NVIDIA solutions can be significant. I recommend starting with a pilot cluster to validate your specific workloads."

Competitive training performance with excellent scaling for large AI clusters

— Robert Zhang, BeiLuo

Frequently Asked Questions

What is the training throughput of Ascend 910 for popular models?

Ascend 910 delivers impressive training throughput: ResNet-50 training achieves ~1800 images/second per card; BERT-Large pre-training processes ~100 sequences/second; GPT-3 style models train at competitive speeds with good scaling efficiency. The HBM2 memory bandwidth (1.2 TB/s) is crucial for large model training, eliminating the memory bottlenecks common with GDDR-based accelerators. When scaled to clusters, Ascend 910 maintains good linear scaling efficiency up to hundreds of cards with proper interconnect configuration. Actual throughput depends on model architecture, batch size, and data pipeline optimization.

Benchmark your specific model on Ascend 910 to verify performance meets your training time requirements.

training throughput ResNet-50 BERT scaling efficiency
How does Ascend 910 compare to NVIDIA A100 for training?

Ascend 910 and A100 are both high-performance AI training accelerators with different strengths: Ascend 910 offers 256 TFLOPS FP16 vs A100's 312 TFLOPS, making A100 about 22% faster in raw compute. However, Ascend 910 has better power efficiency (0.83 vs 0.78 TFLOPS/W) and lower power consumption (310W vs 400W). A100 has advantages in ecosystem maturity (CUDA vs CANN) and memory capacity (80GB vs 32GB). For large models that don't fit in 32GB, A100's larger memory is significant. For models that fit, Ascend 910 offers competitive performance with better TCO. Both support similar precision formats (FP16, BF16, INT8).

Choose Ascend 910 for cost-optimized training of models fitting in 32GB; choose A100 for maximum ecosystem compatibility or very large models.

A100 comparison training performance TCO
What is the scaling efficiency for Ascend 910 clusters?

Ascend 910 clusters achieve excellent scaling efficiency: 16-card clusters typically achieve 90-95% linear scaling; 128-card clusters maintain 85-90% efficiency; 1024-card clusters achieve 80-85% efficiency with proper configuration. The 100Gbps RoCE interconnect provides high-bandwidth, low-latency communication between cards. The CANN software stack includes optimized collective communication libraries (similar to NCCL) for efficient distributed training. For maximum scaling efficiency, use the recommended network topology (fat-tree or dragonfly) and ensure proper data pipeline optimization to keep GPUs/VPUs fed with data.

Design your cluster with proper interconnect topology and work with our solutions team for large-scale deployment optimization.

scaling efficiency distributed training cluster configuration
Does Ascend 910 support inference as well as training?

Yes, Ascend 910 excels at both training and high-throughput inference: For training, the 256 TFLOPS FP16 performance and HBM2 bandwidth accelerate large model training. For inference, the 512 TOPS INT8 performance enables extremely high throughput - up to 100,000+ images/second for ResNet-50. Many customers use Ascend 910 for both training and inference deployment, simplifying their AI infrastructure. The high memory bandwidth is particularly beneficial for large model inference (like GPT-3 style models) where weight loading can be a bottleneck. For pure inference deployments, consider Ascend 310 for edge or cost-optimized data center inference.

Use Ascend 910 for training+inference combined workloads or maximum inference throughput; use Ascend 310 for cost-optimized inference.

inference throughput training and inference
What is the typical power and cooling requirement for Ascend 910 servers?

Ascend 910 servers require significant power and cooling infrastructure: Each Ascend 910 card consumes 310W, so an 8-card server requires ~2500W plus CPU and system power (total ~3000W). For a 42U rack with 8 servers, plan for 25-30kW power capacity. Cooling requirements depend on deployment density: Air-cooled servers need 150-200 CFM per card; Liquid cooling enables higher density with 40-50kW per rack. Data center ambient temperature should be maintained at 18-27°C for optimal performance. The power supplies are typically 80 PLUS Platinum or Titanium for efficiency. Work with our solutions team for data center design guidance.

Ensure your data center can support the power and cooling requirements before deploying Ascend 910 clusters.

power consumption cooling data center requirements