AI Infrastructure Solution

Application

Description

Comprehensive AI infrastructure solution combining Ascend AI accelerators with Kunpeng server processors for scalable AI training and inference deployments. Optimized for data centers requiring high-performance AI computing with excellent power efficiency.

Core Advantages

Industry-Leading AI Performance Ascend 910 delivers 256 TFLOPS FP16 performance with excellent power efficiency, enabling faster training of large AI models while reducing data center power consumption.
Unified Software Stack CANN software architecture provides seamless support for TensorFlow, PyTorch, and MindSpore, enabling easy migration of existing AI workloads.
Scalable Architecture Support for clusters up to 4096 cards with high-speed 100Gbps interconnect enables training of the largest AI models with linear scaling efficiency.
Edge-to-Cloud Continuity Same Da Vinci architecture from edge (Ascend 310) to data center (Ascend 910) ensures model compatibility and simplified deployment.
Advantage 5 This solution provides industry-leading performance and reliability for demanding applications.

Recommended Bill of Materials (BOM)

Item Part Number Description Quantity Datasheet
1 Ascend 910 AI Training Processor 8 📄 Download
2 Kunpeng 920-6426 64-Core Server Processor 2 📄 Download
3 DDR4-2933 128GB Server Memory 16 📄 Download
4 CANN 6.0 AI Software Stack 1 📄 Download

Applications

Large language model training
Computer vision at scale
Recommendation systems
Scientific computing

Technical Specifications

Training Performance
256 TFLOPS FP16 per Ascend 910
Inference Performance
512 TOPS INT8 per Ascend 910
Edge Inference
16 TOPS INT8 per Ascend 310
Server Compute
64 cores @ 2.6GHz per Kunpeng 920
Memory Capacity
32GB HBM2 + 4TB DDR4 per node
Network
100Gbps RoCE v2 interconnect
Power
310W per AI card, 180W per CPU
Form Factor
PCIe cards + 2U servers

Customer Success Stories

Cloud AI Provider

Cloud Computing | Large Language Model Training

Challenge

Training billion-parameter language models required massive compute infrastructure with high power consumption and cooling costs using existing GPU clusters.

Solution

Deployed 512-node Ascend 910 cluster with Kunpeng 920 servers, utilizing high-speed RoCE interconnect and optimized CANN software stack for distributed training.

Results

Achieved 25% better training throughput-per-watt compared to previous GPU infrastructure, reduced training time for 175B parameter model from 3 weeks to 10 days, and lowered data center cooling costs by 30%.

Smart City Initiative

Government | Video Analytics Platform

Challenge

Processing video feeds from 10,000+ cameras in real-time for traffic monitoring and public safety required massive inference capacity with low latency.

Solution

Implemented distributed AI infrastructure with Ascend 310 edge nodes for local processing and Ascend 910 cluster for centralized model training and complex analytics.

Results

System processes 50,000 video streams simultaneously with <50ms latency, achieved 40% cost reduction compared to GPU-based solution, and reduced false positive alerts by 60% through improved model accuracy.

FAE Expert Insights

D

Dr. Michael Zhang

Principal Solutions Architect - AI Infrastructure

18 years

Professional Insights

Key considerations: Ascend 910 delivers competitive training performance with 20-30% better power efficiency than GPUs; Unified Da Vinci architecture simplifies edge-to-cloud model deployment; CANN software stack enables seamless migration from TensorFlow/PyTorch; Kunpeng servers provide cost-effective foundation for AI infrastructure; Scalable to thousands of cards with linear performance scaling. Common pitfalls to avoid: Underestimating interconnect bandwidth requirements for large clusters; Not optimizing data pipeline which can become bottleneck before AI accelerators; Ignoring software migration effort - plan for 2-4 weeks optimization; Inadequate cooling design for dense AI server deployments.

Key Takeaways

  • Ascend 910 delivers competitive training performance with 20-30% better power efficiency than GPUs
  • Unified Da Vinci architecture simplifies edge-to-cloud model deployment
  • CANN software stack enables seamless migration from TensorFlow/PyTorch
  • Kunpeng servers provide cost-effective foundation for AI infrastructure
  • Scalable to thousands of cards with linear performance scaling

Decision Framework

Decision Framework
Steps:
  1. Evaluate requirements
  2. Compare solutions
  3. Consult FAE

Ready to Implement This Solution?

Contact our FAE team for design support and quotes

Contact Us Now

Frequently Asked Questions

What is the recommended cluster size for training large language models?

For large language models (100B+ parameters), we recommend starting with 64-128 Ascend 910 cards, which provides sufficient memory aggregate (2-4TB) and compute for efficient training. Models up to 175B parameters can be trained effectively on 256 cards. For very large models (500B+), scale to 512-1024 cards. The key consideration is not just compute but memory capacity - each card provides 32GB HBM2, and large models require significant memory for parameters, optimizer states, and activations. Our FAE team can help calculate exact requirements based on your model architecture.

Contact our solutions team with your model specifications for cluster sizing recommendations and performance projections.

How does the Ascend solution compare to NVIDIA DGX systems?

Ascend-based solutions offer several advantages over DGX: Cost - typically 30-40% lower acquisition cost for equivalent performance

Power efficiency - 20-30% better TFLOPS-per-watt reduces operating costs

Flexibility - mix Ascend 910 (training) and 310 (inference) based on workload

Software - CANN provides competitive performance with broader framework support than early versions. DGX advantages include mature ecosystem, broader ISV support, and NVLink for fast GPU-to-GPU communication. For organizations not locked into CUDA, Ascend offers compelling TCO advantages. We recommend benchmarking your specific workloads on both platforms.

Evaluate both platforms with your actual workloads. Ascend typically wins on TCO for large-scale deployments.

What storage configuration is recommended for AI training clusters?

AI training requires high-throughput storage to feed data to accelerators: Local NVMe - 8-16TB per server for dataset caching

Parallel file system - Lustre, BeeGFS, or Huawei OceanStor for shared storage

Throughput - plan for 10-20GB/s aggregate for large clusters

Capacity - 100TB to 1PB+ depending on dataset size. The storage architecture should use tiering: hot data (active training sets) on NVMe, warm data on SSD, cold data on HDD. Network storage should use high-speed interconnect (100Gbps+) to prevent I/O bottlenecks. Our solutions include storage sizing tools based on your dataset characteristics.

Size storage for 10-20GB/s throughput per 64-card cluster segment. Contact us for storage architecture design.

How long does it take to migrate models from GPUs to Ascend?

Migration time depends on model complexity: Simple models (ResNet, BERT) - 1-2 days using automatic conversion tools

Complex models with custom ops - 1-2 weeks including operator development

Large models requiring optimization - 2-4 weeks for performance tuning. The CANN toolkit provides: Graph conversion from TensorFlow/PyTorch ONNX

Automatic operator mapping

Profiling tools to identify bottlenecks

Optimization recommendations. Most standard models convert with minimal changes. Custom operators may need reimplementation using CANN APIs. Our FAE team provides migration support and can handle complex conversions.

Plan for 2-4 weeks for initial migration and optimization. Contact our FAE team for migration assistance.

What cooling solutions are available for dense Ascend deployments?

Ascend 910 servers support multiple cooling options: Air cooling - standard for up to 4 cards per server (1.2kW per node)

Enhanced air cooling - up to 8 cards with optimized airflow (2.5kW per node)

Liquid cooling - direct-to-chip for highest density (40-50kW per rack)

Immersion cooling - for extreme density deployments. For a typical deployment: 2U server with 8 Ascend 910 cards requires 2.5kW and standard data center cooling

42U rack with 16 servers (128 cards) requires 40kW cooling capacity. We recommend liquid cooling for new data center builds as it enables highest density and lowest PUE.

Choose air cooling for <30kW per rack, liquid cooling for >30kW or new data center builds.

What network topology is recommended for large Ascend clusters?

For Ascend clusters, we recommend: Small clusters (16-64 cards) - single switch with 100Gbps ports

Medium clusters (128-512 cards) - fat-tree topology with spine-leaf architecture

Large clusters (1024+ cards) - dragonfly or enhanced fat-tree for optimal bisection bandwidth. Key considerations: Bisection bandwidth - maintain 1:1 or better for all-to-all communication

Latency - <2us for optimal training performance

Redundancy - dual paths for high availability. The RoCE v2 protocol provides RDMA capabilities for efficient collective operations. Our network design guides provide switch configurations and cabling diagrams for various cluster sizes.

Use fat-tree topology for clusters up to 512 cards, dragonfly for larger deployments. Contact us for network design assistance.