AI Accelerator IP
Configurable CNN accelerator for edge AI applications with support for popular neural network architectures.
Product Overview
Description
The Gowin AI Accelerator IP provides efficient hardware acceleration for convolutional neural networks on FPGA.
Supports INT8 quantization for efficient inference with minimal accuracy loss.
Optimized for edge applications including image classification, object detection, and keyword spotting.
Product Series
AI
Primary Application
Image classification
Key Features
- Configurable systolic array architecture
- INT8 quantization support
- On-chip weight and activation memory
- AXI4 interface for integration
- Model compiler from TensorFlow/PyTorch
- Batch processing support
- Low-power design
- Real-time inference capability
Specifications
| Voltage Rating | 25V DC |
|---|---|
| Current Rating | 1A |
| Temperature Range | -40°C to +85°C |
| Package | SOP-8 |
| Architecture | Systolic array CNN accelerator |
| MAC Units | 64-1024 configurable |
| Precision | INT8 (weights/activations) |
| On-chip Memory | 64-512KB |
| Throughput | Up to 100 GOPS |
| Latency | <10ms typical |
| Power | <1W typical |
| Supported Networks | MobileNet, ResNet, Custom |
Applications
Image classification
Electronic system design
Object detection
Electronic system design
Keyword spotting
Electronic system design
Anomaly detection
Electronic system design
Smart camera applications
Electronic system design
FAE Expert Insights
"Based on extensive field experience, this product delivers excellent performance across various operating conditions. The design incorporates proven architecture with robust protection features. Customers consistently report high satisfaction with reliability and ease of integration."
Excellent choice for industrial applications
— Michael Chen, BeiLuo
Frequently Asked Questions
What neural networks are supported?
The AI Accelerator supports popular CNN architectures including MobileNet V1/V2/V3 (optimized for edge), ResNet-18/34/50, and custom networks. The model compiler converts TensorFlow Lite and PyTorch models to the accelerator's format. Supported layers include Conv2D, DepthwiseConv, Fully Connected, ReLU, MaxPool, and AvgPool.
Verify your model architecture against supported layers. Contact FAE for model optimization assistance.
What performance can I expect?
Performance depends on FPGA device and configuration: GW2A-LV18 with 64 MACs achieves 10-20 GOPS; GW2A-LV55 with 256 MACs achieves 50-100 GOPS. Typical inference times: MobileNet on 224x224 image < 10ms; ResNet-18 < 20ms. Throughput scales with MAC count and clock frequency. Contact FAE for performance estimation based on your specific model.
Select MAC count based on latency requirements. More MACs = faster inference but higher resource usage.