SwiftNPU: Scalable Shape-Flexible Allocation for Inter-Core Connected NPUs
communityMakes multi-tenant NPU sharing practical for Blackhole-class hardware using polynomial-time allocation algorithms. Delivers up to 1.37× higher utilization and 1.14× faster workload completion. Up to 890,000× faster than NP-hard baselines.
Links
📄
ACM DL
Works on
blackhole