Filtered by: Network × Clear all

Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference Accelerators

Afzal Ahmad, Gaoyu Mao, Shoubo Hu, Hui-Ling Zhen 2026-07-27

Multi-task inference models waste computation by executing identical operations regardless of the active task. We propose a HW/SW co-designed approach where a lightweight gating network predicts per-tile binary execution masks conditioned on the task command, enabling zero-overhead skipping of masked tiles. On a closed-loop visuomotor driving task in CARLA, task-conditional sparsity reduces FLOPs by 66-76%, on-device latency by 51-59% (2.1-2.4x speedup), and energy per inference from 263 to 108-128mJ. This work demonstrates that leveraging the task command as a free signal can significantly improve efficiency of multi-task inference accelerators without altering model architecture or inference pipeline.

PDF