Filtered by: Architecture × Clear all

Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing

Sadegh Majidi, Niloofar Mireshghallah, Kazem Taram 2026-07-27

LeakyLMs introduces a set of attacks that leak proprietary model architecture and inference optimization details from production language models using only per-token generation timing. The method builds a detailed timing model for NVIDIA GPUs and performs a search over the architecture space to recover properties like transformer layers, hidden dimension size, and attention heads. Experimental evidence shows that for Llama models, the near-correct architectural configuration appears in the top-10 guesses over 90% of the time, and the attack successfully detects speculative decoding in Google Gemini Flash 2.5. This matters because it demonstrates that sensitive model and deployment information can be inferred remotely via timing side channels, posing a significant security risk to proprietary language model services.

PDF

Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference Accelerators

Afzal Ahmad, Gaoyu Mao, Shoubo Hu, Hui-Ling Zhen 2026-07-27

Multi-task inference models waste computation by executing identical operations regardless of the active task. We propose a HW/SW co-designed approach where a lightweight gating network predicts per-tile binary execution masks conditioned on the task command, enabling zero-overhead skipping of masked tiles. On a closed-loop visuomotor driving task in CARLA, task-conditional sparsity reduces FLOPs by 66-76%, on-device latency by 51-59% (2.1-2.4x speedup), and energy per inference from 263 to 108-128mJ. This work demonstrates that leveraging the task command as a free signal can significantly improve efficiency of multi-task inference accelerators without altering model architecture or inference pipeline.

PDF