As machine learning (ML) becomes increasingly embedded in hardware, especially through specialized accelerators, new security risks are emerging. Modern ML accelerators rely on complex, globally distributed supply chains, creating opportunities for attackers to insert malicious modifications—known as hardware Trojans (HTs)—into the design or manufacturing process. 

Traditionally, these attacks have focused on obvious disruptions such as backdoors or system failures, which are often detected during testing. But there is a more insidious class of attacks on the horizon: gradual accuracy-degrading attacks. Instead of causing immediate and noticeable failures, these attacks slowly corrupt the model’s performance over time. Because the system continues to function normally, detecting the issue becomes significantly more difficult.

In a paper presented at the International Symposium on Quality Electronic Design, researchers propose a novel sensitivity analysis method based on partial derivatives, which quantitatively assesses the security implications of weights in a hardware-implemented ML model. This analysis not only identifies the most critical parameters to target but also provides a foundation for protecting ML accelerators. 

The Growing Threat Landscape

Existing works have demonstrated various attacks targeting components of ML accelerators. That includes components such as the model’s neurons, memory interfaces, pooling layers, and the reconfigurable interconnects within the ML accelerator. These attacks typically aim to cause specific misclassification decisions (a backdoor) or a denial-of-service attack. While attacks on weights have been considered, they often focus on extracting embedded information rather than degrading performance.

The researchers anticipate that, unlike traditional service-denial attacks and backdoors that exhibit sharp, easily recognizable behavior when triggered, the next generation of ML accelerator attacks will aim to subtly corrupt the model weights over time, gradually eroding its overall performance. Because the accelerator appears to function correctly (the attack is either inactive or it imposes a tiny accuracy drift), these attacks are exceptionally difficult to detect. 

In this regard, a key research gap remains: the absence of a systematic methodology to identify which weights are the most valuable targets for adversaries, enabling them to induce significant accuracy loss through minimal, and therefore stealthy, modifications. To address this gap, the researchers adapt partial derivative (PD) theory to rank all model weights based on their impact on the loss function. This ranking provides a quantitative, systematic approach to identifying the most influential targets.

Proposed Sensitivity-Based Hardware Trojans

The authors propose a sensitivity analysis method based on partial derivatives. This technique ranks model weights by their impact on overall performance, helping identify which parameters are most critical, where small changes can cause large drops in accuracy, and pinpointing high-value targets within an ML model.

Using this method, the researchers designed six different hardware Trojans and tested them on a LeNet-5 neural network implemented on an FPGA.

Block diagram of the proposed weight perturbation Trojans.

 

The tests found that small, targeted changes in ML hardware can lead to major performance losses—without triggering alarms. The tests also demonstrated the framework’s effectiveness by designing and implementing a novel class of hardware Trojans, called accuracy-degrading Trojans.

Key Takeaways

To address emerging attacks that quietly degrade ML performance by tampering with model weights over time, researchers developed a new framework to identify which weights are most vulnerable to hardware-level manipulation. 

The findings highlight a critical shift in ML security:

  • Tiny modifications to key weights can significantly degrade accuracy
  • The hardware cost of inserting these attacks is minimal
  • Standard verification methods are unlikely to detect attacks
  • Future attacks may prioritize stealth over disruption
  • Hardware-level vulnerabilities can have long-term, hidden impacts
  • Traditional testing approaches may not be sufficient

This framework fills a key gap in securing machine-learning accelerators against subtle, hard-to-detect threats. Importantly, this sensitivity-based approach can also be used defensively, helping designers protect critical components and strengthen system resilience. As ML accelerators become more widely used, tackling these hidden vulnerabilities will be essential to ensuring secure and reliable operation.

Interested in learning more about Machine Learning and Security Vulnerabilities? We have thousands of articles related to these industries and more! Also learn more about our eBooks and eLearning collections on cutting edge technologies.

Interested in acquiring full-text access to this collection for your entire organization? Request a free demo and trial subscription for your organization.