HOW TO BUILD RESILIENT AI MODELS?

Adversarial attacks on AI models

Four threat categories and effective defense strategies

This article opens a series dedicated to AI security and adversarial machine learning. It is based on the findings of an MBA thesis in Cybersecurity completed at the Military University of Technology and presents a concise overview of key conclusions and practical recommendations. 

Now that we know AI models can be deceived, the next engineering question is: how exactly can they be defended, and at what cost? In this article, I break adversarial attacks down into four categories, explain why the traditional “gold standard” of defense, adversarial training, is not always cost-effective, and introduce the concept of lightweight defense through data augmentation. You’ll gain a practical threat landscape overview and a decision-making framework for determining when robust defenses are worth the investment and when they are not. This is a challenge increasingly faced by organizations deploying AI systems in business processes.

What Are the Main Types of Adversarial Attacks?

Before designing a defense strategy, we need to understand what we are defending against. In both the literature and my research, adversarial attacks are categorized along three dimensions: the attacker’s level of knowledge, the attack phase, and the attack objective. This taxonomy directly influences which protection and control mechanisms can be implemented.

Adversarial Attack Taxonomy

Based on the Attacker’s Knowledge

White-box
  • Full access to the model
  • Knowledge of the architecture and model weights
  • Examples: FGSM, PGD, CW
Black-box
  • Access only through queries
  • No knowledge of the model’s internal structure
  • Examples: Square Attack, transfer attacks

Based on the Attack Phase

Evasion Attacks
  • Performed during the inference phase
Poisoning Attacks
  • Performed during the training phase
  • Examples: backdoor attacks

Based on the Attack Objective

Targeted Attacks
  • Force the model to classify an input as a specific target class
Untargeted Attacks
  • Cause any form of misclassification

The most challenging scenario is a white-box attack, where the adversary knows the model architecture and parameters and can therefore exploit its gradients. This is the case I investigated experimentally, because if a defense mechanism can withstand white-box attacks, it should also be effective against less capable adversaries.

How Do Attacks Against AI Models Work?

Gradient-based attacks exploit the same mathematical principles used to train neural networks, but in reverse. Rather than modifying model parameters to minimize error, the attacker modifies the input data to maximize it.

FGSM: The Fastest Gradient-Based Attack

FGSM (Fast Gradient Sign Method) performs a single step in the direction of the gradient. It is computationally efficient but relatively simple.

PGD and BIM: The Benchmark for Robustness Evaluation

BIM and PGD are iterative extensions of FGSM. Today, PGD is commonly regarded as the practical benchmark for evaluating model robustness.

Carlini-Wagner (CW): The Ultimate Stress Test

Carlini-Wagner (CW) is an optimization-based attack that searches for the smallest possible perturbation capable of deceiving a model. While slower than other methods, it is exceptionally effective.

Square Attack: A Gradient-Free Approach

Square Attack belongs to the black-box category and operates exclusively through model queries, without requiring access to gradients.

In my experiments, the CW attack proved to be the most effective adversarial method and will therefore serve as the primary reference point throughout the remainder of this series.

Is Adversarial Training Worth the Cost?

The dominant defense strategy remains Adversarial Training, formulated as a Min-Max optimization problem. During training, worst-case perturbations are generated, and the model learns to become robust against them.

While effective, this approach comes with significant computational overhead.

Adversarial Training as the Gold Standard

Adversarial Training (Min-Max)
  • High computational cost
  • Attack generation during every training iteration
  • Significantly increased resource requirements
Data Augmentation

(Jitter, RandomZero, GaussianNoise)

  • Low computational cost
  • Single training pass
  • Additional layer of randomness introduced at the input level

Computational Cost vs. Practical Effectiveness

The challenge with Min-Max training is twofold: it consumes substantial computational resources and often struggles against attacks that were not encountered during training. In environments with limited budgets and relatively small datasets, this creates a genuine barrier to adoption.

How Can Data Augmentation Improve AI Robustness?

An alternative approach relies on a simple idea: if a white-box attack closely follows model gradients, we can disrupt that process by introducing controlled randomness into the input. This strategy can improve AI model robustness without incurring the costs associated with adversarial training.

What Does Lightweight Defense Look Like?

In my research, I evaluated seven techniques based on data augmentation and controlled perturbation of input data.

Jitter, RandomZero, and Other Augmentation Techniques

  • Jitter: Adds random noise to selected signal points.
  • RandomZero: Randomly sets portions of the input to zero.
  • GaussianNoise: Adds normally distributed noise.
  • SmoothTS and ShuffleDef: Alternative approaches with varying levels of aggressiveness and impact on input data.

Each of these methods acts as a lightweight pre-processing layer and introduces virtually negligible computational overhead compared to the Min-Max approach.

How Should AI Security Be Tested and Common Mistakes Avoided?

Understanding attack techniques alone is not enough. Effective AI security requires continuous robustness testing and alignment with recognized standards and frameworks.

What Do NIST AI RMF and the AI Act Recommend?

Within the NIST AI Risk Management Framework (AI RMF), subcategory MEASURE 2.7 recommends red teaming and robustness assessments against threats such as adversarial examples and data poisoning.

Meanwhile, the publication NIST AI 100-2: Adversarial Machine Learning: A Taxonomy and Terminology provides a common vocabulary that organizations can adopt to ensure technical teams and risk managers speak the same language.

From a regulatory perspective, Article 15(3) of the EU AI Act requires AI systems to be resilient against errors and inconsistencies, while Article 15(5) explicitly addresses threats related to model evasion attacks.

In practice, this means model robustness is becoming an integral component of both AI security and regulatory compliance.

Three Common Mistakes Organizations Make

1. Evaluating robustness with only a weak attack

If you test exclusively against FGSM, you may develop a false sense of security. Robustness assessments should also include PGD and CW attacks.

2. Treating Min-Max training as the only viable solution

For many use cases, lightweight data augmentation techniques can deliver comparable improvements at a fraction of the cost.

3. Ignoring black-box and poisoning attacks

A defense strategy designed solely for white-box threats leaves other attack vectors exposed.

How Should Organizations Choose a Defense Strategy?

Approach defense selection as an engineering and economic decision. The question is not “Which method is the best?” but rather: “Which method provides the required level of robustness at an acceptable cost without significantly degrading baseline model performance?”

What Do These Threats Mean for Organizations?

Adversarial attacks form a well-defined taxonomy, with white-box attacks leveraging model gradients representing the most demanding threat scenario.

Traditional Min-Max adversarial training remains highly effective, but it is not always economically justified. As a result, lightweight techniques based on data augmentation are gaining relevance, aligning well with both the principles of the AI Act and the recommendations of the NIST AI RMF.

For organizations seeking to improve AI resilience without dramatically increasing AI cybersecurity costs, these approaches represent a compelling alternative.


In Part 3 of this series, I will present the results of my own experiments, where a single, straightforward technique reduced the effectiveness of the most powerful attack from −26.5 percentage points to just −0.9 percentage points. There was, however, one critical condition. And I’ll reveal exactly what it was.

AUTHOR:
Piotr Hawryło is a Software Team Leader at ALTEN with more than 10 years of experience in embedded systems development. He holds an MBA in Cybersecurity from the Military University of Technology and has participated in machine learning projects, including initiatives related to exoplanet detection. His areas of expertise include software engineering, technical team leadership, and emerging technologies.