Skip to content

Adversarial AI & Security Attacks

← Back to Domain-Specific Families

Abstractions about how adversaries exploit machine-learning systems and security boundaries, including attacks on training data and models (data poisoning attack, model theft, transfer-learning attack), attacks that manipulate deployed model inputs or outputs (evasion attack, prompt injection, model inversion attack), and foundational security design principles (Kerckhoffs's principle, blue-on-blue risk).

14 abstractions in this family — domain-specific abstractions that sit near one another in structural-signature space (k-means over structural-signature embeddings). Each is shown with its short description.

  • AI Supply-Chain Attack — The compromise of an upstream AI-pipeline component — training data, weights, packages, eval sets — by an adversary who exploits the deployer's trust in the producer's channel rather than breaching the perimeter, so the poisoned artefact is imported voluntarily.
  • Blue-on-Blue Risk — The risk that a force's own friend-or-foe classification mechanism misfires and its engagement rule fires on a friendly — harm that is symmetric and non-recoverable, so the only leverage is preventive, at the identification layer, never at the trigger-puller.
  • Bounded Storage Model — A cryptographic security model that uses temporary public randomness and a bound on adversarial retained information, rather than computational limits, to evaluate specified protocol goals.
  • Cyberattack — Conduct a deliberate hostile action through or against digital systems to gain unauthorized access, steal or manipulate information, disrupt availability, or maliciously control computing resources.
  • Data Extraction Through Prompting — The AI-security failure mode in which an adversary crafts inputs that make a deployed language model emit private content — training text, system prompts, retrieved documents — by exploiting the legitimate inference interface, whose breadth exceeds the access-control envelope around the data.
  • Data Poisoning Attack — The adversarial-ML failure mode where an attacker corrupts training inputs so the model learns an attacker-chosen behaviour — an upstream contamination that runtime defenses are structurally powerless against and that clean benchmark accuracy cannot detect.
  • Evasion Attack — Craft inputs that fall on the benign side of a deployed classifier's learned decision boundary while preserving their malicious payload — exploiting the gap between the boundary and the true concept it was meant to represent, so near-perfect test accuracy coexists with near-zero robustness under attack.
  • Input Manipulation Attack — Recognize an adversary crafting inputs to a fixed decision system so it produces a preferred output, by taking as the unit of analysis the gap between the input manifold the system was specified against and the manifold the attacker can construct.
  • Kerckhoffs's principle — Design a cryptosystem to stay secure even when the entire algorithm is public, confining secrecy to the key alone — because a leaked key is cheaply rotated while a leaked algorithm cannot be replaced without rebuilding the whole system.
  • Model Inversion Attack — Recover the content of individual training records by treating a deployed model as a constrained inverse oracle — combining predictions-as-evidence with a prior over the input space to solve for the generating input, so 'the model is not the data' becomes a quantitative claim.
  • Model Skewing — The adversarial AI-security attack in which a stream of individually-unflaggable, region-biased inputs to an online learning loop gradually pulls a model's decision boundary off the deployer's criterion toward one the attacker prefers — capturing a deployment that launched clean.
  • Model Theft — Clone a deployed model's valuable function from its outputs alone by treating it as a high-bandwidth oracle — querying it, harvesting (input, prediction) pairs, and fitting a substitute — so the IP boundary becomes a price set by extraction cost, not a wall.
  • Prompt Injection — Untrusted content delivered to a language model through a data channel is interpreted as instructions rather than material to process, so an embedded directive executes at the model's privilege — the failure being the absence of any in-band boundary between instructions and data, not a lapse in alignment.
  • Transfer-Learning Attack — Vulnerabilities, backdoors, or poisoned representations baked into an upstream pretrained model ride intact into every downstream system built on it, because the downstream team's audit boundary encloses only its new layers while the attack surface spans the whole inherited substrate.