#llm-security
-
Best AI Red Teaming Tools for ML Models in 2026
A practitioner's guide to the tools used to red team machine learning models: PyRIT, garak, ART, ModelScan, and Giskard, and where each fits a pipeline.
-
Adversarial Attacks on Vision-Language Models: CLIP, LLaVA, GPT-4
Vision-language models widen the adversarial attack surface: crafted images can steer text output, carry typographic payloads, and jailbreak the model.
-
Training Data Extraction from LLMs: The Carlini Results Explained
Carlini et al. demonstrated verbatim extraction of training data from GPT-2. The results have been widely misread. What the extraction numbers actually show.