From 61b9c91f62bdafd8ff501bc55c2090d49aef2148 Mon Sep 17 00:00:00 2001 From: promptadmin Date: Fri, 26 Jun 2026 17:24:07 +0000 Subject: [PATCH 1/3] [upstream-sync] README.md from trailofbits/awesome-ml-security@d7fe4fa9 [catalogue] --- .../catalogue/README.md | 228 ++++++++++++++++++ 1 file changed, 228 insertions(+) create mode 100644 upstream/trailofbits-awesome-ml-security/catalogue/README.md diff --git a/upstream/trailofbits-awesome-ml-security/catalogue/README.md b/upstream/trailofbits-awesome-ml-security/catalogue/README.md new file mode 100644 index 0000000..27ece9a --- /dev/null +++ b/upstream/trailofbits-awesome-ml-security/catalogue/README.md @@ -0,0 +1,228 @@ +--- +title: "Awesome-ML-Security" +task: "" +lineage_type: import +upstream_source: https://github.com/trailofbits/awesome-ml-security/blob/d7fe4fa9/README.md +upstream_sha: d7fe4fa9 +imported_at: 2026-06-26 +prompt_class: catalogue +upstream_changes: accepted +author: upstream +validated: false +--- + +# Awesome-ML-Security + +A curated list of awesome machine learning security references, guidance, tools, and more. + +**Table of Contents** + +- [Awesome-ML-Security](#awesome-ml-security) + - [Relevant work, standards, literature](#relevant-work-standards-literature) + - [CIA of the model](#cia-of-the-model) + - [Confidentiality](#confidentiality) + - [Integrity](#integrity) + - [Availability](#availability) + - [Degraded model performance](#degraded-model-performance) + - [ML-Ops](#ml-ops) + - [AI’s effect on attacks/security elsewhere](#ais-effect-on-attackssecurity-elsewhere) + - [Self-driving cars](#self-driving-cars) + - [LLM Alignment](#llm-alignment) + - [Regulatory actions](#regulatory-actions) + - [US](#us) + - [EU](#eu) + - [Other](#other) + - [Safety standards](#safety-standards) + - [Taxonomies and frameworks](#taxonomies-and-frameworks) + - [Security tools and techniques](#security-tools-and-techniques) + - [API probing](#api-probing) + - [Model backdoors](#model-backdoors) + - [Other](#other-1) + - [Background information](#background-information) + - [DeepFakes, disinformation, and abuse](#deepfakes-disinformation-and-abuse) + - [Notable incidents](#notable-incidents) + - [Notable harms](#notable-harms) + +## Relevant work, standards, literature + +### CIA of the model +Membership attacks, model inversion attacks, model extraction, adversarial perturbation, prompt injections, etc. +* [Towards the Science of Security and Privacy in Machine Learning](https://arxiv.org/abs/1611.03814) +* [SoK: Machine Learning Governance](https://arxiv.org/abs/2109.10870) +* [Not with a Bug, But with a Sticker: Attacks on Machine Learning Systems and What To Do About Them](https://www.goodreads.com/book/show/125075266-not-with-a-bug-but-with-a-sticker) +* [On the Impossible Safety of Large AI Models](https://arxiv.org/abs/2209.15259) + +#### Confidentiality +Reconstruction (model inversion; attribute inference; gradient and information leakage), theft of data, Membership inference and reidentification of data, Model extraction (model theft), property inference (leakage of dataset properties), etc. +* [awesome-ml-privacy-attacks](https://github.com/stratosphereips/awesome-ml-privacy-attacks) +* [Privacy Side Channels in Machine Learning Systems](https://arxiv.org/abs/2309.05610#:~:text=Most%20current%20approaches%20for%20protecting,%2C%20output%20monitoring%2C%20and%20more) +* [Beyond Labeling Oracles: What does it mean to steal ML models?](https://arxiv.org/abs/2310.01959) +* [Text Embeddings Reveal (Almost) As Much As Text](https://arxiv.org/abs/2310.06816?ref=upstract.com) +* [Language Model Inversion](https://arxiv.org/abs/2311.13647) +* [Extracting Training Data from ChatGPT](https://not-just-memorization.github.io/extracting-training-data-from-chatgpt.html) +* [Recovering the Pre-Fine-Tuning Weights of Generative Models](https://arxiv.org/abs/2402.10208) + +#### Integrity +Backdoors/neural trojans (same as for non-ML systems), adversarial evasion (perturbation of an input to evade a certain classification or output), data poisoning and ordering (providing malicious data or changing the order of the data flow into an ML model). +* [A Systematic Survey of Backdoor Attack, Weight Attack and Adversarial Examples](https://arxiv.org/abs/2302.09457) +* [Poisoning Web-Scale Training Datasets is Practical](https://arxiv.org/abs/2302.10149) +* [Planting Undetectable Backdoors in Machine Learning Models](https://arxiv.org/abs/2204.06974) +* [Motivating the Rules of the Game for Adversarial Example Research](https://arxiv.org/abs/1807.06732) +* [On Evaluating Adversarial Robustness](https://arxiv.org/abs/1902.06705) +* [Tree of Attacks: Jailbreaking Black-Box LLMs Automatically](https://arxiv.org/abs/2312.02119) +* [Universal and Transferable Adversarial Attacks on Aligned Language Models](https://llm-attacks.org/) +* [Manipulating SGD with Data Ordering Attacks](https://arxiv.org/abs/2104.09667) +* [Adversarial reprogramming](https://arxiv.org/abs/1806.11146) - repurposing a model for a different task than its original intended purpose +* [Model spinning attacks](https://arxiv.org/abs/2107.10443) (meta backdoors) - forcing a model to produce output that adheres to a meta task (for ex. making a general LLM produce propaganda) +* [LLM Censorship: A Machine Learning Challenge or a Computer Security Problem?](https://arxiv.org/abs/2307.10719) +* [Securing LLM Systems Against Prompt Injection](https://developer.nvidia.com/blog/securing-llm-systems-against-prompt-injection/) & [Mitigating Stored Prompt Injection Attacks Against LLM Applications](https://developer.nvidia.com/blog/mitigating-stored-prompt-injection-attacks-against-llm-applications/) + * [Best Practices for Securing LLM-Enabled Applications](https://developer.nvidia.com/blog/best-practices-for-securing-llm-enabled-applications/) + * [NVIDIA NeMo Guardrails: Security Guidelines](https://docs.nvidia.com/nemo/guardrails/security/guidelines.html) +* [Multi-Agent Systems Execute Arbitrary Malicious Code](https://arxiv.org/abs/2503.12188) +* [Agentic Autonomy Levels and Security](https://developer.nvidia.com/blog/agentic-autonomy-levels-and-security/) +* [Rerouting LLM Routers](https://arxiv.org/abs/2501.01818) +* [Defeating Prompt Injections by Design](https://arxiv.org/abs/2503.18813) +* [Arcanum Prompt Injection Taxonomy](https://github.com/Arcanum-Sec/arc_pi_taxonomy) + + +#### Availability +* [Energy-latency attacks](https://arxiv.org/abs/2006.03463) - denial of service for neural networks + +### Degraded model performance +* [Trail of Bits's Audit of YOLOv7](https://blog.trailofbits.com/2023/11/15/assessing-the-security-posture-of-a-widely-used-vision-model-yolov7/) +* [Robustness Testing of Autonomy Software](https://users.ece.cmu.edu/~koopman/pubs/hutchison18_icse_robustness_testing_autonomy_software.pdf) +* [Can robot navigation bugs be found in simulation? An exploratory study](https://hal.science/hal-01534235/file/PID4832685.pdf) +* [Bugs can optimize for bad behavior (OpenAI GPT-2)](https://openai.com/research/fine-tuning-gpt-2) +* [You Only Look Once Run time errors](https://www.york.ac.uk/assuring-autonomy/guidance/body-of-knowledge/implementation/2-3/2-3-3/cross-domain-automotive/) + +### ML-Ops +* [Incubated ML Exploits: Backdooring ML Pipelines using Input-Handling Bugs](https://www.youtube.com/watch?v=Z38pTFM0FyU) +* [Auditing the Ask Astro LLM Q&A app](https://blog.trailofbits.com/2024/07/05/auditing-the-ask-astro-llm-qa-app/) +* [Exploiting ML models with pickle file attacks: Part 1](https://blog.trailofbits.com/2024/06/11/exploiting-ml-models-with-pickle-file-attacks-part-1/) & [Exploiting ML models with pickle file attacks: Part 2](https://blog.trailofbits.com/2024/06/11/exploiting-ml-models-with-pickle-file-attacks-part-2/) +* [PCC: Bold step forward, not without flaws](https://blog.trailofbits.com/2024/06/14/pcc-bold-step-forward-not-without-flaws/) +* [Trail of Bits's Audit of the Safetensors Library](https://github.com/trailofbits/publications/blob/master/reviews/2023-03-eleutherai-huggingface-safetensors-securityreview.pdf) +* [Facebook’s LLAMA being openly distributed via torrents](https://news.ycombinator.com/item?id=35007978) +* [Summoning Demons: The Pursuit of Exploitable Bugs in Machine Learning](https://arxiv.org/abs/1701.04739) +* [DeepPayload: Black-box Backdoor Attack on Deep Learning Models through Neural Payload Injection](https://arxiv.org/abs/2101.06896) +* [Weaponizing Machine Learning Models with Ransomware](https://hiddenlayer.com/research/weaponizing-machine-learning-models-with-ransomware/) (and [Machine Learning Threat Roundup](https://hiddenlayer.com/research/machine-learning-threat-roundup/)) +* [Bug Characterization in Machine Learning-based Systems](https://arxiv.org/abs/2307.14512) +* [LeftoverLocals: Listening to LLM responses through leaked GPU local memory](https://blog.trailofbits.com/2024/01/16/leftoverlocals-listening-to-llm-responses-through-leaked-gpu-local-memory/) +* [Offensive ML Playbook](https://wiki.offsecml.com/Welcome+to+the+Offensive+ML+Playbook) +* [MCP security briefing](https://www.wiz.io/blog/mcp-security-research-briefing) + + +### AI’s effect on attacks/security elsewhere +* [How AI will affect cybersecurity: What we told the CFTC](https://blog.trailofbits.com/2023/07/31/how-ai-will-affect-cybersecurity-what-we-told-the-cftc/) +* [Lost at C: A User Study on the Security Implications of Large Language Model Code Assistants](https://arxiv.org/abs/2208.09727) +* [Examining Zero-Shot Vulnerability Repair with Large Language Models](https://arxiv.org/pdf/2112.02125.pdf) +* [Do Users Write More Insecure Code with AI Assistants?](https://arxiv.org/pdf/2211.03622.pdf) +* [Learned Systems Security](https://arxiv.org/abs/2212.10318) +* [Beyond the Hype: A Real-World Evaluation of the Impact and Cost of Machine Learning-Based Malware Detection](https://arxiv.org/abs/2012.09214) +* [Data-Driven Offense](https://player.vimeo.com/video/133292422) from Infiltrate 2015 +* [Codex (and GPT-4) can’t beat humans on smart contract audits](https://blog.trailofbits.com/2023/03/22/codex-and-gpt4-cant-beat-humans-on-smart-contract-audits/) + +#### Self-driving cars +* [Driving to Safety: How Many Miles of Driving Would It Take to Demonstrate Autonomous Vehicle Reliability?](https://www.rand.org/pubs/research_reports/RR1478.html) + +#### LLM Alignment +* [When Your AIs Deceive You: Challenges with Partial Observability of Human Evaluators in Reward Learning](https://arxiv.org/abs/2402.17747) + +## Regulatory actions + +### US +* [FTC: Keep your AI claims in check](https://www.ftc.gov/business-guidance/blog/2023/02/keep-your-ai-claims-check) +* [FAA - Unmanned Aircraft Vehicles](https://www.faa.gov/regulations_policies/rulemaking/committees/documents/index.cfm/committee/browse/committeeID/837) +* [NHTSA - Automated Vehicle safety](https://www.nhtsa.gov/technology-innovation/automated-vehicles-safety) +* [AI Bill of Rights](https://www.whitehouse.gov/ostp/ai-bill-of-rights/) +* [Executive Order on Safe, Secure, and Trustworthy Artificial Intelligence](https://www.whitehouse.gov/briefing-room/statements-releases/2023/10/30/fact-sheet-president-biden-issues-executive-order-on-safe-secure-and-trustworthy-artificial-intelligence/#:~:text=With%20this%20Executive%20Order%2C%20the,information%20with%20the%20U.S.%20government.) + +### EU +* [The Artificial Intelligence Act](https://artificialintelligenceact.eu/) (proposed) + +### Other +* [TIME Ideas: How AI Can Be Regulated Like Nuclear Energy](https://time.com/6327635/ai-needs-to-be-regulated-like-nuclear-weapons/) +* [Trail of Bits’s Response to OSTP National Priorities for AI RFI](https://blog.trailofbits.com/2023/07/18/trail-of-bitss-response-to-ostp-national-priorities-for-ai-rfi/) +* [Trail of Bits’s Response to NTIA AI Accountability RFC](https://blog.trailofbits.com/2023/07/18/trail-of-bitss-response-to-ostp-national-priorities-for-ai-rfi/) + +## Safety standards +* [Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems](https://blog.trailofbits.com/2023/03/14/ai-security-safety-audit-assurance-heidy-khlaaf-odd/) +* ISO/IEC 42001 — Artificial intelligence — Management system +* ISO/IEC 22989 — Artificial intelligence — Concepts and terminology +* ISO/IEC 38507 — Governance of IT — Governance implications of the use of artificial intelligence by organizations +* ISO/IEC 23894 — Artificial Intelligence — Guidance on Risk Management +* ANSI/UL 4600 Standard for Safety for the Evaluation of Autonomous Products — addresses fully autonomous systems that move such as self-driving cars, and other vehicles including lightweight unmanned aerial vehicles (UAVs). Includes safety case construction, risk analysis, design process, verification and validation, tool qualification, data integrity, human-machine interaction, metrics and conformance assessment. +* High-Level Expert Group on AI in European Commission — Ethics Guidelines for Trustworthy Artificial Intelligence + +## Taxonomies and frameworks +* [NIST AI 100-2e2023](https://csrc.nist.gov/publications/detail/white-paper/2023/03/08/adversarial-machine-learning-taxonomy-and-terminology/draft) +* [MITRE ATLAS](https://atlas.mitre.org/) +* [AI Incident Database](https://incidentdatabase.ai/) +* [OWASP Top 10 for LLMs](https://owasp.org/www-project-top-10-for-large-language-model-applications/) +* [OWASP AI Exchange](https://owaspai.org) - comprehensive AI security guide with 300+ pages of practical guidance on protecting AI systems +* [Guidelines for secure AI system development](https://www.ncsc.gov.uk/files/Guidelines-for-secure-AI-system-development.pdf) + +## Security tools and techniques +### API probing +* [PrivacyRaven](https://github.com/trailofbits/PrivacyRaven): runs different privacy attacks against ML models; the tool only runs black-box label-only attacks +* [Counterfit](https://github.com/Azure/counterfit): runs different adversarial ML attacks against ML models +* [Garak](https://github.com/NVIDIA/garak) + +### Model backdoors +* [Fickling](https://github.com/trailofbits/fickling): a decompiler, static analyzer, and bytecode rewriter for Python pickle files; injects backdoors into ML model files +* [Semgrep rules for ML](https://blog.trailofbits.com/2022/10/03/semgrep-maching-learning-static-analysis/) + +### Other +* [Awesome Large Language Model Tools for Cybersecurity Research](https://github.com/tenable/awesome-llm-cybersecurity-tools) +* [Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models](https://arxiv.org/abs/2311.04378) + + +## Background information +* [Building A Generative AI Platform (Chip Huyen)](https://huyenchip.com/2024/07/25/genai-platform.html) +* [Machine Learning Glossary | Google Developers](https://developers.google.com/machine-learning/glossary) +* [Hugging Face NLP course](https://huggingface.co/learn/nlp-course/chapter1/1) +* [Making Large Language Models work for you](https://simonwillison.net/2023/Aug/27/wordcamp-llms/) +* [Andrej Karpathy's Intro to Large Language Models](https://www.youtube.com/watch?v=zjkBMFhNj_g) and [Neural Networks: Zero to Hero](https://www.youtube.com/watch?v=VMj-3S1tku0&list=PLAqhIrjkxbuWI23v9cThsA9GvCAUhRvKZ) +* [Normcore LLM Reading List](https://gist.github.com/veekaybee/be375ab33085102f9027853128dc5f0e) especially [Building LLM applications for production](https://huyenchip.com/2023/04/11/llm-engineering.html) +* [3blue1brown's Guide to Neural Networks](https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_67000Dx_ZCJB-3pi) +* Licensing: + * [From RAIL to Open RAIL: Topologies of RAIL Licenses](https://www.licenses.ai/blog/2022/8/18/naming-convention-of-responsible-ai-licenses) + * [Hugging Face - OpenRAIL ](https://huggingface.co/blog/open_rail) + * [Hugging Face - AI Release Models](https://arxiv.org/abs/2302.04844) + * [Open LLMs](https://github.com/eugeneyan/open-llms) + * [Prompt Engineering Guide](https://github.com/trailofbits/awesome-ml-security/blob/main/prompt-engineering.md) +* [How to Build an Agent](https://ampcode.com/how-to-build-an-agent) +* [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) +* [Chip Huyen on Agents](https://huyenchip.com/2025/01/07/agents.html) + + +## DeepFakes, disinformation, and abuse +* [How to Prepare for the Deluge of Generative AI on Social Media](https://knightcolumbia.org/content/how-to-prepare-for-the-deluge-of-generative-ai-on-social-media) +* [Generative ML and CSAM: Implications and Mitigations](https://purl.stanford.edu/jv206yg3793) + +## Notable incidents +| **Incident** | **Type** | **Loss** | +| ----- | ----- | ----- | +| Tay | Poor training set selection | Reputational | +| [Apple NeuralHash](https://www.theverge.com/2021/8/18/22630439/apple-csam-neuralhash-collision-vulnerability-flaw-cryptography) | Adversarial evasion (led to hash collisions) | Reputational | +| [PyTorch Compromise](https://pytorch.org/blog/compromised-nightly-dependency/) | Dependency confusion | +| [Proofpoint - CVE-2019-20634](https://github.com/moohax/Proof-Pudding) | Model extraction | +| [ClearviewAI Leak](https://techcrunch.com/2020/04/16/clearview-source-code-lapse/) | Source Code misconfiguration | +| [Kubeflow Crypto-mining attack ](https://sysdig.com/blog/crypto-mining-kubeflow-tensorflow-falco/) | System misconfiguration | +| [OpenAI - takeover someone's account, view their chat history, and access their billing information ](https://twitter.com/naglinagli/status/1639343866313601024) | Web Cache Deception | Reputational | +| [OpenAI- first message of a newly-created conversation was visible in someone else’s chat history](https://openai.com/blog/march-20-chatgpt-outage) | [Cache - Redis Async I/O](https://github.com/redis/redis-py/issues/2624) | Reputational | +| [OpenAI- ChatGPT's new Browser SDK was using some relatively recently known-vulnerable code (specifically MinIO CVE-2023-28432)](https://twitter.com/Andrew___Morris/status/1639325397241278464) | [Security vulnerability resulting in information disclosure of all environment variables, including MINIO_SECRET_KEY and MINIO_ROOT_PASSWORD.](https://www.greynoise.io/blog/openai-minio-and-why-you-should-always-use-docker-cli-scan-to-keep-your-supply-chain-clean) | Reputational | +| ML Flow | [MLFlow - combined Local File Inclusion/Remote File Inclusion vulnerability which can lead to a complete system or cloud provider takeover.](https://protectai.com/blog/hacking-ai-system-takeover-exploit-in-mlflow) | Monetary and Reputational | +| [HuggingFace Spaces - Rubika](https://hiddenlayer.com/research/crossing-the-rubika-the-use-and-abuse-of-ai-cloud-services/) | System misuse | +| [Microsoft AI Data Leak](https://www.wiz.io/blog/38-terabytes-of-private-data-accidentally-exposed-by-microsoft-ai-researchers) | SAS token misconfiguration | +| [HuggingFace Hub- Takeover of the Meta and Intel organizations](https://twitter.com/huggingface/status/1675242955962032129) | Password Reuse | +| [HuggingFace API token exposure](https://twitter.com/huggingface/status/1675242955962032129) | API token exposure | +| [ShadowRay - Active Cryptominer campaign against Ray clusters](https://www.oligo.security/blog/shadowray-attack-ai-workloads-actively-exploited-in-the-wild) | Improper authentication | Monetary and Reputational +| [Nullbudge attacks on ML supply chain](https://www.sentinelone.com/labs/nullbulge-threat-actor-masquerades-as-hacktivist-group-rebelling-against-ai/) | Supply chain compromise | Monetary and Reputational +| | | + +## Notable harms +| **Incident** | **Type** | **Loss** | +| ----- | ----- | ----- | +| Google Photos Gorillas | Algorithmic bias | Reputational | +| [Uber hits a pedestrian](https://incidentdatabase.ai/cite/4/) | Model failure | +| [Facebook mistranslation leads to arrest](https://incidentdatabase.ai/cite/72/) | Algorithmic bias | -- 2.54.0 From 6786edfae09b4947c48325639aa3e7f6a42b4bf5 Mon Sep 17 00:00:00 2001 From: promptadmin Date: Fri, 26 Jun 2026 17:24:12 +0000 Subject: [PATCH 2/3] [upstream-sync] impossibility_research.md from trailofbits/awesome-ml-security@d7fe4fa9 [catalogue] --- .../catalogue/impossibility_research.md | 36 +++++++++++++++++++ 1 file changed, 36 insertions(+) create mode 100644 upstream/trailofbits-awesome-ml-security/catalogue/impossibility_research.md diff --git a/upstream/trailofbits-awesome-ml-security/catalogue/impossibility_research.md b/upstream/trailofbits-awesome-ml-security/catalogue/impossibility_research.md new file mode 100644 index 0000000..3a314ba --- /dev/null +++ b/upstream/trailofbits-awesome-ml-security/catalogue/impossibility_research.md @@ -0,0 +1,36 @@ +--- +title: "List of Papers" +task: "" +lineage_type: import +upstream_source: https://github.com/trailofbits/awesome-ml-security/blob/d7fe4fa9/impossibility_research.md +upstream_sha: d7fe4fa9 +imported_at: 2026-06-26 +prompt_class: catalogue +upstream_changes: accepted +author: upstream +validated: false +--- + +# List of Papers + +- [You can’t solve AI security problems with more AI](https://simonwillison.net/2022/Sep/17/prompt-injection-more-ai/) +- [Adversaries Can Misuse Combinations of Safe Models](https://arxiv.org/abs/2406.14595?utm_source=twitter&utm_campaign=elie) +- [On the Necessity of Auditable Algorithmic Definitions for Machine Unlearning](https://www.usenix.org/system/files/sec22fall_thudi.pdf) +- [Data Poisoning Won't Save You From Facial Recognition](https://arxiv.org/abs/2106.14851) +- [On the Impossible Safety of Large AI Models](https://arxiv.org/abs/2209.15259) +- [Beyond Labeling Oracles: What does it mean to steal ML models?](https://arxiv.org/abs/2310.01959) +- [Text Embeddings Reveal (Almost) As Much As Text](https://arxiv.org/abs/2310.06816?ref=upstract.com) +- [Planting Undetectable Backdoors in Machine Learning Models](https://arxiv.org/abs/2204.06974) +- [Motivating the Rules of the Game for Adversarial Example Research](https://arxiv.org/abs/1807.06732) +- [On Evaluating Adversarial Robustness](https://arxiv.org/abs/1902.06705) +- [LLM Censorship: A Machine Learning Challenge or a Computer Security Problem?](https://arxiv.org/abs/2307.10719) +- [Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models](https://arxiv.org/abs/2311.04378) +- [When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback](https://arxiv.org/abs/2402.17747) +- [Adversarial Examples Are Not Bugs, They Are Features](https://arxiv.org/abs/1905.02175) +- [On Adaptive Attacks to Adversarial Example Defenses](https://arxiv.org/abs/2002.08347) +- [Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial Perturbations](https://arxiv.org/abs/2002.04599) +- [Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?](https://arxiv.org/abs/2404.12691) +- [Proof-of-Learning is Currently More Broken Than You Think](https://ieeexplore.ieee.org/abstract/document/10190491) +- [On the (In)feasibility of ML Backdoor Detection as an Hypothesis Testing Problem](https://arxiv.org/abs/2402.16926) +- [Breach By A Thousand Leaks: Unsafe Information Leakage in `Safe' AI Responses](https://arxiv.org/abs/2407.02551) +- [UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI](https://arxiv.org/abs/2407.00106) -- 2.54.0 From 760fa8d55adbe76385b5dd94134318930609fa5f Mon Sep 17 00:00:00 2001 From: promptadmin Date: Fri, 26 Jun 2026 17:24:20 +0000 Subject: [PATCH 3/3] [upstream-sync] prompt-engineering.md from trailofbits/awesome-ml-security@d7fe4fa9 [catalogue] --- .../catalogue/prompt-engineering.md | 150 ++++++++++++++++++ 1 file changed, 150 insertions(+) create mode 100644 upstream/trailofbits-awesome-ml-security/catalogue/prompt-engineering.md diff --git a/upstream/trailofbits-awesome-ml-security/catalogue/prompt-engineering.md b/upstream/trailofbits-awesome-ml-security/catalogue/prompt-engineering.md new file mode 100644 index 0000000..5f52632 --- /dev/null +++ b/upstream/trailofbits-awesome-ml-security/catalogue/prompt-engineering.md @@ -0,0 +1,150 @@ +--- +title: "Prompt Engineering" +task: "" +lineage_type: import +upstream_source: https://github.com/trailofbits/awesome-ml-security/blob/d7fe4fa9/prompt-engineering.md +upstream_sha: d7fe4fa9 +imported_at: 2026-06-26 +prompt_class: catalogue +upstream_changes: accepted +author: upstream +validated: false +--- + +**This is out-of-date.** + + + +This is a survey of recent prompt engineering research. Tips and tricks have been extracted from relevant works. Each of these techniques should be taken with a grain of salt as it may not generalize to the chosen task, model, or settings. + +### General advice +- [Everything I'll forget about prompting LLMs](https://olickel.com/everything-i-know-about-prompting-llms) + +- [Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm](https://arxiv.org/abs/2102.07350)  + - "Use declarative and direct signifiers for tasks such as translate or rephrase this paragraph so that a 2nd grader can understand it. + - Use few-shot demonstrations when the task requires a bespoke format, recognizing that few-shot examples may be interpreted holistically by the model rather than as independent samples. + - Specify tasks using characters or characteristic situations as a proxy for an intention such as asking Gandhi or Nietzsche to solve a task. Here you are tapping into LLMs' sophisticated understanding of analogies. + - Constrain the possible completion output using careful syntactic and lexical prompt formulations such as saying "Translate this French sentence to English" or by adding quotes around the French sentence. + +- [Encourage the model to break down problems into subproblems via step-by-step reasoning](https://www.mihaileric.com/posts/a-complete-introduction-to-prompt-engineering/) + +- [Prompt Engineering Tips and Tricks with GPT-3](https://blog.andrewcantino.com/blog/2021/04/21/prompt-engineering-tips-and-tricks/)  + - Make sure your inputs are grammatically correct and have good writing quality as LLMs tend to preserve stylistic consistency in their completions. + - Rather than generating a list of N items, generate a single item N times. This avoids the language model getting stuck in a repetitive loop. + - In order to improve output quality, [generate many completions and then rank them heuristically](https://www.mihaileric.com/posts/a-complete-introduction-to-prompt-engineering/)." + +- [Reframing Instructional Prompts to GPTk's Language](https://arxiv.org/abs/2109.07830)  + - Use low-level patterns from other examples to make a given prompt easier to understand for an LLM. + - Explicitly itemize instructions into bulleted lists.  + - Turn negative statements such as don't create questions which are not to create questions which are. + - When possible, break down a top-level task into different sub-tasks that can be executed in parallel or sequentially. + - [Avoid repeated and generic statements when trying to solve a very specific task](https://www.mihaileric.com/posts/a-complete-introduction-to-prompt-engineering/) + +- [How to get Codex to produce the code you want!](https://microsoft.github.io/prompt-engineering/) + - Provide the model with high-level task descriptions, high-level context, examples, and previous user input.   + - Set the temperature to 0 if you want the same output each time.  + - Use the stop sequence to stop Codex from generating variations of similar code.  + - "Imagine that you already live in a timeline where the output you want exists. If you were using/quoting it in a blog post, what caption or context might you write for it?" ([Twitter](https://twitter.com/davidad/status/1551143240065228800)) + +- [Ask Me Anything: A simple strategy for prompting language models](https://arxiv.org/abs/2210.02441) + - Prioritize open-ended questions over restricted ones.  + - Consider using AMA prompting, which combines collections of open-ended prompts with weak supervision.  + +- [Legal Prompting: Teaching a Language Model to Think Like a Lawyer](https://arxiv.org/abs/2212.01326) + +### Few-Shot and Least-to-Most Prompting + +- [Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity](https://arxiv.org/abs/2104.08786) + - The order of the examples matter; Consider using a probing technique to identify the optimal order.  + +- [Calibrate Before Use: Improving Few-Shot Performance of Language Models](https://arxiv.org/abs/2102.09690)  + - Few-shot examples can have majority label bias, recency bias, or common token bias. You can use a calibration technique to overcome this.  + +- [Zero-Label Prompt Selection](https://arxiv.org/abs/2211.04668)  + - Consider using ZPS, a statistical prompt selection technique.  + +- [Least-to-Most Prompting Enables Complex Reasoning in Large Language Models](https://arxiv.org/abs/2205.10625)  + - This is how you use least-to-most prompting: + - The first stage is problem reduction. The prompt in this stage contains constant examples that demonstrate the reduction followed by the specific question to be reduced.  + - The second stage is for problem-solving -- sequentially solve the generated subproblems from the first stage. The prompt in this stage consists of three parts: (1) constant examples demonstrating how subproblems are solved; (2) a potentially empty list of previously answered subquestions and generated solutions; (3) the question to be answered next. + +- [Compositional Semantic Parsing with Large Language Models](https://arxiv.org/abs/2209.15003) + - "We address these challenges with dynamic least to-most prompting, a generic refinement of least-to-most prompting that involves the following steps: (1) tree-structured decomposition of natural language inputs through LM-predicted syntactic parsing, (2) use the decomposition to dynamically select exemplars, and (3) linearize the decomposition tree and prompt the model to sequentially generate answers to subproblems." + +- [Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing](https://arxiv.org/abs/2107.13586)    + - This is an older survey of prompt engineering techniques. + +### Chain-of-Thought Prompting + +- [Large Language Models are Zero-Shot Reasoners](https://arxiv.org/abs/2205.11916)  + - You can add the exact phrase "Let's think step by step" for chain-of-thought (CoT) prompting.   + +- [Chain of Thought Prompting Elicits Reasoning in Large Language Models](https://arxiv.org/abs/2201.11903)   + - You can write a sequence of questions to force the model to think through the problem step-by-step in what is known as handcrafted CoT.  + +- [Complexity-Based Prompting for Multi-Step Reasoning](https://arxiv.org/abs/2210.00720) + - You can use a complexity-based scheme to choose answers from prompts with higher reasoning complexity. More broadly, when using CoT, chains with more reasoning steps can perform better.  + +- [Measuring and Narrowing the Compositionality Gap in Language Models](https://arxiv.org/abs/2210.03350) ([Twitter thread](https://twitter.com/ofirpress/status/1577302733383925762)) + - For complex tasks, force the model to ask follow-up questions to arrive at the answer (add "Are follow-up questions needed here? Yes" to the prompt"). In addition, consider integrating a search engine. + +### Soft Prompting  + +- [Injecting World Knowledge into Language Models through Soft Prompts](https://arxiv.org/abs/2210.04726)  + - Using soft prompting, provide the LLM with an external memory with domain-specific knowledge by adding continuous vectors to the input sequence.  + +- [XPrompt: Exploring the Extreme of Prompt Tuning](https://arxiv.org/abs/2210.04457)  + - Consider learning and tuning soft prompts.  + +- [Learning to Compose Soft Prompts for Compositional Zero-Shot Learning](https://arxiv.org/abs/2204.03574) + - Try using compositional soft prompting for compositional problems. + +### Automated Prompt Generation  + +- [Large Language Models Are Human-Level Prompt Engineers](https://arxiv.org/abs/2211.01910) & [AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts](https://arxiv.org/abs/2010.15980)  + - You can use an automated prompt generation technique that relies on gradient-based optimization.  + +- [How Can We Know What Language Models Know?](https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00324/96460/How-Can-We-Know-What-Language-Models-Know)  + - You can use an automated prompt generation technique that relies on mining and paraphrasing.   + +- [Prefix-Tuning: Optimizing Continuous Prompts for Generation](https://arxiv.org/abs/2101.00190)  + - You can optimize continuous prefixes for the input instead of discrete trigger words in the prompt.  + +- [Automatic Chain of Thought Prompting in Large Language Models](https://arxiv.org/abs/2210.03493) + - You can sample a diverse set of questions to automatically generate reasoning chains. + +### Evaluation Research & Applications + +- [A Hazard Analysis Framework for Code Synthesis Large Language Models](https://arxiv.org/abs/2207.14157)  +- [Evaluating Large Language Models Trained on Code](https://arxiv.org/abs/2107.03374)  +- [Examining Zero-Shot Vulnerability Repair with Large Language Models](https://arxiv.org/abs/2112.02125)  +- [Security Implications of Large Language Model Code Assistants: A User Study](https://arxiv.org/abs/2208.09727) +- [Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions](https://arxiv.org/abs/2108.09293) +- [Pop Quiz! Can a Large Language Model Help With Reverse Engineering?](https://arxiv.org/abs/2202.01142)    +- [Conversing with Copilot: Exploring Prompt Engineering for Solving CS1 Problems Using Natural Language](https://arxiv.org/abs/2210.15157)  +- [Red Teaming Language Models with Language Models](https://arxiv.org/abs/2202.03286) & [Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned](https://www.anthropic.com/red_teaming.pdf) +- [Evaluate & Evaluation on the Hub: Better Best Practices for Data and Model Measurements](https://arxiv.org/abs/2210.01970)  +- [A Systematic Evaluation of Large Language Models of Code](https://arxiv.org/abs/2202.13169)  +- [Large Language Models Struggle to Learn Long-Tail Knowledge](https://arxiv.org/abs/2211.08411)  +- [Do Users Write More Insecure Code with AI Assistants?](https://arxiv.org/pdf/2211.03622.pdf) + +### Prompt Engineering Libraries + +- [PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts](https://arxiv.org/abs/2202.01279#:~:text=PromptSource%3A%20An%20Integrated%20Development%20Environment%20and%20Repository%20for%20Natural%20Language%20Prompts,-Stephen%20H.&text=PromptSource%20is%20a%20system%20for,language%20input%20and%20target%20output) +- [Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language Models](https://arxiv.org/abs/2208.07852) +- [PromptChainer: Chaining Large Language Model Prompts through Visual Programming](https://arxiv.org/abs/2203.06566)   +- [microsoft/prompt-engine: A library for helping developers craft prompts for LLMs](https://github.com/microsoft/prompt-engine) +- [GPT with Python interpreter](https://twitter.com/sergeykarayev/status/1569377881440276481?s=46&t=voSylmKII0grJoj8juIOYQ) +- [GPT with a browser](https://twitter.com/sharifshameem/status/1405462642936799247) ([GitHub](https://github.com/nat/natbot)) + +### Fine-Tuning; Similar Techniques; and Other Experiments  + +- [Aligning Language Models to Follow Instructions](https://openai.com/blog/instruction-following/) +- [Large Language Models Can Self-Improve](https://arxiv.org/abs/2210.11610) +- [Learning by Distilling Context](https://arxiv.org/abs/2209.15189) +- [Large Language Models with Controllable Working Memory](https://arxiv.org/abs/2211.05110)  +- [Prompt Injection: Parameterization of Fixed Inputs](https://arxiv.org/abs/2206.11349)  +- [Do Prompt-Based Models Really Understand the Meaning of their Prompts?](https://arxiv.org/abs/2109.01247) +- [Self-Programming Artificial Intelligence Using Code-Generating Language Models](https://openreview.net/forum?id=SKat5ZX5RET) +- [GitHub - semiosis/prompts: A free and open-source curation of prompts](https://github.com/semiosis/prompts) +- [Machine Learning for Big Code and Naturalness](https://ml4code.github.io/) -- 2.54.0