Advisor
Wiki AI, Automation & Emerging Tech LLM Threats

LLM Threats 18 articles

01
Indirect Prompt Injection via External Content
Overview Indirect prompt injection via external content is a security risk in AI-driven systems where malicious actors manipulate inputs sourced from third-party or external data to influence the behavior of…
02
Jailbreaking and Safety Bypass Techniques
Overview Jailbreaking and safety bypass techniques refer to methods used to circumvent built-in restrictions or security controls in AI systems, particularly large language models and automated platforms. These techniques pose…
03
LLM Abuse for Social Engineering
Overview Large Language Models (LLMs) have become integral to AI-driven automation, enabling sophisticated natural language generation and interaction. However, their capabilities can be exploited for social engineering attacks, where adversaries…
04
LLM Abuse in Phishing and Fraud
Overview Large Language Models (LLMs) have become integral to AI-driven automation and natural language processing tasks, but their capabilities also present novel risks in cybersecurity. In particular, LLMs can be…
05
LLM Hallucinations and Security Impact
Overview Large Language Model (LLM) hallucinations refer to instances where AI language models generate outputs that are factually incorrect, misleading, or fabricated despite appearing plausible. In modern security operations, these…
06
LLM Model Extraction Attacks
Overview LLM model extraction attacks involve adversaries attempting to replicate or steal large language models (LLMs) by querying them extensively and analyzing their outputs. This risk is significant in AI-driven…
07
LLM-Assisted Malware Development
Overview LLM-Assisted Malware Development refers to the use of large language models (LLMs) to facilitate the creation, modification, or enhancement of malicious software. This emerging risk area impacts modern security…
08
Model Memorization and Privacy Exposure
Overview Model memorization and privacy exposure refer to the phenomenon where machine learning models, particularly large language models (LLMs), inadvertently retain and reproduce sensitive information from their training data. This…
09
Model Output Manipulation and Steering
Overview Model Output Manipulation and Steering refers to techniques and risks associated with influencing or controlling the outputs generated by AI models, particularly large language models (LLMs) and automated decision…
10
Monitoring and Logging of LLM Usage
Overview Monitoring and logging of Large Language Model (LLM) usage involve the systematic collection and analysis of interaction data between users and AI-driven language models. This practice is critical in…
11
Prompt Injection Attacks
Overview Prompt injection attacks are a class of adversarial techniques targeting AI systems, particularly large language models (LLMs), by manipulating input prompts to alter or subvert intended outputs. These attacks…
12
Prompt-Based Denial of Service
Overview Prompt-Based Denial of Service (DoS) is a cybersecurity risk emerging from the exploitation of AI systems, particularly large language models (LLMs), through malicious or excessive input prompts. This attack…
13
Secure Prompt Engineering Practices
Overview Secure prompt engineering practices involve designing, validating, and managing inputs to AI systems, particularly large language models (LLMs), to mitigate security risks and ensure reliable, trustworthy outputs. In modern…
14
Sensitive Data Leakage via Prompts
Overview Sensitive data leakage via prompts refers to the inadvertent or intentional exposure of confidential or private information through inputs provided to AI systems, particularly large language models (LLMs). This…
15
Supply Chain Risks in LLM APIs
Overview Supply chain risks in Large Language Model (LLM) APIs refer to vulnerabilities and threats arising from dependencies on third-party AI services and components within the AI supply chain. These…
16
Training Data Leakage Risks
Overview Training data leakage risks refer to the unintended exposure or incorporation of sensitive, proprietary, or confidential information within AI training datasets, which can compromise the security and privacy of…
17
Trust Boundaries in LLM Integrations
Overview Trust boundaries in Large Language Model (LLM) integrations define the points at which data, commands, or outputs cross between systems or components with differing levels of trust and security…
18
Unauthorized Fine-Tuning Risks
Overview Unauthorized fine-tuning refers to the modification or adaptation of pre-trained AI models without proper authorization or oversight, often leading to unintended or malicious outcomes. In modern security operations, where…