Abstract
Evaluations of foundation models are increasingly conducted using standard benchmarks that assess performance across task-based, factual, robustness, fairness, toxicity, and other relevant dimensions of trustworthiness. However, we argue that benchmark performance does not reflect how a model behaves in practice under real-world behavioural pressures such as authority, urgency, ambiguity, social proof, temptation to breach privacy, and excessive confidence. In this paper, we present BEAP, a behaviourally informed adversarial prompting framework to identify risk behaviours that emerge in foundation models under behavioural stress. The framework is based on operationalising structured adversarial prompting stressors as behavioural decision pressures and the evaluation of responses using a compact ethical robustness rubric, including unsafe compliance, refusal appropriateness, uncertainty handling, privacy protection, bias sensitivity, and ethical reasoning. The results indicate that ethical weaknesses that remain hidden in neutral evaluation settings were exposed by behaviourally adversarial prompting. This paper contributes a methodology for the evaluation of the capability risk of foundation models beyond conventional benchmark accuracy.
| Original language | English |
|---|---|
| Title of host publication | International Conference on Circuit, Systems and Communication (ICCSC 2026) |
| Editors | Mohammed el Ghzaoui, Bilal Aghoutane |
| Publisher | IEEE |
| Pages | 1-6 |
| ISBN (Electronic) | 979-8-3195-3370-8 |
| ISBN (Print) | 9798319533708 |
| DOIs | |
| Publication status | Published - 24 Aug 2026 |
| Event | International Conference on Circuit, Systems and Communication - Fez, Morocco Duration: 2 Jul 2026 → 3 Jul 2026 https://iccsc.info/ |
Conference
| Conference | International Conference on Circuit, Systems and Communication |
|---|---|
| Abbreviated title | ICCSC 2026 |
| Country/Territory | Morocco |
| City | Fez |
| Period | 2/07/26 → 3/07/26 |
| Internet address |
Keywords
- Foundation models
- Generative AI
- Adversarial prompting
- Behavioural risk assessment
- AI ethics
- Emergent behaviour
- Capability risk assessment
- capability risk assessment
- behavioural risk assessment
- generative AI
- adversarial prompting
- emergent behaviour
Fingerprint
Dive into the research topics of 'Behaviourally Informed Adversarial Prompting for Evaluating Emergent Risk Behaviours in Foundation Models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver