Skip to main navigation Skip to search Skip to main content

Behaviourally Informed Adversarial Prompting for Evaluating Emergent Risk Behaviours in Foundation Models

Research output: Chapter in Book/Report/Conference proceedingConference proceeding (ISBN)peer-review

8 Downloads (Pure)

Abstract

Evaluations of foundation models are increasingly conducted using standard benchmarks that assess performance across task-based, factual, robustness, fairness, toxicity, and other relevant dimensions of trustworthiness. However, we argue that benchmark performance does not reflect how a model behaves in practice under real-world behavioural pressures such as authority, urgency, ambiguity, social proof, temptation to breach privacy, and excessive confidence. In this paper, we present BEAP, a behaviourally informed adversarial prompting framework to identify risk behaviours that emerge in foundation models under behavioural stress. The framework is based on operationalising structured adversarial prompting stressors as behavioural decision pressures and the evaluation of responses using a compact ethical robustness rubric, including unsafe compliance, refusal appropriateness, uncertainty handling, privacy protection, bias sensitivity, and ethical reasoning. The results indicate that ethical weaknesses that remain hidden in neutral evaluation settings were exposed by behaviourally adversarial prompting. This paper contributes a methodology for the evaluation of the capability risk of foundation models beyond conventional benchmark accuracy.
Original languageEnglish
Title of host publicationInternational Conference on Circuit, Systems and Communication (ICCSC 2026)
EditorsMohammed el Ghzaoui, Bilal Aghoutane
PublisherIEEE
Pages1-6
ISBN (Electronic)979-8-3195-3370-8
ISBN (Print)9798319533708
DOIs
Publication statusPublished - 24 Aug 2026
EventInternational Conference on Circuit, Systems and Communication - Fez, Morocco
Duration: 2 Jul 20263 Jul 2026
https://iccsc.info/

Conference

ConferenceInternational Conference on Circuit, Systems and Communication
Abbreviated titleICCSC 2026
Country/TerritoryMorocco
CityFez
Period2/07/263/07/26
Internet address

Keywords

  • Foundation models
  • Generative AI
  • Adversarial prompting
  • Behavioural risk assessment
  • AI ethics
  • Emergent behaviour
  • Capability risk assessment
  • capability risk assessment
  • behavioural risk assessment
  • generative AI
  • adversarial prompting
  • emergent behaviour

Fingerprint

Dive into the research topics of 'Behaviourally Informed Adversarial Prompting for Evaluating Emergent Risk Behaviours in Foundation Models'. Together they form a unique fingerprint.

Cite this