错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ASTRA: Adversarial Stealthy Trigger Reasoning Attacks for Black-Box LLMs

  • Seong-Gyu Park,
  • Sohee Park,
  • Daeseon Choi

摘要

Large language models (LLMs) are increasingly deployed in safety-critical contexts, yet their vulnerability to inference-time backdoor attacks remains insufficiently understood. We present Adversarial Stealthy Trigger Reasoning Attacks, a reinforcement learning-based framework for automatically discovering stealthy, query-agnostic reasoning triggers that compromise LLM inference in a fully black-box setting. ASTRA formulates trigger construction as a sequential decision-making problem and uses PPO to optimize both character-level and word-level triggers without access to model parameters or training data. Leveraging a perplexity-guided reward, ASTRA generates semantically natural word-level triggers that evade detection mechanisms while reliably activating malicious reasoning chains. Across five reasoning benchmarks and three representative LLMs, ASTRA consistently outperforms existing heuristic methods, achieving substantially higher attack success rates (ASR) and lower detection rates. Our results highlight that word-level triggers, optimized under ASTRA, provide particularly strong stealthiness by exploiting the model’s inherent language priors. This work provides a new methodology for systematically evaluating LLM safety under realistic black-box adversarial conditions and reveals that inference-time reasoning-chain manipulation poses a significant yet underexplored threat. ASTRA contributes a practical and extensible framework for future research in LLM vulnerability assessment and defense.