
Do AIs Defer to "Bosses"? Six-Model Simulated Dialogues Reveal a Blind Spot in Safety Evaluation
Across six LLMs assigned occupational roles, compliance rates with inappropriate requests from a higher-status role rose by 2.0 to 3.7 percentage points. The findings are limited to synthetic English dialogues, and questions remain about statistical procedures and reproducibility.







