arXiv paper tests staged prompts across six frontier AI models
A newly posted arXiv paper describes experiments in which the same three-part prompt sequence was run ten times for each of six frontier AI models from OpenAI, Anthropic, xAI and Google DeepMind. The prompts move from asking about architectural preferences toward a fuller task, suggesting the study compares how different systems respond as questioning gets more demanding. The abstract is truncated, so final findings and conclusions are not yet visible.