AI在战略决策上能做得有多好?基于战略模拟的实证基准测试

How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations

Strategy Science · 2026
被引 1 · 同刊同年前 6%
ABS 3

中文导读

本文提出用战略教学模拟来评估大型语言模型的战略决策能力,测试了34个模型在Back Bay Battery模拟中的表现,发现前沿模型在管理战略不确定性上存在弱点。

Abstract

Benchmarks have helped fuel rapid progress in large language models (LLMs) across a variety of domains including math, science, dialogue, and coding. Yet no existing benchmark adequately captures the defining elements of strategic decision making: uncertainty, complexity, irreversible multiperiod moves, and delayed or noisy feedback. This gap limits our ability to assess and guide LLMs’ capabilities in strategy. We propose that established strategy teaching simulations provide an ideal benchmarking approach because (1) they approximate the essential features of real-world strategy, and (2) they do so in a controlled, replicable environment suitable for evaluation. To demonstrate this, we assess the performance of 21 proprietary and 13 open-source LLMs on the Back Bay Battery (BBB) simulation, a widely used exercise in strategy and innovation courses. The simulation requires balancing short-term profitability against long-term competitive positioning while integrating complex information about customer preferences and technological change. We built an interface enabling LLMs to interact with the simulation as though encountering it for the first time, masking identifiers to reduce contamination from prior training data. Our results show clear progress in composite BBB performance: Later models generally outperform earlier versions, and reasoning-focused models from late 2024–early 2025 (e.g., o4-mini, Claude Sonnet 4, Gemini 2.0 Flash) exceed even the average scores of historical MBA student cohorts. However, frontier models from mid-to-late 2025 (e.g., GPT-5, Claude Opus 4.5, Gemini 3) have declined, underperforming both earlier LLMs and MBA students. This decline is partially explained by a systematic bias toward exploiting the core business at the expense of investing in future growth. Overall, these findings highlight impressive advances in LLMs’ strategic abilities since their inception. At the same time, we document current frontier models’ surprising weakness in managing strategic uncertainty. This paper pioneers and provides guidance for using simulation-based benchmarking as a productive framework for strategy researchers to track progress, identify blind spots, and shape the trajectory of strategy-specific LLM capabilities. History: Accepted for the Special Issue: Can AI Do Strategy? Supplemental Material: The online appendix is available at https://doi.org/10.1287/stsc.2025.0444 .

战略管理人工智能大型语言模型基准测试战略模拟