AI paper index

BOXBOX: Evaluating Large Language Models on Sequential Race-Strategy Decisions.

2026-08-25 · Zenodo (CERN European Organization for Nuclear Research)

One-line summary

An AI research paper on BOXBOX: Evaluating Large Language Models on Sequential Race-Strategy Decisions..

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

Large language models are increasingly proposed asdecision-makers in sequential, high-stakes settings, yet evaluatingtheir decision quality is difficult: benchmarks risk contaminationfrom training data, and the quality of a real decision is hardto score objectively. We introduce BOXBOX, a benchmark thatevaluates language models on Formula 1 race-strategy decisions,where choices are discrete, outcomes are objectively quantifiablein time, and a fresh post-cutoff season provides test data nocurrent model could have memorised. From eleven races weextract 196 decision points by fixed rules, score each model’scall against an ex-ante optimum computed by a calibrated racesimulator, and compare against the decisions taken by professional team strategists. On the primary evaluation set of 125 drydecision points from the 2026 season, we report three findings.First, every model evaluated falls well short of the human pitwall, beating the real team call on between 16 and 24 percent ofdecisions. Second, model price does not reliably predict decisionquality: the cheapest model tested, an open-weight system pricedroughly two orders of magnitude below the flagships, achieves thelowest mean distance from the ex-ante optimum, and we find nostatistically significant evidence that the more expensive flagshipmodels decide better. Third, accuracy and self-consistency divergesharply: the most accurate model reverses its own call on identicalinputs in roughly two of five cases, while one flagship reversesitself in half. A test for training-data recall finds a weak andinconsistent signal in two of five models, not corroborated bya same-circuit comparison, indicating at most a modest effectrather than the strong memorisation that would invalidate thebenchmark. Code, data, and the full preregistration are public.

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment