arXiv — cs.AI preprintsInternational8 October 2026
SWE-Game: Can Coding Agents Build the Games We Want?
This is an official announcement record
Firsthand records what arXiv — cs.AI preprints announced and links to the original. The wording below is theirs, not ours.
arXiv:2609.33678v3 Announce Type: replace Abstract: We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemented games. Evaluation combines engine-state chec
Read the official announcement
Opens arxiv.org
More from arXiv — cs.AI preprints
- COMPASS: Finding Where Reasoning Lives in Language Models8 October 2026
- Decoupled Multi-Agent Orchestration8 October 2026
- OTel: Open Telco AI Datasets, Benchmarks, and Models8 October 2026
- Comparative review of hybrid forecasting models for short-term prediction of building thermal load8 October 2026
- Dynamical low-rank equilibrium computation for stochastic games between advanced persistent threats and moving target defense8 October 2026
This content is for informational purposes only and is not professional advice. Specifications, prices, plan tiers, and features change frequently and may differ from what is shown here; verify current details on the manufacturer's or company's official page before purchasing. Ratings are based on analysis of published documentation, not independent lab testing.