LLM benchmark for long-horizon sequential decision-making — language models play the turn-based strategy game LGeneral against the CPU over a REST API, scored on a graded victory scale.
benchmark rest-api game-ai ai-agents turn-based-strategy sequential-decision-making llm agentic llm-evaluation llm-benchmark lgeneral
-
Updated
Jul 16, 2026 - C