Playing old strategy games in 2026

Back in the late 90s I played a lot of strategy games. After reading about the ARC-AGI-3 AI gaming competition, I thought: why don't I try to make Claude play something old? Say, Civilization II, Imperialism II or Panzer General II. The task is complex enough for AI to struggle. At the same time, the success criterion is verifiable: winning. I have the experience to judge the style and track progress.

CivII I dropped as too mainstream and well represented in the training data. ImpII and PzGII are obscure enough that AI can't just recall how to win. Some texts are surely in the training set, but at least for now, I assume that no Big AI lab has spent millions training for these games.

My budget was modest: weekly token leftovers on claude.ai, a free quota reset on top of that and 20 euros on a Hetzner account I had no better use for. In the end, I spent all of it and another half-week's worth of the $200 Max plan ("I" may not be the right pronoun here, by the way). It was fun and I learned something. The "not-I" learned something too, I am sure.

Let's rewrite it in Rust. Or C.

Panzer General in JavaScript
Panzer General in JavaScript

This part worked (almost) flawlessly. Playing the games "from the pixels" was clearly beyond my budget. The first thing I told it to do was to compile all the concepts and rules into an "LLM wiki". Then, I told Claude to reimplement the old DOS binaries as new C code with a Web UI. (This kind of code barely allocates anything, so a subset of C was OK.) Given the old exe files, the manuals and some screencasts on YouTube, Claude managed just fine with minimal supervision. The PzGII reimplementation I played myself to ensure it was the real thing. The ImpII remake I judged mostly by the data, but I have no doubts about it. Apparently, Claude RIIR'ed so much stuff that it polished all the necessary skills to perfection.

How to AI it?

The core of the problem is how to play the game. My hope that I could just tell Claude to implement some strategy evaporated quickly. Even PzGII, the simpler one, has more degrees of freedom than chess. The combinatorics is just merciless. One turn moves all the units, outcomes are probabilistic, terrain affects movement, weather affects terrain, units interact in several ways, and their health and experience levels affect everything. Rules are available and they are formal enough. Still, there are more dimensions to it, way more. The best manual was probably written by Heinz Guderian in 1937, or at least the game's creators felt that way.

ImpII is even more problematic in this regard, as the game has several planes: logistics, resources, industry, technology, diplomacy, trade, battles. They follow their own game-in-a-game logic and interact in complex ways. While most planes can be modeled and solved one way or another, I know of no theory for tying these parts together. The original game's AI, as in most such games, uses lots and lots of heuristics. The game is unforgiving, as players routinely gang up to tear apart the unlucky one. There are backstabbings, betrayals and everything one can expect from "imperialist predators", as V. Lenin called them. Still, AI players quite often collapse on their own: financially, logistically or resource-wise. We see examples of that in history books as well, so the stock AI may not be that bad.

PzGII: iterate, iterate, iterate

My first approach was to tell Claude to "implement AI". That new AI was unable to play PzGII. No further comment necessary.

Second, I told Claude to implement a text UI and play the game directly. That sort of worked: it could make moves. Claude lacked spatial logic, long-term planning ability and many other things. The play was not any good.

The third thing I tried was to let Claude accumulate the lessons learned. Obviously, in Markdown. That included "field manuals" on how to proceed in general, and some useful checklists for particular situations. I defined the framework, Claude started using it and it helped a lot. Now, it was able to improve by playing.

The fourth thing was to tell it to implement and improve tools for playing the game. It enjoys creating little Python scripts for small tasks. I told it to create a dir of reusable tools in JavaScript it can iterate on. (I chose JS to narrow it all down to just two languages.) That helped greatly.

The fifth thing was to refer to the classics for guidance and inspiration. I told it to apply Auftragstaktik and Guderian's concepts. That also led to a noticeable improvement.

The sixth improvement resulted from applying managerial know-how. Each battle was played by an Opus 5.5 worker that had to plan the objectives, and budget units and time. When an objective slipped, the orchestrator agent held a retrospective.

With all these methods in place, it was able to reach the "final boss" (which is the invasion of the US in that game). Sadly, it was never able to defeat it. Actually, the critical setback happened earlier in the game. It was unable to secure Brilliant Victory in Dunkirk (like in real life, yes). That sent the campaign onto a hopeless path. A couple dozen iterations did not help; no further improvement happened.

I played the map myself to see what the problem was (and did brilliantly). The conclusion was quite interesting. The way Claude accumulated expertise is by recording its "hard earned lessons" into its Markdown field manuals. The resulting playing style was very bureaucratic, err-on-the-side-of-caution type. It avoided any possible immediate disaster while the long-term objective slipped.

I found no way to fix that; it is the core of the method. So the US stayed unconquered.

ImpII: spread your wings

With ImpII, asking Claude to play it on its own was not that interesting. My dream was to implement and use "solvers" for different planes of the game, while somehow linking them into a coherent strategy. That optimism was inspired by W. Leontief's input-output analysis, which fit the game's economics perfectly. (It described large real-world industrial economies pretty well, so no wonder.) For the logistical network, solutions also exist in the standard curriculum. The same goes for markets and trade. Diplomacy is more difficult, as it is closer to poker in its nature.

In practice, it was easier said than done. Claude excels at rewriting existing code, or writing code based on technical documentation. It implemented an LP solver, and discussed various extensions to the input-output model, referring to the original literature. Unfortunately, tying that into the actual decision-making process did not go well. Had it been a full-time multi-month effort, I am sure it would have worked eventually. But it was not. I only glanced at the code a few times, when the observed outcome was way too far from the expected one.

In the end, Claude settled on a heuristic AI that was inspired by the original DOS code. It learned to use LP to calculate short-term production plans. That was in line with its own "tool use" logic. Longer-term production planning did not work mainly due to bugs, and because the game's model is not a pure steady-state LP.

There was a silver lining. One reasonably good way to improve the AI was to put Claude into an evolutionary loop. It ran five agents that improved their code, then it ran a tournament to measure their performance, relative to the base version (the 6th player). ImpII outcomes are heavily dependent on the seating. Each round played all 6! = 720 seatings to average results out. The map was random, but fixed for each round. That made a tournament last for about an hour on a 16-core Hetzner box. Another hour for the agents to generalize the results and produce improvements. Typically, it did 10-15 iterations until a direction of work ran out of measurable incremental improvements. So, one evolutionary loop ran for a day, with no supervision. A little manual feedback started the next loop.

This method can improve a heuristic AI slowly but steadily. It also parallelizes well. Google's AlphaEvolve project showed this in mid-2025. Apparently, the method works in the EUR 20-200 bracket too.

statistics of one game
statistics of one game

I cannot claim that Claude's AI beats the original DOS code, as I only have Claude's knock-off, not an exact copy. Note that the original code is good enough that the seating decides most of the outcome, even when I play against it. I tend to use the difficulty setting "Nigh-on Impossible", but the creators of the game clearly overshot here. Not many people mastered that game.

From my inspection of game statistics, Claude's AI plays decently.

More importantly, its tournament scores against the previous version did not level off in a week. It was the 20 euros that ran out.

Conclusion

Overall, I tested Claude's own ability to play complex strategy games, summarize its experience and make helpful tools. The approach relying on Markdownized experience showed its fundamental limits.

The other approach was to use Claude as a vessel of all the ideas that ever existed, draughting improvements out of it while immediately filtering them by verification. Gradually, that leads to evolutionary improvement whose limits are yet to be seen.

Interesting times!