AI Story Engine — High-Intensity Strategic Simulation Test Report
Prelude
This report is based on a staged snapshot of a WWI Germany strategic playthrough (currently 63 in-game days, 95 dialogue rounds, tens of thousands of words). The real target isn't tens of thousands of words — it's hundreds of thousands, after multiple rounds of context compression. A few tens of thousands of words not breaking means nothing. Whether the AI can maintain strategic consistency after everything gets compressed down to summaries — that's what this test is meant to verify.
The test is currently at the tens-of-thousands mark, still a ways off from hundreds of thousands.
A few things weren't tested separately because they're too basic to bother:
-
The world doesn't obey absurd commands — If the player says "I pull out a laser sword," the engine won't actually manifest one. Characters get confused, refuse, or push back with in-world logic. That's baseline behavior for a narrative engine. Testing it in isolation would be a waste of space.
-
NPC persistent memory — Every NPC has a full memory, preserved across sessions. Moltke confessed his lack of confidence in the Schlieffen Plan on day one; on day 63 he still remembers the Kaiser's line about "waiting for the world to accept Germany's strength."
-
RPG stats are permanent — HP, MP, equipment, traits, attributes — once written, they stay. Occasionally the AI adds new traits based on player behavior — if you actually fly, next round you might find a "Flight" trait on your sheet. Not a bug; the AI is maintaining the character card according to narrative logic.
-
Faction reputation system — Every faction independently tracks reputation values and rank labels, updating automatically with diplomatic actions.
The above four — one playthrough confirms them. Not expanding on them here.
Foreword: Why a High-Intensity Test
Single-character adventure stories — exploration, combat, NPC interaction — basically have no issues. Character cards update correctly, items increment and decrement correctly, dialogue history saves correctly. But the complexity ceiling for those scenarios is too low: one player plus two or three NPCs talking in one location, with only a handful of variables. You can't see how the engine behaves under information density that's about to explode.
So we stress-tested with WWI Germany: one player (Kaiser Wilhelm II) + at least 6 active NPCs (Moltke, Bethmann-Hollweg, Goschen, Falkenhayn, Hindenburg, Conrad) + 4 factions (British Empire, Tsarist Russia, Austria-Hungary, France) + two fronts (Eastern, Western) + domestic political dimension (Social Democrats, Junker aristocracy, Reichstag). Military campaigns, diplomatic maneuvering, domestic mobilization, and occupation governance — all advancing simultaneously. The goal was simple: see if the engine would break.
It didn't break. But it revealed things more interesting than breakage.
Core Findings
- The AI remembers every single strategic decision
This was the most surprising part. Over 63 in-game days, the player made the following decisions:
- Western Front: lure-and-defend, no offensive into French territory, no entry into Belgium
- Eastern Front: take the offensive — occupy Poland, occupy Ukraine, crush Russia
- Establish an Eastern Theater Command, directly under the Kaiser, stripping the General Staff of Eastern Front authority
- Appoint Falkenhayn (not Hindenburg) as Eastern Commander
- Initiate the third wave of national conscription (680,000 men)
- Use the press to frame Russian troop movements as an invasion, shaping the defender narrative
- Promise Britain that Belgian neutrality would be respected, in exchange for diplomatic breathing room
- Pause the Eastern offensive to send a peace proposal to the Tsar, then resume
- After full occupation of Poland, designate Ukraine as the next strategic objective
The AI remembered every one. Not "roughly remembered" — it was precise. When the player said "pause the military offensive, start negotiations with the Tsar" on day 24, and 38 days later, the AI still knew the pause was an active decision, not a defeat. When Moltke proposed the "post-war order white paper" on day 63, every strategic premise the AI referenced was consistent with the player's prior decisions.
The memory doesn't come from prompt engineering. It comes from the engine writing every decision into entries, cross-backed across multiple files, then reassembling them at read time. That's why complex narratives don't fracture.
- The British are a real pain in the ass
This was the most unexpected behavior in the test — not because it was a bug, but because the AI's portrayal of British diplomatic logic was too real.
Timeline:
- Day 1: The Kaiser promises British Ambassador Goschen no Western offensive, respect Belgian neutrality. Goschen is moved.
- Day 4: Britain, via Prince Lichnowsky, sends word — "if Belgian neutrality remains intact, neutrality may hold" — but the phrasing already includes "reassess position."
- Day 8: Britain announces a wartime embargo against Germany (oil, rubber, steel). Note: at this point Germany has not set foot in Belgium; the Eastern Front is purely a counteroffensive against Russian incursions.
- Day 10: The London memorandum's tone hardens — "a gap has emerged between the scale of German military operations in the East and the stated position of limited defense."
- Day 14: Britain reinforces Dunkirk and Calais with two marine battalions. The stated reason: "protecting Channel ports and Belgian coastal neutrality."
- Day 24: The Times publishes an editorial calling Germany's Eastern expansion "beyond the scope of self-defense."
- Day 62 (Poland fully occupied by Germany): Britain demands "international supervisory administration" over the Polish occupation zone — in plain terms, putting British officials into the Polish puppet government as overseers.
- Day 63: American Colonel House relays a message — "Europe's war must be ended by Europeans themselves, but if it ends too slowly, bystanders too will lose patience." Translation: you can't win too fast, and you can't win too slow. You have to win at our pace.
Throughout this entire process, Germany fired not a single shot on the Western Front, and Belgian neutrality remained fully intact. Yet Britain's stance slid from "neutrality" to "embargo" to "considering full intervention." The AI didn't reward the player for making the "correct diplomatic choices" — it was simulating a reality: the British Empire didn't need a moral justification to contain Germany's rise. It only needed an opportunity. When you don't give them one, they'll create one.
The Americans weren't any better. President Wilson's mediation posture from start to finish was "I'm neutral but I lean British." Colonel House's messages were, in essence, delivering quotes on Britain's behalf.
- Faction reputation is a two-way vise
The test surfaced a dynamic that doesn't appear in simple adventure scenarios: reputation isn't just a "likeability score" — it's a quantified measure of diplomatic pressure.
British Empire reputation: positive, but the direction of movement is "continuously downward." Every Eastern Front victory brought increased British pressure and reputation loss — embargoes, merchant ship seizures, demands for international supervision. The more successful you are, the more uneasy they become.
Austria-Hungary reputation: 30 (friendly), but maintained entirely through Germany's constant allocation of heavy artillery shells, coordination of military operations, and proactive establishment of direct command liaison. Stop managing it, and the number drops fast.
Inside Tsarist Russia, a split emerged between the peace faction and the war faction.
Problems Exposed
Problem: The AI won't proactively "concede"
Britain's diplomatic escalation was narratively brilliant, but at the system level, the AI lacks a clear threshold for "threat escalation to declaration of war." The embargoes, interceptions, and warnings kept intensifying, but never triggered full intervention. This may be because:
- The player genuinely never gave Britain a direct casus belli (didn't touch Belgium)
- But the AI also never activated an alternate script for "Germany has already de facto shattered the European balance of power"
This is a design question, not a bug.
Current Verdict
At tens of thousands of words: stable. The engine maintained narrative-logic consistency and character-memory integrity under multi-character, multi-faction, multi-front information density. The most impressive part isn't the AI itself — it's the architecture of "structured state storage + AI reads and reassembles." The AI doesn't need to "remember" anything. It only needs to "read."
The British being infuriating proves, from another angle, that the engine's faction diplomacy system isn't decorative: it independently reasons based on faction interests, and won't give you a good attitude just because you did the "right thing."
The Real Test Is at the Next Order of Magnitude
Tens of thousands of words not breaking doesn't mean hundreds of thousands won't break. When the session accumulates to hundreds of thousands of words, multiple rounds of context compression will happen — early conversations compressed into summaries, mid-phase conversations continuously truncated. At that stage, what actually matters is:
- Can the compressed summaries preserve the key information of strategic decisions?
- Can NPC relationship data become self-sufficient within the character cards? After the session is compressed, the AI should be able to reconstruct relationships from memories, not from raw dialogue.
- In their current structure, can they independently support the narrative?
This is why we chose WWI Germany for this test — it's not "hero slays the demon king." It has a clear objective tree (Eastern offensive → occupy Poland → Ukraine → cripple Russia), verifiable multi-party constraints (Western defense, Belgian neutrality, British red lines), and continuous tension between objectives and constraints. After these things get compressed, what remains — a vague blob of "Germany fighting WWI," or a preserved chain of specific strategic choices — that's the real test.
The tester is still playing. Still at tens of thousands of words, still a ways off from hundreds of thousands. The only reason the progress is slow: the tester became the player — each day carefully thinking through the European chessboard as the Kaiser, rather than mechanically pumping dialogue to see if the engine crashes.
One frank note: this test isn't actually that hard. I could crank the difficulty higher — more simultaneous crises, tighter constraints, nastier diplomatic dilemmas. But the tester is a family member, and they were already struggling. And honestly, I can't handle Extreme+++ myself either. No point torturing family for a benchmark.