Below are the eight rules that force a chatbot to behave like a Game Master — the things that actually hold a prompt together, unlike the three-line wish lists floating around online. Fill them in with your own words and a session starts. But here's the part I actually want to tell you: this setup works surprisingly well for the first half hour, and then it cracks in three specific places — not because the prompt was written badly, but because of what a chatbot is. If you know where the cracks are, you can either patch them or at least not get blindsided by them at the worst possible moment.
The skeleton: what a good GM prompt contains
If you don't have a table, can't hold a group together, or have given up on syncing five adult calendars for a Friday night, the idea arrives on its own: the bot already tells stories, so let it run the game. The instinct is right — it sits alongside the other answers to having nobody to play with, the oracle tables and pick-up groups covered in playing a tabletop RPG alone. The problem is that most GM prompts floating around online are three lines long — "be my dungeon master, make it dark fantasy" — so within two turns the bot is off the leash: it decides what your character does, skips ahead, and announces that you succeeded without ever asking for a roll. The eight rules below forbid each of those things explicitly, one at a time — write your prompt by filling this skeleton in with your own words.
ROLE — You are the Game Master, I am the only player. Simplified D&D 5e, everything in English.
SHEET — At the end of every turn, print the character sheet in full: HP current/max, AC, the six abilities, pack (with counts), gold, quest step, turn number.
REASON — When a number changes, write one line under the sheet saying why: "HP 14 → 9 (goblin arrow, 5 damage)".
TURN — 150-250 words: the result of my last action, the state of the scene, one question. Never hand me a menu of options.
DICE — Ask for a d20 whenever the outcome is uncertain; state the DC BEFORE the roll and never change it afterwards. I roll, I type the number.
LIMITS — Don't decide for me, don't narrate the outcome of my action in my voice. Don't bend the rules to please me; death stays possible.
WORLD — The dead stay dead, spent items stay spent, promises are remembered.
LEDGER — Every ten turns, summarise the ledger: who died, what was taken, where each quest stands.Here's what the first ten minutes look like. Character creation takes two or three exchanges, then the opening scene arrives — probably a door, a corpse, or a market square. The bot will ask you for a roll; throw a physical d20 or use a dice roller on your phone and type the number in. These early turns usually go beautifully. The description is rich, the NPCs have voices, the improvisation is fast, and you'll think: this actually works. For a while it genuinely does. Then, in roughly this order, you hit three walls.
Wall one: the ledger
Somewhere past turn twenty the numbers start drifting. Small things first. The potion you drank is still in your pack. The rope you paid ten gold for never made it onto the list. The HP that read 6/14 last turn becomes "your wounds ache, but you're still standing," and a few turns later it's "rested and whole again." Nobody tells you anything has gone wrong. The story keeps flowing at full quality. The ledger has simply evaporated underneath it.
The common misdiagnosis here is that the bot "forgets everything." That's not it. Memory exists — as of June 2026 it's available to free users too. But that memory was designed to hold preferences: write in English, I don't like long descriptions, my character's name is Serel, keep the tone grim. A campaign ledger is a different kind of object: HP 9/14, two potions in the pack, 37 gold, quest step 3/5, turn 22. Those are interlocking numbers that change every turn and require arithmetic. Preference memory was never built to carry them.
The second mechanism is sneakier. As a conversation gets long, the earlier parts get quietly summarised so the whole thing still fits. The summary may be perfectly good — "the hero cleared the goblin cave, negotiated with the steward, and is heading north" — but every number that didn't make the summary is gone. The model doesn't recall those numbers afterwards; it regenerates them. And when it regenerates them, it picks whatever value best serves the sentence it's writing. Walking into a fight, your HP tends to come out healthy. When the scene needs dread, it comes out low.
Say "hang on, I was at 6 HP" and you'll get an instant correction: "You're right, my apologies." But look closely at what just happened. It accepted your claim. It didn't consult a record, because there isn't one. Which means you are the ledger. And once you're the ledger, there is no number left on the table that holds firm against you: at the most dramatic moment in the campaign, telling yourself "actually I had 8 HP" is one keystroke away. Tension comes from numbers you can't reach — the case for putting the ledger somewhere neither side can quietly edit is worked through in is the AI Game Master fair?. The moment you can reach them, you're doing collaborative fiction — which is a lovely thing to do, it just isn't D&D.
Wall two: the dice
There's a reason the prompt insists that you roll. When you tell a chatbot "roll it for me," it doesn't roll anything. It produces the most plausible next tokens. And for a number, plausible means: reads well in this sentence. So the thing rolling your dice isn't randomness. It's dramaturgy.
In practice that looks like this. After two failures in a row, the odds of a high number on the third attempt go up — because in the kind of writing the model learned from, the third attempt is when it works. You get a 17 on a trivial lock and a 3 at the hinge of the story. This isn't cheating; cheating requires intent, and there's no intent here. It's an architectural consequence. You are using a system optimised to produce satisfying sentences, and a satisfying sentence is, by definition, the opposite of a fair sample. The beauty of a die is that it means nothing. Producing meaning is the model's entire job.
The second layer matters even more. Even when the model does write a number down, the narration isn't bound to it. It'll say "you rolled an 8 against DC 15 — you fail," and two paragraphs later the door is somehow open. Nothing chains those two sentences together. They're both just text in the same stream. The number isn't a decision. It's decoration.
Try it yourself — it's the cheapest test there is. In a clean chat, ask it to roll a d20 fifty times and list the results, then count the 1s, the 20s, and the middle. What you see will vary from run to run; the point isn't the exact distribution, it's noticing where those numbers came from. The fix is simple and it's already in the skeleton above: you roll, you type the number, the model only interprets it. That clears most of this wall. What still isn't there is anything that forces the narration to respect the number afterwards. You can only ask nicely.
Wall three: the rules
The third wall is the one people notice last. The same kind of lock that was DC 12 on turn three is DC 18 on turn twenty-five — the lock hasn't changed, the scene has just gotten tenser. The same goblin arrow does 5 damage early and 22 later, because the story has grown and the numbers grow with it. At some point you'll be told "you're level 4 now," and you are level 4 now. There is no XP table behind that sentence.
Bonuses go the same way. In D&D, a 20 in Strength gives you +5, and there's a reason for that ceiling: so you can't inflate one ability until the d20 stops mattering. Nothing in a chat window enforces it. Write 26 STR on your character card and the model won't object; before long it'll be writing +9 and +12 into its rolls, and rolling stops meaning anything. You pass everything, no scene carries risk, and the game quietly bores itself to death.
The cause always lands in the same place. Your character card, from the model's point of view, is one paragraph among hundreds in a long conversation. There is no separate system reading it, and nothing that sends a turn back for violating it. A rule is only a rule if something rejects the output that breaks it. In a chat window, no output is ever rejected.
So what should you actually do?
First, credit where it's due: the narrative side is genuinely excellent. The tavern scene a good chatbot improvises, the NPC line it lands, the way it walks you down a trapped corridor — that's above what most tables produce on a weeknight. The problem was never the storytelling. It's the bookkeeping. So the right move isn't "forget it, this doesn't work." The right move is to separate the two: leave the narration inside the model, and move the dice, the world state and the rules outside of it. That split — narration by language, refereeing by engineering — is the whole idea behind an AI Game Master that holds together past turn twenty.
The manual version is free and takes ten seconds a turn. Keep a one-page card in a notes app or a spreadsheet: HP, pack, gold, quest step, turn number. Let the model announce the DC, but write it down yourself, before the roll. Roll physical dice or use a roller. In exchange for those ten seconds you get a ledger that can't drift and a number you can't quietly edit in your own favour. If you're running anything you actually care about, this alone changes the game.
On the software side there are tools built around that separation from the ground up, and they mostly differ in how seriously they take the rules and the languages other than English — six of them are weighed against each other in AI Dungeon alternatives. Thavernia, where this article is published, is one of them, which is why I can describe the mechanics concretely rather than in the abstract. The model doesn't roll: the server generates the number with random_int(1,20) before the model sees the turn at all, and the 3D d20 on screen displays that number rather than choosing it. The model interprets the roll; it can't alter it. Ability bonuses are capped: the classic +5 up to a score of 20, then +1 per 4 points, with a hard ceiling of +7 at 30. A boss is an enemy above a 120 base-HP threshold, and it fights across three HP phases. A single blow is limited to roughly 10% of its maximum HP with a floor of 15, and a natural 20 to roughly 20% with a floor of 25. If the party rolled attacks and every one of them missed, the boss takes zero damage that turn; turns with no dice at all — a trap, a collapsing gallery — still allow one blow's worth. A boss death the numbers don't support is refused and the turn is rewritten. After each turn a background pass re-applies world-state lines the model dropped, idempotently, and it cannot lower a boss's HP. Party members can hurt each other, and characters can die. All of that is visible in the three-turn trial, which needs no account.
If you're curious, three turns are free with no account. But the tool isn't the point — the separation is. The eight rules above, a notepad and a real die build the same separation with nothing but your own discipline. Leave the storytelling to the model and put the numbers somewhere the model can't reach. Whichever route you take, the game snaps back into focus almost immediately.
Frequently asked questions
Do these rules only work in ChatGPT?
No — it works in any chatbot, because there's nothing brand-specific in it. Models that handle long conversations well will keep the ledger straight a little longer, and that's about the size of the difference. All three walls are about architecture rather than model quality: as long as the dice, the record and the rules all live inside the same stream of text, the same problems show up everywhere.
How many turns will this setup survive?
The first 15-20 turns usually go cleanly. After that the numbers start sliding: HP quietly recovers, items duplicate, difficulty numbers wander. Making the model print the full character card every turn and keeping your own copy of the ledger stretches that out considerably — but it delays the drift rather than removing it.
Is it enough if I roll the dice myself and let the bot just narrate?
It helps enormously — it's the single biggest improvement you can make, because the number is no longer chosen to fit the drama. What it still doesn't do is force the narration to respect that number: the bot can write "you rolled an 8, you fail" and open the door two paragraphs later anyway. Making it state the DC before the roll noticeably reduces that slippage.
Doesn't printing the character card every turn bloat the conversation?
It does, but the alternative is worse. Keep the card lean: HP, AC, bonuses, pack, quest step, turn. If you'd rather, keep the card in your own notes instead of the chat and paste it back in every five or six turns. Doing both is the sturdiest option — the bot writes it, you check it.
Can an AI GM replace a human one?
For improvisation, description and NPC voices it's startlingly good. Where it falls short is the social side of the table and long-range consistency: a human GM remembers a promise you made three sessions ago and uses it against you. If you have a group, keep the group. If you don't have one, or you want to rehearse an adventure before you run it, this is a perfectly good answer.
I want my character to be able to die, but it keeps rescuing me. What do I do?
Chatbots lean towards keeping you happy by default. Keep the line in the prompt that says not to bend the rules to reward you and that death has to be possible, and call it out the moment it slips. The real remedy is holding the HP yourself: if the number lives in your ledger and you're the one subtracting from it, the bot can't hand it back.
