I had a great time playing Monster of the Week with GPT-4 when it first came out and the message limits were much higher. I loved how fluid it was as a GM and story partner. I had to teach it to push back and to realize characters with more interesting internal states, but that wasn't hard to do: I just coached it to reveal the character's inner position and desires as if it were discussing backstory.
It immediately knew all the rules (suggesting perhaps that GPT-4 was trained on some PDF libraries) and could both apply them literally and flexibly. My character's background was one where they deliberately repressed aspects which would govern their game moves, making all of my rolls go poorly. GPT-4 flexibly thought of ways to fold that idea in to its judgements on roll failures (in MotW, when a roll fails the GM "makes a move", generally progressing the story by raising the stakes). For instance, my character had a nascent form of prescience that they accessed by flipping a coin to make hard decisions. When the roll would go poorly, GPT-4 would induce visions in my character which would mislead them. We essentially brainstormed this move as a variant on the built-in one to make it more interesting given the weaknesses of my character's recalcitrance.
GPT-4 did an amazing job building environments. It proposed setting it in a sleepy town in the PNW. I've never been, so I thought it'd be interesting. GPT hallucinated a sleepy fishing town replete with a worn out lighthouse and a tourist trap B&B. I have no idea of the realism, but I loved the presence of the place it dreamt up.
Finally, it was all cut short after about an hour. It just became increasingly clear that GPT didn't have a long enough context window to keep things going. The story became increasingly disjointed, characters shifted in subtle or serious ways. No narrative cohesion was possible.
But man was it fun for about an hour. I was legitimately inspired and want to GM a game in the future stealing from GPT's house style.