Ask either model for a game design document and you'll get something that looks finished: sections, headers, a neat little table of mechanics. The first draft is never the hard part. The hard part is the fifth revision, after you've changed the core mechanic twice and the document needs to still make sense.
What the first draft is good for
Either tool is fine for turning a vague pitch into a structured skeleton fast: core loop, win/lose conditions, a rough scope list. Treat this output as a table of contents you're allowed to gut, not a document you polish. Its real value is forcing you to answer questions you were avoiding, like what actually happens when the player loses.
Where the difference shows up
Long documents that get revised in place, where a change on page one has to stay consistent with page eight, are where the tools diverge. Long-context consistency and following an established doc's voice on edit are the two things worth testing directly on your document, because they vary by how you're using the tool, not just by benchmark scores you'll read elsewhere.
The test that actually tells you something
Don't compare them on a fresh pitch. Paste in a design doc you've already revised twice, ask for a targeted change to one system, and check whether the rest of the document still agrees with itself afterward. That's the failure mode that costs real time: a doc that silently contradicts itself after an edit, and nobody notices until a level designer builds the wrong thing.
The part neither one does
Neither tool knows what's fun. They can format your idea, spot an inconsistency, and generate the boring boilerplate sections (control scheme, platform targets) you don't want to write by hand. The core loop, the thing that makes someone want to play a second time, still has to come from you. Use the model to get out of the blank page faster, not to decide what the game is.