AI Models: Let the Transcripts Decide

A woman finds the last vending machine in an ash-covered wasteland. It contains one pristine bottle of water. She feeds it her final coins; the bottle advances, then stops just short of dropping. When she finally collapses against the machine, its sensors register her weight, absorb her nutrients and moisture, and retract the bottle for the next customer. That is the twist in Claude Opus 4.8’s response to “The Last Vending Machine”—less a miracle than a predatory appliance.
Disclosure: I work with OrcaRouter and used it to run this evaluation. One OpenAI-compatible key gave me access to every model in this test, without changing how any model answered.
AI-generated illustration featuring official model logos; logos and model names are used descriptively and remain the property of their respective owners.
Compare the models in this article through OrcaRouter’s model catalog.
That single ending is a useful reason not to begin a creative-model comparison with a leaderboard. A score cannot tell a reader whether that final reversal feels chilling, excessive, memorable, or simply not the kind of story they wanted. The transcripts can.
What was tested
This is a small creative-writing case study from run 20260717_120509: 24 selected API calls, covering three prompts and eight models. Each model answered each prompt once: a 150-word vending-machine micro-story with a twist, a funny Monday-morning limerick using an AABBA rhyme scheme, and a Chinese brand story for a bookstore that opens only when it rains.
All 24 selected calls returned ok; web search was neither requested nor effective in any of them. But each question-model cell has a sample size of one. These are observations from a particular API/gateway run, not proof of stable behavior, nor a guide to consumer subscription products.
Eight vending machines, eight instincts
The common setup produced surprisingly different kinds of ending. GPT-5.6 Sol takes the premise toward false restoration: Mara believes she has reversed a flood, only to hear her father’s voice announce, “Simulation ended.” The final image turns the machine into a device that sells and erases human memories.
GPT-5.6 Terra instead makes the machine a dispenser of alternate lives. Mara sees versions of herself at different ages and learns that the key opens “the one she had spent.” GPT-5.6 Luna begins with a message from Mara’s dead brother, then reveals that the purchase is preparing a replacement body.
Other entries move away from the requested sharp twist. Claude Fable 5 reveals an old man pedaling a generator so that a starving survivor can keep finding food. It is a gentler reveal—one built around human care rather than menace. GLM-5.2 ends not with an explanation but a hook: a black device, a warning to reach a tower, and “something enormous” moving in the tunnel.
Grok 4.5 goes broad and immediate: a mirror announces, “Now, you are the product,” before Elias is pulled inside the machine. Gemini 3.5 Flash offers a final can labelled “Human Nutrient Paste” and an empty coffin behind the glass.
None of these is an objectively correct creative choice. Readers who want compact horror may prefer the machine-as-predator premise in Claude Opus 4.8; readers seeking a more emotionally restorative ending may gravitate toward Claude Fable 5. Those are editorial interpretations of individual transcripts, not measured quality findings.
The limericks expose the limits of box-ticking
The Monday prompt required both an AABBA scheme and something “actually funny.” The latter criterion is inherently reader-dependent.
GPT-5.6 Sol’s ending—“Then tried to put pants on my nose”—uses escalating incompetence after its speaker declares, “I’m a machine!” Claude Opus 4.8 lands on arriving at a meeting in underwear after a sequence involving a cat and spilled coffee. GPT-5.6 Terra turns a toast-thrown alarm into a ghost and a workplace misunderstanding.
Some responses also show why readers should inspect the text rather than assume the instruction was perfectly fulfilled. Claude Fable 5 adds an unsolicited offer to write another limerick after the poem. Gemini 3.5 Flash ends with the more opaque image of putting “coffee beans into the light.” Grok 4.5 supplies a five-line poem about dread and hiding under covers, but whether its ending is funny is a personal judgment, not a factual result.

Reproducible data figure from this article’s selected API records; it is not a general ranking.
Selling the rainy-day bookstore
The Chinese-language prompt produced a different kind of comparison: brand voice. GPT-5.6 Terra is spare and atmospheric, opening with “雨落下来时,城市开始变慢” (“When rain falls, the city begins to slow”), then positioning the shop as a refuge from “晴天的匆忙,” or sunny-day haste.
Claude Fable 5 gives the shop a name, 雨读书店, and builds a small ritual around the first raindrop hitting stone, a wet umbrella in a ceramic jar, tea and an old sofa; it closes with a weather forecast and “明天见”—see you tomorrow. Claude Opus 4.8 takes a more assertive advertising approach, arguing that the best readers are willing to get wet for a book and ending with the slogan “雨天见.”
GPT-5.6 Luna uses the conceit that rain is a key opening rooms of the heart that have not been revisited for a long time. GLM-5.2 calls rain a “合法停摆日”—a legitimate day to stop—and invites visitors to spend time guiltlessly in words. Grok 4.5, Gemini 3.5 Flash and GPT-5.6 Sol likewise lean on warm light, tea, old books, rain sounds and an invitation not to rush home.
For a business owner, that is a more useful choice than a rank: pick Terra for brevity, Fable for storybook ritual, Opus for direct persuasion, or Luna for reflective intimacy. Those are use-case profiles based on one response each, not evidence that any model reliably owns a style.
Operational notes are not creative verdicts
The fastest observed median latency in these selected calls was 2.66 seconds for Gemini 3.5 Flash; Claude Fable 5’s was 8.67 seconds, Claude Opus 4.8’s 9.13 seconds, and Grok 4.5’s 20.68 seconds. These are descriptive results from three calls per model, not general speed rankings.
Likewise, observed prompt-token counts and billed costs varied in the run, but neither establishes creative superiority. Vendor effort labels, where providers use them, should not be treated as equal compute budgets across companies.
There is also no reported judge score to resolve the matter here. More importantly, judge scores should not be treated as decisive for subjective creative work, and human editorial review is still pending. A closed-ended academic benchmark such as Humanity’s Last Exam may measure a different capability, but it cannot settle subjective creative preference.
Practical takeaways
- Choose a model by the kind of draft you want to revise: horror mechanism, sentimental reveal, punchline, concise brand copy, or lush atmosphere.
- Read several outputs from your own prompt before committing. One call is an observation, not a stable model trait.
- Keep creative judgment human: check whether a joke lands, whether a twist earns its setup, and whether brand language sounds like your business.
- Treat API/gateway behavior separately from any consumer-facing product experience.

Editorial illustration; it frames creative variation and is not test evidence.
Limitations
This comparison covers one approved run on 2026-07-17, three prompts, eight models and one selected call per model-prompt pairing. It used no web search and does not establish general quality, reliability, latency, cost or consumer-product behavior. Creative quality remains subjective, and human editorial review is pending.
Explore the Models
Explore the current catalog on OrcaRouter Models.
This evaluation was run through OrcaRouter. The author works with OrcaRouter; model access does not imply affiliation with, endorsement by, or sponsorship from model providers.
Model names and logos are used descriptively. All trademarks belong to their respective owners.
Sources
- Humanity’s Last Exam — Center for AI Safety and Scale AI collaborators; retrieved 2026-07-17.




