Eleven Models, One Coffee Shop Prompt. The Results Differ Wildly. Choice Is Not The Same As Quality.
Netlify wired up OpenRouter to support eleven different AI models, including Kimi K3, GLM 5.2, and DeepSeek V4. They fed each the identical prompt: build a one-page coffee shop site with hours, address, menu, and a photo, explicitly hinting that no CMS was needed. The outputs varied enough to warrant a detailed comparison, with the team also steering models away from what they call the dreaded purple AI slop via built-in UI design guidance.
This is a live demonstration of what I would call prompt determinism, or rather the lack thereof. The same instruction across eleven models produces eleven interpretations, which teaches you that the model is not a vending machine. It is a negotiator. Understanding that the choice of model shapes the output as much as the prompt itself is the single most important mental shift for anyone trying to use these tools deliberately rather than superstitiously.
Netlify ran the comparison using OpenRouter as the routing layer, testing models including Kimi K3, GLM 5.2, and DeepSeek V4 on an identical coffee shop site prompt.
- Go to openrouter.ai and create a free account, which gives you access to dozens of models through a single interface.
- Type the coffee shop prompt from the story verbatim into the chat box, then send it to one model and note the result.
- Switch to a different model from the dropdown, paste the identical prompt again, and compare the two outputs side by side. You will immediately see what Netlify saw. The model is a variable, not a constant.