OpenAI Hits 750 Tokens Per Second. Speed Is Not Intelligence. Read The Footnotes.
OpenAI previewed an "Ultrafast" mode for GPT-5.6 Sol, claiming up to 14x faster processing than standard, reaching 750 output tokens per second. DeepMind is separately working on sign-language access improvements. The safety discussion notes that models do not need malicious intent to cause harm. They need only the wrong permissions and an imperfect understanding of user intent.
This illustrates the principle of capability asymmetry. A faster model amplifies both correct outputs and misinterpretations equally. The mechanism here is permission scope. When a model processes information faster than a human can review it, the window for catching errors shrinks proportionally. Speed without guardrails is not progress. It is acceleration toward whatever cliff is nearest.
OpenAI published the Ultrafast preview at openai.com/index/previewing-ultrafast. DeepMind is pursuing sign-language accessibility. Meta is separately bringing capable models to local hardware.
- Open ChatGPT and ask it a multi-part question with three sub-tasks. Time how long the full response takes to appear on your screen.
- Ask the same question again but request the answer in a numbered list format. Compare the time and notice how output structure affects perceived speed.
- Read your first answer carefully and flag any sentence where the model made an assumption you did not ask for. That gap between what you meant and what it did is the exact risk profile that speed amplifies.