Tencent Drops Hy4 Preview. Edges GLM-5.3 And Kimi K3. The Margin Is A Tenth Of A Point.
Tencent released a preview of Hy4, an open-source model it claims outperformed GLM-5.3 and Kimi K3 in internal blind testing. The evaluation involved 163 experts and 203 engineering tasks, with Hy4 scoring 2.99 out of 4.00 against 2.92 for GLM-5.3 and 2.94 for Kimi K3. Tencent conducted the evaluation itself. Read that sentence again.
The principle is self-evaluation bias. When a company runs its own benchmark and wins by 0.05 points, the result tells you less about model quality and more about competitive positioning. The mental model: always ask who administered the test. Independent benchmarks matter. Internal ones are marketing with decimals.
Tencent released the Hy4 preview as open-source, claiming it scored 2.99 out of 4.00 in a blind evaluation of 203 engineering tasks judged by 163 experts. Z.ai's GLM-5.3 scored 2.92 and Moonshot's Kimi K3 scored 2.94 in the same internal Tencent evaluation.
- Open Hugging Face and search for open-source models from Chinese AI labs. You will find several available for free download or browser-based inference.
- Pick two models and run the same coding prompt through each, such as writing a Python function to sort a dictionary by values.
- Compare the outputs side by side. You are now doing what Tencent did, except honestly. Notice how small the quality differences often are between competent models.