<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>国产算力 on 繁星物语</title><link>https://sql668.github.io/blog/tags/%E5%9B%BD%E4%BA%A7%E7%AE%97%E5%8A%9B/</link><description>Recent content in 国产算力 on 繁星物语</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Fri, 18 Sep 2026 12:45:00 +0800</lastBuildDate><atom:link href="https://sql668.github.io/blog/tags/%E5%9B%BD%E4%BA%A7%E7%AE%97%E5%8A%9B/index.xml" rel="self" type="application/rss+xml"/><item><title>GLM 自建推理栈：把「3× 吞吐、10 万国产卡、两周上线」拆开核一遍</title><link>https://sql668.github.io/blog/posts/glm-inference-stack-audit/</link><pubDate>Fri, 18 Sep 2026 12:45:00 +0800</pubDate><guid>https://sql668.github.io/blog/posts/glm-inference-stack-audit/</guid><description>&lt;p&gt;智谱 2026-09-17 在 z.ai 官方博客发了《Toward Recursive Self-Improvement: How GLM Built Its Own Inference Infrastructure》，讲 GLM-5.3-Flash 怎么在十万张国产加速器上自建推理栈。中文圈当天就转开了，但对不齐的不是媒体、是口径：官方中文博客稿写「端到端服务性能约 3 倍」「部分场景的性能差距超过 20%」，唐杰 X 写「3.2× 端到端吞吐」「over 30% to under 1%」。两家转载又都同时用了两套数——量子位导语跟唐杰用 3.2×、正文转官方中文博客稿用 3×；爱范儿同一篇里 3.2× 与约 3 倍并存。同一件事，两套数。&lt;/p&gt;</description></item></channel></rss>