<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>延迟预算 on 繁星物语</title><link>https://sql668.github.io/blog/tags/%E5%BB%B6%E8%BF%9F%E9%A2%84%E7%AE%97/</link><description>Recent content in 延迟预算 on 繁星物语</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 16 Sep 2026 14:46:00 +0800</lastBuildDate><atom:link href="https://sql668.github.io/blog/tags/%E5%BB%B6%E8%BF%9F%E9%A2%84%E7%AE%97/index.xml" rel="self" type="application/rss+xml"/><item><title>System One Models 与 Jev：不生成字符串的模型，凭什么快两个数量级</title><link>https://sql668.github.io/blog/posts/system-one-jev-typed-output/</link><pubDate>Wed, 16 Sep 2026 14:46:00 +0800</pubDate><guid>https://sql668.github.io/blog/posts/system-one-jev-typed-output/</guid><description>&lt;p&gt;请求路径里放一个模型做判断——这笔转账要不要放行、这条工单要不要升级。现在的做法是发 prompt、解析返回的 JSON：单次请求几秒到几分钟，输出 token 按字计费，而 SLO 是五百毫秒。这个差距调参补不回来。&lt;/p&gt;</description></item></channel></rss>