claude -p · streaming · thinking

把「终端 Claude 的真思维链」接进网页,并做成流式

在自己的网页里调用本机的 claude -p(终端 Claude Code 本人,走你自己的订阅), 让回答一个字一个字流式冒出来;并把它开口前的「思维链」(扩展思考)也单独抠出来、流式显示。 全套在 macOS + FastAPI + 原生 JS 上实测跑通,每段代码都能直接抄。

—— 作者 · 离

FastAPISSE stream-jsonthinking_delta 原生 JS

0一张图看懂架构

浏览器  ⇄  (SSE 流)  ⇄  你的后端(FastAPI)  ⇄  subprocess:
claude -p --output-format stream-json

浏览器发消息给后端;后端起一个 claude -p 要它用 stream-json 边想边吐; 后端把这股 JSON 流逐块转成 SSE(text/event-stream)推给浏览器;浏览器边读边渲染。

为什么要个后端中转?因为 claude -p 是本机命令、走你机器上登录的 Claude 订阅,浏览器没法直连。 好处是 不花 API / 中转的钱,而且它带着你本机的上下文。

1后端:把 claude -p 的流转成 SSE

三个必加的命令行参数:--output-format stream-json、--verbose、--include-partial-messages。 少了 --include-partial-messages 就拿不到逐字增量,更拿不到思维链。

app.pyimport json, os, subprocess
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from pydantic import BaseModel

app = FastAPI()
CLAUDE_BIN = "/opt/homebrew/bin/claude"   # which claude

class Req(BaseModel):
    prompt: str = ""
    system: str = ""
    rawThink: bool = True   # 是否想看「真思维链」(开=给思考预算+触发词,慢半拍)

@app.post("/cc-real")
def cc_real(inp: Req):
    env = dict(os.environ)
    # subprocess 默认环境很干净,常找不到 node/claude,手动补 PATH、HOME
    env["PATH"] = "/opt/homebrew/bin:/usr/local/bin:" + env.get("PATH", "/usr/bin:/bin")
    env["HOME"] = os.path.expanduser("~")
    # 想看思维链:给思考预算 + 在用户消息尾巴塞触发词(只抬概率,强不来,见「坑 2」)
    if inp.rawThink:
        env["MAX_THINKING_TOKENS"] = "4000"
    user = inp.prompt + ("\n\n(开口前先 ultrathink,认真想透再说。)" if inp.rawThink else "")

    def gen():
        proc = subprocess.Popen(
            [CLAUDE_BIN, "-p",
             "--output-format", "stream-json",
             "--verbose",
             "--include-partial-messages",
             "--append-system-prompt", inp.system,
             user],
            cwd=env["HOME"], env=env,
            stdout=subprocess.PIPE, stderr=subprocess.DEVNULL,
            text=True, bufsize=1,            # 行缓冲,配合逐行读
        )
        try:
            for line in proc.stdout:         # 逐行读 claude 吐的 JSONL
                line = line.strip()
                if not line:
                    continue
                try:
                    j = json.loads(line)
                except Exception:
                    continue
                # 增量都包在 stream_event.event.content_block_delta 里
                if j.get("type") != "stream_event":
                    continue
                ev = j.get("event") or {}
                if ev.get("type") != "content_block_delta":
                    continue
                d = ev.get("delta") or {}
                # ① 正文:text_delta
                if d.get("type") == "text_delta" and d.get("text"):
                    yield "data: " + json.dumps({"delta": d["text"]}, ensure_ascii=False) + "\n\n"
                # ② 真思维链:thinking_delta(模型扩展思考的原生事件,没整理过的脑内)
                elif d.get("type") == "thinking_delta" and d.get("thinking"):
                    yield "data: " + json.dumps({"think": d["thinking"]}, ensure_ascii=False) + "\n\n"
        finally:
            try:
                proc.terminate()
            except Exception:
                pass
            yield "data: " + json.dumps({"done": True}) + "\n\n"

    return StreamingResponse(
        gen(),
        media_type="text/event-stream",
        headers={
            "X-Accel-Buffering": "no",   # ★ 关键!见「坑 1」:让 nginx 别缓冲这条流
            "Cache-Control": "no-cache",
        },
    )
两类「思维」别混

thinking_delta =模型真·扩展思考,没替谁整理过的原始推理(本教程主角)。

你在 system 里要它在正文写 <思绪>…</思绪> =它写给人看的内心独白, 是普通 text_delta。两者可同时存在,渲染时分开显(见第 3 节)。

2前端:边读边显(解析 SSE 流)

原生 fetch + ReadableStream 就够,不用任何库。要点:自己按 \n 切行、认 data: 前缀、 把 think 和 delta 分别累加,各自回调出去。

app.jsasync function askCcReal(prompt, system, { onBody, onThink } = {}) {
  const res = await fetch("/cc-real", {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ prompt, system, rawThink: true }),
  });
  if (!res.ok) throw new Error("接不通 " + res.status);

  const reader = res.body.getReader();
  const dec = new TextDecoder();
  let buf = "", body = "", think = "";

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    buf += dec.decode(value, { stream: true });

    let nl;
    while ((nl = buf.indexOf("\n")) >= 0) {       // 一行行抠出来
      const line = buf.slice(0, nl).trim();
      buf = buf.slice(nl + 1);
      if (!line.startsWith("data:")) continue;
      let j;
      try { j = JSON.parse(line.slice(5).trim()); } catch { continue; }
      if (j.think) { think += j.think; onThink && onThink(think); }  // 思维链累加
      if (j.delta) { body  += j.delta; onBody  && onBody(body);  }   // 正文累加
    }
  }
  return { body, think };
}

// 用法:思维链一格、正文一格,各自实时刷新
askCcReal("今天累了,安慰我一下", "你是一个温柔的助手。", {
  onThink: (t) => { thinkBox.textContent = t; thinkBox.hidden = !t; }, // 有就显、没有就藏
  onBody:  (b) => { bodyBox.textContent  = b; },
});

体验小贴士:用 requestAnimationFrame 合帧,别每个增量都直接改 DOM(高频会卡)。 累加的是「目前全文」,每次整段赋值即可,不用自己拼。

3(可选)把内联 <思绪> 标签边流边折叠

如果你也让模型在正文里写 <思绪>…</思绪>,难点是:流式时闭合标签还没到。 解法是支持「开标签已出现、还没闭合」的中间态——把开标签后面的内容先当思绪流着显,闭合了再定稿。

function splitReply(text) {
  text = text || "";
  // 已闭合:正常切出思绪 / 正文
  const closed = text.match(/<(思绪|思考|thinking)>([\s\S]*?)<\/\1>/);
  if (closed) return { think: closed[2].trim(),
                       body: text.replace(closed[0], "").trim(), open: false };
  // 只有开标签(还在流):开标签后面整段先当思绪
  const open = text.match(/<(思绪|思考|thinking)>/);
  if (open) return { think: text.slice(open.index + open[0].length),
                     body: text.slice(0, open.index).trim(), open: true };
  return { think: "", body: text, open: false };
}

坑 1 · 公网上 nginx 把 SSE 憋成一坨(看着「不流式」)

🩹 现象 → 根因 → 修法

现象:本机直连流式好好的,一挂到公网域名(前面有 nginx 反代)就变成「整段啪一下蹦出来」。

根因:nginx 默认 proxy_buffering on,把整条响应缓冲到结束、再一坨吐给客户端,SSE 当场报废。

修法(不用改 nginx 配置):在后端这条响应上加响应头 X-Accel-Buffering: no, nginx 认这个头、会对这条响应单独关掉缓冲。(第 1 节代码里已加。)

自测:curl 掐时间戳,对比本机 vs 公网——

curl -sN -X POST http://本机:端口/cc-real \
  -H 'Content-Type: application/json' -d '{"prompt":"测试流式"}' \
| while IFS= read -r l; do [ -n "$l" ] && date "+%H:%M:%S  $l"; done

# 本机:每行隔着几百毫秒,一股一股  = 流式正常
# 公网:所有行挤在最后一瞬间全到    = 被反代缓冲了 → 加 X-Accel-Buffering: no

坑 2 · claude -p 的扩展思考「强不来」(最重要的大实话)

⚠️ 别指望 100% 每次都有思维链

实测:MAX_THINKING_TOKENS 环境变量、prompt 里的 think / think hard / ultrathink 关键词, 都只是「抬高概率」,不是开关。

同一句话连发,可能这几次每次都想(几大块 thinking_delta),下几次一块都没有、直接开口。 一个常见原因:你若同时要求它写 <思绪>,模型常把「写思绪」当成思考、就跳过了真扩展思考(两者抢同一件事)。 重度使用撞限流降档时,思考也会被勒住。

所以正确做法:思维链有就漂亮地流出来、没有就那一格干脆别显(永不破 UI), 不要在产品里承诺「每次都能看到它的思考」。把 rawThink 做成开关: 开=注入预算+触发词(慢半拍、更可能出),关=不强制(回得快)。

坑 3 · 拿不到 thinking_delta?检查这几样

想确认到底吐了哪些事件类型,跑这条自检:

claude -p --output-format stream-json --verbose --include-partial-messages \
  "随便说句话,先 ultrathink 想想" 2>/dev/null \
| python3 -c "import sys,json
for l in sys.stdin:
  try: j=json.loads(l)
  except: continue
  if j.get('type')=='stream_event':
    d=j.get('event',{}).get('delta',{})
    if d.get('type'): print(d['type'])" | sort | uniq -c
# 看到 thinking_delta 计数 > 0,就说明这次它真的思考了

安全 & 注意

一句话总结

claude -p --output-format stream-json --verbose --include-partial-messages 把字和思考逐块吐出来, 后端按 text_delta / thinking_delta 分两路转成 SSE(记得 X-Accel-Buffering: no), 前端 getReader() 边读边显、思维链单独一格。思维链能不能出,看模型当下,强不来——有就显、没有就藏。