fix(deepseek-harness): user turns are aborted ~3 ms after turn/start once the bundle is loaded

#2437 · open · 1 comments

View on GitHub ↗

a1554154492-dev

## Pre-submission checklist - [x] I have searched existing issues and this hasn't been mentioned before - [x] I have read the project documentation and confirmed this issue doesn't already exist - [x] This issue is specific to MemOS and not a general software issue --- ## Bug Description | 问题描述 Once `@memtensor/memos-local-plugin` is present in the DSH profile's `dsh.profile.bundles`, **every direct-user turn is aborted ~3 ms after `turn/start`**: * no `step/start` is ever emitted, and * no `request/header` is ever written (no model request is assembled at all), so the UI never produces a reply. The host process itself stays alive and idle: main-thread CPU over a 4 s sample was **0.14 s**, no crash log, no `CrashDumps` entry, no WER/Application-Error event. The practical symptom is "DSH froze / cannot converse", but the host is not frozen — the turn is simply killed before it starts. The automatic recall is **not** the cause: with the plugin loaded, the recall runs and completes normally *after* the turn has already ended (`[core.retrieval] done reason="turn_start" … totalMs=2710`), so the trigger is earlier than any memory work. Correlation is clean: in a single DSH session with 11 completed turns, **all 4 turns that ran with the plugin loaded were interrupted with 0 steps**, and **all 7 turns with the plugin removed completed normally**. Expected: with `recallEnabled: true` the turn should proceed to step 1 and receive the source-labelled recall context in the same model request. ## How to Reproduce | 如何重现 1. DSH desktop 0.1.7-rc.2, add `@memtensor/[email protected]` to the profile `dsh.profile.bundles`, with a deliberately conservative config (write it while DSH is closed — never hot-load): ```yaml - id: memos-local-memory config: enabled: true recallEnabled: true captureEnabled: false viewerEnabled: false toolsEnabled: true failOnStartupError: false ``` 2. Fully quit DSH, start it again (clean start), wait for the normal boot. 3. Send one short direct message in the GUI (e.g. `test`). 4. Inspect the session log (JSONL) — the turn dies immediately: ```json {"type":"agent/inbox/spliced","data":{"target":"next-turn","inserted":[{"content":[{"type":"text","text":"test"}],"source":{"kind":"user"}}]}} {"type":"turn/start","data":{"turn":11}} {"type":"agent/inbox/spliced","data":{"target":"next-turn","removedCount":1,"inserted":[]}} {"type":"turn/end","data":{"turn":11,"reason":{"kind":"interrupted"}}} ``` (timestamps: `turn/start` at `.840`, `turn/end` at `.843` — 3 ms apart) 5. Remove `@memtensor/memos-local-plugin` from `dsh.profile.bundles`, restart, and send the same message: the turn runs (`step/start` → `request/header` → reply within ~1 s). The same happens with `captureEnabled: true` (+ all LLM stages disabled) and with the plugin's default `--full` config, and it also happens when the bundle is added while DSH is running (hot-load). Only removing the bundle restores normal turns. ## Environment | 环境信息 * DSH (DeepSeek Harness) **0.1.7-rc.2**, desktop build — Electron **44.0.0**, Node **24.18.1**, Windows 11 x64 * `@memtensor/memos-local-plugin` **2.0.19** (npm; adapter files verified byte-identical to **2.0.20**, and `adapters/deepseek-harness` on repo `main` is identical to tag `v2.0.34`) * `better-sqlite3` **13.0.3** (N-API prebuild, loads fine in Electron), `@huggingface/transformers` **4.2.0**, `onnxruntime-node` **1.24.3** * Memory home on a normal local profile; `memos.db` ≈ 23 MB (735 traces / 64 episodes / 36 policies / 7 skills); boot log: `migrations.summary total=13 applied=0 skipped=13` * Local embeddings (`Xenova/all-MiniLM-L6-v2`) work: `loading` → `ready durationMs=548`, then embedding calls complete in 7–382 ms * Config: `recallEnabled: true`, `captureEnabled: false`, `viewerEnabled: false`; algorithm LLM stages left at defaults (all `true`) * Launch used for the capture: host stdout/stderr redirected to files, plus `--inspect` (V8 inspector) and `--enable-logging` ## Additional Context | 其他信息 ### Turn-by-turn correlation (one DSH session, same profile) | turn | plugin | result | steps | |---|---|---|---| | 1, 4, 5, 8, 9, 10, 12 | removed | `completed` | 57, 15, 26, 7, 19, 17, 5 | | 2 | loaded (recall-only, clean start) | `completed` | 13 — *single historical exception* | | 3 | loaded (`--full`) | **`interrupted`** | **0** | | 6 | loaded (capture-lite, hot-load) | **`interrupted`** | **0** | | 7 | loaded (recall-only, clean start) | **`interrupted`** | **0** | | 11 | loaded (recall-only, clean start, capture/viewer off) | **`interrupted`** | **0** | ### Plugin host log around the failure (sanitized) ``` <boot> [storage] sqlite.open … wal=true busyTimeoutMs=5000 <boot> [storage.migration] migrations.summary total=13 applied=0 skipped=13 <boot> [core.pipeline.bootstrap] hostLlmBridge.registered id="deepseek-harness.host.v1" <boot> [core.pipeline.bootstrap] pipeline.ready home="<MEMOS_HOME>" <boot> dsh web: http://127.0.0.1:<port>/?token=<redacted> t+20.880s [embedding.local] loading model="Xenova/all-MiniLM-L6-v2" ← 40 ms AFTER the turn was already dead t+21.428s [embedding.local] ready durationMs=548 t+23.587s [core.retrieval] done reason="turn_start" … tier2=12 tier3=3 kept=1 totalMs=2710 (no further output) ``` So: the turn ends first, the recall starts shortly afterwards and completes within its 3 000 ms budget. Nothing errors, nothing throws visibly. ### Host state while "frozen" (sampled at failure time) ``` MAIN-HOST cpu_delta_4s=0.14s threads=21 handles=549 ws_mb=178 renderer cpu_delta_4s=0.03s gpu cpu_delta_4s=0s ``` V8 inspector `Debugger.pause` captured a **single native frame** (`#0 onStreamRead @ (native)`), i.e. the main thread was parked in the event loop with **no JS frame running** — consistent with an awaited promise that never settles / a lifecycle decision that ends the turn, not with synchronous blocking or a busy loop. ### What we already ruled out (with measurements) * **Crash** — no crash logs, no `CrashDumps`, no WER/Application-Error events. * **Slow capture pipeline** — with all LLM stages disabled (`synthReflections`, `llmScoring`, `useLlm` ×4, `llmFilterEnabled`) every engine call was ≤ 1.4 s; the turn is still interrupted, and it is interrupted *before* any of them run. * **Slow automatic recall** — `deadline.ts` is correct (timer + abort + `Promise.race`, 3 000 ms cap); the recall finished in 2 710 ms, i.e. after the turn had already ended. * **Database / FTS cost** — measured on a copy of the real DB: a 500-char × 6-column `trigram` FTS insert costs **~0.4 ms**; a `MATCH` query 0 ms. * **Other DSH plugins** — in the affected profile, no plugin other than memos registers `agent/pre-step` (the installed circuit breaker only hooks tool calls via `inject=["tools"]`). * **Storage layer** — `[email protected]` (N-API) resolves and runs inside the host; `PRAGMA integrity_check` on the DB returns `ok`. * **Installation hygiene** — profile bundle list, `overrides: better-sqlite3: 13.0.3` and all migrations are consistent; `migrations applied=0`. ### Where it points (needs maintainer knowledge of DSH's pre-step contract) `adapters/deepseek-harness/index.ts` registers the only lifecycle hook: ```ts ctx.on("agent/pre-step", async (payload, next) => bridge.beforeStep(payload, next)); ``` and `bridge.beforeStep()` starts with `const decision = await next();`, then performs up to `min(recallTimeoutMs, 3000)` ms of recall before returning an `{ kind: "enter", messages: [...] }` decision. The observed interruption happens **before the recall's first async work** (the embedding model starts loading ~40 ms after `turn/start`, while `turn/end` was already written at +3 ms), so the abort is not produced by the recall result. Questions that would let us fix it on our side: 1. Does DSH require `agent/pre-step` handlers to settle (or at least return the decision) within a strict budget, and does an async/blocking pre-step handler cause `turn/end { kind: "interrupted" }`? 2. If so, is the intended pattern to return the decision first and inject the recall context through a different channel (e.g. `session/event` or a follow-up splice), instead of extending the pre-step waterfall? 3. Is there a supported way to observe *why* a turn ended as `interrupted` (a reason/diagnostic we can log) so adapter authors can self-diagnose? We can run any instrumentation you want on the failing setup (V8 inspector stack dumps at hang time, hook-level tracing in `dist/`, extra env flags) and report back. ### Willingness to Implement | 实现意愿 Happy to test patches and provide full diagnostics (host logs, stack dumps) for whatever instrumentation you need; no problem implementing the adapter-side change if you confirm the intended pre-step contract.

Comments

a1554154492-dev

## Update: root cause found — this is a **session-format v4 violation**, not a turn abort Following up on this report with a corrected analysis. The original core observation — "the turn is aborted ~3 ms after `turn/start`" — is an artifact of crash-repair logging. Here is what actually happens, why it looks like a freeze, and a one-line fix. ### 1. `turn/end {kind:"interrupted"}` is written by crash recovery, not by the live turn `@deepseek-ai/dsh-session` only produces `interrupted` from `openTurnClosers()` (crash recovery). It synthesizes the closer and **reuses the timestamp of the last real event** (`const time = last.time`) with `seq = last.seq + 1`; if a step was still open it synthesizes `step/end` first. In our log the `turn/end interrupted` of turn 11 has `seq = 1184 = last real seq (1183) + 1`, with the *same millisecond* as the preceding `agent/inbox/spliced`. So the turn was never aborted 3 ms later — it never closed at all, and was closed retroactively the next time DSH started. The same pattern reproduces on every failing turn. ### 2. What actually fails: the automatic-recall message cannot be persisted Every failing turn's durable log ends at `agent/inbox/spliced {removedCount:1}` (the pre-step claim). After that: no `step/start`, no `session-log-*/delivery-accepted`, nothing for the 30–311 s until the process was killed. `dist/adapters/deepseek-harness/index.js` builds the recall message like this: ```js source: { kind: "plugin", plugin: "memos-local-memory", form: "recall" } ``` DSH 0.1.7-rc.2 persists session logs in **format v4**, which refuses the retired v3 `plugin` wrapper: ``` @deepseek-ai/dsh-session-format-v3-to-v4: SessionFormatError: format v4 message requires a producer-owned source kind ``` Native DSH plugins use producer-owned kinds (`runtime-context`, `skill-catalog`, `time-context`, `agent-instructions`, `skill-invocation`, …). The v3→v4 migration rewrites `{kind:"plugin", plugin:X}` into `{kind:"plugin:X"}`, so the equivalent legal shape here is `{ kind: "plugin:memos-local-memory", form: "recall" }`. ### 3. Why it only bites when recall returns content The recall message is injected only when recall returns a non-empty context inside the 3 s budget. If recall is empty, or exceeds `min(recallTimeoutMs, 3000)` and the host-side guard fails open (original decision kept), nothing is injected and the turn is fine. That matches the corpus exactly: the single "successful" turn with the plugin loaded had `memos_search dur=3018 ms` (guard fired), while every failing turn had a normal 2.1–2.7 s recall with `kept=1`. ### 4. Why it presents as a freeze instead of a clean error The invalid message enters the in-memory log, then the v4 encoder refuses it; from that point the session's persistence is wedged, so the request/reply are never durably written and the UI shows nothing. On the next start, crash recovery closes the still-open turn with the synthetic `interrupted` — which is what the original report measured as a "3 ms abort". ### 5. Evidence * Codec unit test: `releasedV4SessionFormatCodec.encodeEvent({type:"user/message", source:{kind:"plugin",plugin:"memos-local-memory",form:"recall"}})` throws `format v4 message requires a producer-owned source kind`; the same event with `kind:"plugin:memos-local-memory"` encodes fine. * Isolated reproduction (`dsh headless` with the bundle loaded, namespace aligned so recall hits real memories): with recall returning `kept=1` the process exits with `dsh: format v4 message requires a producer-owned source kind`, and the session file holds 6 events whose last one is exactly `agent/inbox/spliced {removedCount:1}` — the same shape as the desktop failures. * After the one-line fix, the same scenario runs normally: `user/message source={"kind":"plugin:memos-local-memory","form":"recall"}` is persisted, followed by 17 `step/start`, 27 `tool/call`, 18 `delivery-accepted`. ### 6. Suggested fix ```diff source: { - kind: "plugin", - plugin: DEEPSEEK_HARNESS_PLUGIN, + kind: `plugin:${DEEPSEEK_HARNESS_PLUGIN}`, form: "recall", }, ``` `dist/adapters/deepseek-harness/host-llm.js:220` uses the same `kind:"plugin"` wrapper for host-LLM user messages. It does not reach a session log today, but it should be normalized the same way. Environment: DSH 0.1.7-rc.2 (Electron 44 / Node 24.18.1), `@memtensor/memos-local-plugin` 2.0.19. The DSH adapter is byte-identical in 2.0.20, so 2.0.20 is affected as well. Please consider retitling: this is a format-v4 compatibility bug, not a turn-lifecycle abort. Happy to send a PR for the one-line change. --- **中文摘要**:原报告里的「turn/start 后 3ms 被中断」是崩溃恢复补记的合成事件(时间戳复用最后一条真实事件),不是真的中断。真正原因是自动召回注入的 user 消息用了 v3 已废弃的 `source.kind:"plugin"`,被 DSH 的 session format v4 编解码器硬拒绝(`format v4 message requires a producer-owned source kind`);这条消息写不进会话日志,从此整个会话持久化卡死,重启后被补记成 interrupted。只在「召回在 3s 内返回非空内容」时触发。修复:把 source 改成 `kind:"plugin:memos-local-memory"` 并去掉 `plugin` 字段(一行)。