t3po-ws is a local C++ streaming text translation server for the
Confucius4-T3PO GGUF model.
It follows the upstream interleaved-history protocol and exposes a small
WebSocket interface suitable for desktop applications.
The executable loads the model once with llama.cpp. It does not require Python, PyTorch, transformers, vLLM, or a separate model server at runtime.
- Chinese-to-English and English-to-Chinese streaming translation
- Append-only translation segments with the upstream EOS-as-WAIT policy
- Upstream
low,native, andhighlatency operating points - Interleaved source/translation history for stable context
- Session-scoped terminology constraints
- A compact text control protocol on
/ws/translate - Optional local bearer token in the WebSocket query string
- Metal, CUDA, Vulkan, and CPU llama.cpp backends
T3PO is text-to-text. This server intentionally does not accept PCM audio.
For speech translation, connect a streaming ASR service such as r2t2-ws
and forward its committed text to this server.
Download a quantized model from the official Confucius4-T3PO-GGUF repository. Keep model weights outside this source repository.
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j
ctest --test-dir build --output-on-failureFor an offline build, reuse source checkouts at the exact revisions listed in
THIRD_PARTY_NOTICES.md:
cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Release \
-DT3PO_LLAMA_CPP_SOURCE_DIR=/path/to/llama.cpp \
-DT3PO_IXWEBSOCKET_SOURCE_DIR=/path/to/IXWebSocket \
-DT3PO_NLOHMANN_JSON_SOURCE_DIR=/path/to/jsonMetal is enabled by default on macOS. On Linux, use
-DT3PO_ENABLE_CUDA=ON or -DT3PO_ENABLE_VULKAN=ON when the corresponding
toolchain is installed. Pass -DT3PO_ENABLE_NATIVE=ON for a CPU build tuned
to the current machine instead of a portable binary.
brew tap ZPVIP/t3po https://github.com/ZPVIP/T3PO-ws.git
brew install --HEAD ZPVIP/t3po/t3po-wsThe Formula installs the executable but not the model weights.
./build/t3po-ws \
--model /path/to/Confucius4-T3PO-Q4_K_M.gguf \
--host 127.0.0.1 \
--port 8273 \
--auth autoThe process writes a machine-readable readiness object to stdout:
{
"authToken": "generated-token",
"endpoint": "/ws/translate",
"event": "ready",
"host": "127.0.0.1",
"port": 8273,
"url": "ws://127.0.0.1:8273/ws/translate?token=generated-token"
}Connect to /ws/translate, then initialize the session:
{"type":"init","direction":"zh2en","latency_mode":"native"}The optional terms array contains session terminology:
{
"type": "init",
"direction": "zh2en",
"latency_mode": "low",
"terms": [{"src":"NetEase Youdao","trg":"NetEase Youdao"}]
}Send source increments as text messages:
{"type":"text","text":"Good"}
{"type":"text","text":"morning,"}
{"type":"text","text":"everyone."}The server responds with a wait event when the model emits EOS without
translation, or an append-only translation event when text is committed:
{"type":"wait","source_units":2}
{"type":"translation","source":"Good morning","text":"Good morning","source_units":2}Finish and force the remaining buffered source text with:
{"type":"end"}The server emits any final translation, a metrics event, and an ended event
before closing normally. Only one session is accepted at a time because a
single llama.cpp context owns the model's active decode state.
The Node.js example reads one source chunk per stdin line:
printf 'Good\nmorning,\neveryone.\n' | \
node examples/node-client.mjs ws://127.0.0.1:8273/ws/translate en2zh native| Mode | Tau | Behavior |
|---|---|---|
low |
0.9375009536743164 |
Suppresses terminal logits and commits earlier |
native |
0 |
Uses the model's native decision |
high |
-0.39 |
Raises terminal logits and waits longer |
The bias is applied only to non-forced decisions. Final flushes require at least one generated token so buffered source text is not intentionally left untranslated.
The server source is distributed under the GNU General Public License v3.0. Model weights remain subject to the model publisher's license.