eerwitt/qwentext

Qwentext is a hackathon project for Track 1: MemoryAgent in the Global AI Hackathon Series with Qwen Cloud. The project demonstrates long-term memory for Qwen agents without relying on lossy semantic chunk RAG as the main memory layer.

★ 1Forks 0PythonGitHub ↗Compare

README

Qwentext

License: MIT Python 3.11+ Tests: unittest

Qwentext is a hackathon project for Track 1: MemoryAgent in the Global AI Hackathon Series with Qwen Cloud. The project demonstrates long-term memory for Qwen agents without relying on lossy semantic chunk RAG as the main memory layer.

Qwentext is not a bigger prompt and not a vector scrapbook. It is compact durable memory state where newer evidence can supersede stale facts while preserving the audit trail that explains why the current branch is trusted.

The core idea is event-sourced durable memory. User interactions are treated as dataclass-backed events, and background workers consolidate those events into a structured, persistent memory tree. The agent uses only the relevant branches of that memory state at answer time, while the background worker resolves contradictions, asks the configured model to connect related branches, ranks evidence by confidence and importance, and lowers uncertainty in high-entropy memory branches.

Winning Features

  • Stale facts are superseded instead of duplicated.
  • Current memory remains compact enough to send only relevant branches to Qwen.
  • Previous values and superseded event ids remain auditable in branch metadata.
  • Background consolidation links related memories and adjusts branch importance.
  • Qwen answer generation receives relevant branch context, and chat also exposes the same answer request without memory context for comparison.
  • Tair stores durable memory, events, search metadata, vector routing hints, and metrics.
  • JSONL traces expose every memory mutation.

Qwentext demo UI with chat, compact memory context, and entropy graph

Branch-level supersession metadata is visible when a graph node is selected:

Selected memory branch showing current value, previous values, superseded events, and resolution metadata

Judge Quick Start

The full demo runs offline in under five minutes, no cloud credentials needed:

python -m pip install -e .
qwentext-trace-sample > sample.jsonl
qwentext-local-e2e sample.jsonl --query "travel"
qwentext-web --host 127.0.0.1 --port 8000

Open http://127.0.0.1:8000, chat "I prefer hiking now" then "Actually I prefer cycling now", and watch the preferences.travel branch supersede the old value with an audit trail. Deeper mechanism documentation lives in docs/design-notes.md.

Proof of Alibaba Cloud Deployment

The cloud runtime is real adapter code, not configuration stubs:

  • src/qwentext/memory.py — TairMemoryStore and RedisTairClient using TairDoc JSON documents, Tair streams, TairSearch metadata indexes, TairVector routing (TVS.CREATEINDEX/TVS.HSET/TVS.KNNSEARCH), and TairTS metrics.
  • src/qwentext/models.py — DashScopeQwenClient, QwenHostedModel, and QwenAnswerModel calling Qwen through Alibaba Cloud Model Studio (DashScope).
  • src/qwentext/embeddings.py — QwenEmbeddingModel using the DashScope text-embedding-v4 OpenAI-compatible endpoint.
  • src/qwentext/runtime.py — the factory that swaps the local adapters for the Alibaba Cloud ones when QWENTEXT_ENV=cloud or TAIR_HOST is set.
  • terraform/main.tf — the VPC, ECS, security group, and Tair infrastructure envelope.

Goal

Build a deployment-ready demonstration of long-term memory that can:

  • preserve user and world facts over long time horizons;
  • update old facts when newer evidence supersedes them;
  • avoid stuffing raw chat history into the prompt;
  • show exactly how each memory mutation happened;
  • run locally for fast iteration and deploy to Alibaba Cloud for judging.

Architecture

The deployed Alibaba Cloud shape is:

Qwentext deployed architecture on Alibaba Cloud with Qwen Code

The diagram uses Alibaba Cloud Draw.io icon assets for Alibaba Cloud, VPC, ECS, ACR, Tair, and RAM/STS so the deployed products are visible at a glance. Qwen Cloud / Alibaba Cloud Model Studio is consumed as a managed DashScope API, while Qwen Code is an optional background-investigation wrapper inside the worker runtime.

The system has two paths.

Fast Path

The fast path handles the interactive user request.

  1. The frontend sends a chat message or a selected Teich-style JSONL agent trace to the Python backend.
  2. The backend loads a compact context slice from the memory store.
  3. In cloud mode, TairSearch branch metadata and TairVector routing hints are used to select candidate branches before the lexical fallback is considered.
  4. Qwen receives the user request plus only the relevant durable memory branches through the hosted answer adapter.
  5. The backend returns the Qwen answer with model name, context branch paths, context size, approximate token count, trace metadata, and degraded-mode status.
  6. The interaction is appended to the event stream for later consolidation.

For cloud runtime requests, the backend does not wait for memory mutation or idle investigation. It returns consolidation: queued; the UI polls the worker and refreshes memory, graph, context, and trace state after consolidation. Deterministic local and comparison paths retain synchronous processing for reproducible tests and evaluation.

Slow Path

The slow path is the background memory worker.

  1. The worker reads pending event records from the message bus.
  2. It maps each event to affected memory branches.
  3. It calculates normalized Shannon entropy over each branch's observed value history, then combines it with evidence volume for investigation priority.
  4. It asks the configured MemoryModel adapter for a structured state delta or investigation plan.
  5. If the model asks for external evidence and QWENTEXT_WEB_SEARCH_ENABLED=true, the backend executes the search tool and calls the model again with the search results.
  6. If QWENTEXT_QWEN_CODE_ENABLED=true, the worker can optionally ask Qwen Code through the Python SDK to inspect project context in plan mode and return the same structured memory-investigation schema.
  7. It applies model-returned branch, relationship, and importance changes to durable memory state through deterministic repository methods.
  8. It records a JSONL trace showing the reasoning, tool call, and resulting mutation.

Alibaba Cloud Design

The target deployment uses Alibaba Cloud services directly relevant to the hackathon.

Capability Service Purpose
Model inference Qwen Cloud / Alibaba Cloud Model Studio Generates responses and state deltas.
Code investigation Qwen Code SDK Optionally inspects project context in plan mode and returns structured MemoryInvestigation data.
Compute ECS Runs the Python backend and background worker.
Memory engine Tair Enterprise Edition Stores durable memory state, streams, search metadata, vector routing hints, and temporal metrics.
Container image Container Registry Enterprise Edition, or an external private registry Provides the image pulled by ECS at startup.
Deployment identity RAM and STS AssumeRole Lets Terraform deploy without committing long-lived credentials.
Networking VPC, VSwitch, security group Isolates backend and Tair.
IaC Terraform Creates and destroys the cloud stack.

The Qwen model itself is not provisioned by Terraform. It is consumed as a managed API. Terraform handles the application infrastructure and passes the Qwen API key into compute through sensitive variables or environment configuration. Do not hard-code API keys, passwords, tokens, or secrets in Terraform files.

Tair Data Model

The Tair-backed adapter uses the following data model:

Tair capability Use
TairDoc / JSON document operations Current consolidated memory state through JSON.SET and JSON.GET.
Streams Interaction events and background queue messages through XADD and XRANGE.
TairVector Real Qwen branch embeddings written with TVS.CREATEINDEX and TVS.HSET for semantic context routing before lexical fallback.
TairSearch Searchable branch metadata hashes indexed with FT.CREATE.
TairTS Temporal memory metrics written with TS.ADD, including event counts and memory branch counts.
exHash / hash fields Branch metadata such as path, confidence, importance, evidence count, and searchable text.

TairMemoryStore and TairMessageBus also keep Redis-compatible string/list mirrors so a developer can run tests against simpler Redis-compatible endpoints without changing the core MemoryAgentService contract. The Alibaba Cloud judging runtime uses TairMemoryStore, TairMessageBus, QwenHostedModel, QwenAnswerModel, QwenEmbeddingModel, RedisTairClient, and DashScopeQwenClient; they expose the same dataclass-facing interfaces and serialize to JSON-compatible dictionaries only at adapter boundaries. Local deterministic setup is documented in LOCAL_DEVELOPMENT.md.

Cloud retrieval embeds branch text with Alibaba Cloud Model Studio text-embedding-v4 through the OpenAI-compatible embeddings API, creates a TVS.CREATEINDEX <index> 128 FLAT COSINE index by default, stores branch vectors with TVS.HSET, and routes answer context with TVS.KNNSEARCH plus TairSearch keyword candidates. Configure it with DASHSCOPE_EMBEDDING_ENDPOINT, DASHSCOPE_EMBEDDING_MODEL, and DASHSCOPE_EMBEDDING_DIMENSIONS; defaults are https://dashscope.aliyuncs.com/compatible-mode/v1/embeddings, text-embedding-v4, and 128. If vector commands or embedding calls fail, the trace reports degraded vector status and falls back to TairSearch or lexical ranking.

QWENTEXT_TAIR_PREFIX is the deployment-level Tair application prefix. The web UI namespace field is different: it sends the selected value as backend user_id for chat, memory state, graph, context, trace playback, imports, worker ticks, and idle investigation. In Tair this means one running deployment can isolate demo memories under keys such as qwentext:users:demo:... and qwentext:users:review-a:... without restarting or changing the Tair prefix.

Core Python code uses dataclass schemas such as WorldState, MemoryBranch, MemoryLink, MemoryEvent, QueueMessage, MemoryMutation, MemoryLinkMutation, MemoryInvestigation, ContextResponse, SessionEvent, TraceMessageEvent, and ModelChangeEvent; JSON dictionaries are kept at serialization boundaries such as Teich JSONL traces, client payloads, Hugging Face dataset rows, hosted model clients, and store adapters.

MemoryBranch serialization includes both value and current_value, plus previous_values, superseded_event_ids, resolution_reason, resolved_by_model, last_updated_event_id, and resolved_at. When a later mutation changes the same path, old current evidence moves into the supersession audit trail instead of being silently overwritten. Judges can inspect these fields in /api/world, graph nodes, and the UI node inspector.

Backend Flow

MemoryAgentService composes the repository, message bus, model adapter, and trace sink:

  1. ingest_trace_jsonl(user_id, text) parses Teich-style JSONL message records into MemoryEvent dataclasses.
  2. Each memory event is appended to the configured MemoryStore.
  3. A QueueMessage is published to the configured MessageBus.
  4. process_all(consumer_id) simulates the background worker and asks the configured model for MemoryMutation records.
  5. Event mutations are applied deterministically to WorldState branches with event ids recorded as evidence.
  6. Idle investigation asks the configured model for MemoryInvestigation records that may include branch mutations, relationship links, importance deltas, or a web search query.
  7. When Qwen Code is enabled, the configured model is wrapped by QwenCodeMemoryModel; it uses qwen-code-sdk only during idle investigation, defaults to permission mode plan, and must return structured memory deltas rather than editing the store directly.
  8. When search is enabled, the backend executes the search tool and sends the returned results back to the model for follow-up integration before applying any external-reference links.
  9. get_context(user_id, query) asks the store for branch candidates first; Tair-backed stores use TairSearch and TairVector routing, while local stores fall back to deterministic lexical ranking.
  10. chat(user_id, message) sends the message and compact context to the configured AnswerModel; it also asks the same adapter for a no-context answer variant so the UI can compare memory impact.
  11. Hosted Qwen calls retry transient 429 and 5xx failures with backoff. If answer generation still fails, the API returns a degraded local answer with the error recorded in the answer trace.
  12. Ingest, processing, investigation, context, and answer generation requests are logged through the trace sink as JSONL-compatible AgentTraceEvent records.

Agent Trace Format

Agent traces are read and written as newline-delimited JSON. This keeps compatibility with Teich-generated Hugging Face datasets and with local text-chat uploads to the backend. Once a JSONL trace is loaded, the backend converts common event groups into dataclasses:

  • session becomes SessionEvent;
  • nested message events become TraceMessageEvent with TraceMessage and TextContent;
  • model_change becomes ModelChangeEvent;
  • less common groups such as turn_context, event_msg, response_item, session_meta, session_info, and thinking_level_change are preserved as AgentTraceEvent payloads until a narrower schema is needed.

Serialize these dataclasses back to JSONL when returning traces to the client or preparing dataset input. Do not pass raw JSON dictionaries through memory processing when a dataclass schema exists.

Frontend

The hosted demo is served by qwentext.web with FastAPI and package-local static assets. It intentionally uses plain HTML, CSS, JavaScript, and canvas so the memory system remains the focus. Run it locally with:

qwentext-web --host 127.0.0.1 --port 8000

The first screen is a three-column app:

  1. Chat:

    • choose a memory namespace, defaulting to demo;
    • type a memory-bearing prompt;
    • see the assistant response;
    • expand native thinking details to inspect the compact context sent with the request;
    • inspect the answer trace showing model name, context branches, context size, approximate token count, provider trace, and degraded-mode status.
  2. Memory context and updates:

    • show the context branches used for the current answer;
    • show stored durable memory branches;
    • show startup cloud health for the memory store and answer adapter;
    • show JSONL memory evolution trace records emitted by the backend.
  3. Entropy graph:

    • show branches as nodes and model-returned relationships as graph edges;
    • load existing durable memory branches on first page load, including Tair-backed state from earlier judging or remote Tair runs;
    • show loading and error status while startup memory is being fetched;
    • select a branch node or stored-branch row to inspect storage metadata, search inclusion, evidence event ids, entropy, and update priority;
    • show a bottom status strip for startup loading, background worker activity, selected nodes, and idle memory investigations;
    • tick the background worker every 10 seconds, then ask the backend to rotate across stored branches only when no queued messages are available;
    • report model-guided branch investigation, relationship linking, importance changes, and external search enrichment when a search adapter is configured;
    • color high-entropy branches as unstable;
    • animate a sprite-backed background agent walking toward a rest point below the current target branch;
    • use the real MemoryAgentService and BackgroundMemoryWorker path, backed by local adapters by default.

The chat panel also includes a Hugging Face trace playback flow, collapsed by default behind a compact top bar so the main judging view stays focused on manual chat, context, and graph evolution. Expand the bar, select a namespace first if you want an isolated review run, enter a dataset repository id or URL, load available .jsonl files through the Hugging Face Hub tree API, select a session from the chosen file, then load it into a fixed playback tray above the chat transcript. Play next sends one trace message at a time through the same memory-event queue used by chat, so judges can watch context, stored branches, traces, and the graph evolve incrementally. Import all remains available for the faster demo path that processes the full selected conversation at once.

The graph layout is topology-driven. It computes connected components from real memory links, ignores synthetic world contains links, attracts nodes in the same component toward a shared area, and pushes unconnected components apart. There are no domain-specific graph clusters for programming, travel, identity, or other branch names.

The frontend should remain a demonstration surface, not a large application. The backend memory behavior is the product.

Python Project Rules

Python code is a package under src/qwentext.

Required practices:

  • write or update tests before changing implementation behavior;
  • use argparse for command-line entry points;
  • prefer dataclasses for Python schemas instead of passing raw dictionaries through core APIs;
  • expose focused qwentext-* CLIs instead of one command with many subcommands;
  • keep external dependencies low;
  • prefer Python standard library modules;
  • use unittest for Python tests;
  • keep logic modular and shared through small reusable modules;
  • use type hints on method and function declarations;
  • use the Python logging module for command and service output;
  • never call print() from code under src/;
  • keep comments minimal and only where they clarify non-obvious behavior;
  • avoid broad exception handling and avoid try/except unless there is a concrete recovery path.

Repository Layout

.
├── AGENTS.md
├── CLAUDE.md
├── LICENSE
├── README.md
├── pyproject.toml
├── src/
│   └── qwentext/
│       ├── __init__.py
│       ├── cli_entropy.py
│       ├── cli_compare_rag.py
│       ├── cli_local_e2e.py
│       ├── cli_trace_sample.py
│       ├── entropy.py
│       ├── memory.py
│       ├── messaging.py
│       ├── models.py
│       ├── qwen_code.py
│       ├── runtime.py
│       ├── schema.py
│       ├── service.py
│       ├── static/
│       │   ├── app.js
│       │   ├── index.html
│       │   └── styles.css
│       ├── web.py
│       ├── worker.py
│       └── traces.py
├── tests/
│   ├── test_adapters.py
│   ├── test_cli.py
│   ├── test_e2e_local.py
│   ├── test_entropy.py
│   ├── test_logging_policy.py
│   ├── test_memory.py
│   ├── test_qwen_code.py
│   ├── test_runtime.py
│   ├── test_traces.py
│   └── test_web.py
└── terraform/
    ├── main.tf
    ├── variables.tf
    ├── outputs.tf
    ├── terraform.tfvars.example
    ├── templates/
    │   └── user_data.sh.tftpl
    └── tests/
        └── qwentext_stack.tftest.hcl

Local Development

See LOCAL_DEVELOPMENT.md for editable installs, deterministic offline adapters, environment variables, local commands, and guidance on when to validate with Qwen plus Tair during development.

CLIs

The package exposes focused argparse entry points.

qwentext-entropy '{"a": 3, "b": 1}'
qwentext-synthesize-evolution --pairs 1000 --seed 42 --dataset ds.json --traces-dir traces
qwentext-generate-evolution --pairs 10 --seed 42 --dataset generated.json --traces-dir generated-traces
qwentext-mine-evolution all_conversations.jsonl --dataset ds.json --traces-dir traces
qwentext-compare-rag --output-dir out --limit 3  # small budget-capped smoke run
qwentext-export-evolution --dataset ds.json --traces-dir traces --parquet evolution.parquet
qwentext-compare-rag --output-dir out --evolution-dataset ds.json --traces-dir traces
qwentext-compare-rag --output-dir out --evolution-dataset ds.json --traces-dir traces --evolution-only --limit 3
qwentext-container
qwentext-compare-rag --output-dir rag-comparison-output
qwentext-local-e2e path/to/sample.jsonl --query "travel"
qwentext-trace-sample
qwentext-web --host 127.0.0.1 --port 8000
qwentext-worker --interval-seconds 2

python -m qwentext.cli_entropy '{"a": 3, "b": 1}'
python -m qwentext.cli_local_e2e path/to/sample.jsonl --query "travel"
python -m qwentext.cli_trace_sample

qwentext-local-e2e wires only local adapters. It imports a JSONL trace, publishes memory-event messages, processes the queue with deterministic local rules, and logs the compact context returned for the query. Add --show-agent-trace to log the generated memory evolution trace.

qwentext-compare-rag runs 20+ scripted scenarios through the real Qwentext cloud runtime (remote Qwen mutations and Tair storage; the command refuses to run with local adapters) against deterministic in-process baselines named chunk_rag_top1, chunk_rag_top3, recency_rag, summary_memory, vector_like_rag, and latest_fact_profile, then writes rag_comparison.csv, chart PNGs (rag_comparison_overall.png, rag_comparison_efficiency.png, rag_comparison_context_size.png, and rag_comparison_tradeoff.png), and rag_comparison_report.md. Latency remains in the CSV but is not charted because remote Qwen/Tair and in-process baselines are not comparable. The CSV is seaborn-ready with columns system,scenario,metric,value; metrics include answer correctness, contradiction handling, context size, retrieved/stored item count, measured persistence after a real store reload, measured retrieval latency, prompt size, stale-fact leakage, current-fact inclusion, and estimated context-cost proxy, with aggregate ALL rows for charting and a model column recording which Qwen model produced each Qwentext row. Pass --evolution-dataset and --traces-dir to also run every memory-evolution pair through all baselines; the CSV gains ALL_<tier> aggregate rows and the Markdown report gains per-tier correctness/leakage tables so the synthetic, manual, and mined tiers can be compared directly. Add --prompt-artifacts to write per-scenario with-context and without-context prompt files. Use --evolution-only with --evolution-dataset to exclude the built-in scripted scenarios. --limit N is applied after this filtering, so a small synthetic smoke test runs only the first N evolution pairs.

Long runs are resumable: each completed Qwentext scenario is appended to <output-dir>/qwentext_cache.jsonl (override with --cache) together with the model that produced it and a timestamp. Stop the process at any point and re-run the same command — cached scenarios are skipped and only the remainder costs tokens. The cache is keyed by scenario and model, so switching DASHSCOPE_MODEL re-evaluates under the new model while keeping the earlier model's results for comparison. Use --limit N first to smoke-test the pipeline on a handful of scenarios before committing a large token budget.

QWENTEXT_MODEL_FALLBACKS accepts a comma-separated ordered model list (for example qwen3.6-flash,qwen3.5-flash,qwen-flash). All Qwen adapters share one ModelRotation: when the active model returns a quota-exhaustion error code (DashScope Throttling.*, Arrearage, or OpenAI-compatible insufficient_quota), every service in the process permanently switches to the next model in the list and continues, and the answer trace plus compare cache record the model actually used.

Model Studio API keys are region-scoped. The remote-Tair setup script uses Singapore endpoints under dashscope-intl.aliyuncs.com. A Beijing key must instead use the corresponding dashscope.aliyuncs.com endpoints. Mixing a key and endpoint from different regions returns HTTP 401.

The memory-evolution dataset supports four provenance tiers; every record carries a source field. qwentext-synthesize-evolution generates thousands of exact-label synthetic pairs from a seeded slot grammar (the generator knows which facts it superseded, so labels are ground truth by construction). qwentext-generate-evolution asks the configured Qwen model to construct realistic earlier/later conversations plus expected and faded phrase labels. These records use source=generated; their labels are model-generated and must not be described as exact construction labels or human ground truth. Generated records include a neutral current-state query. Expected labels cover retained, replacement, and newly added durable facts. Faded labels are never added to that query. Answer-time compact context is capped at 2,000 characters. Each completed pair is saved immediately so a partially completed run remains inspectable. qwentext-mine-evolution mines mined silver-label pairs from consecutive conversations in a local or Hugging Face JSONL trace file using a deterministic keyword diff. manual records come from human curation on the /dataset page.

Evolution pairs are replayed by qwentext-compare-rag --evolution-dataset — trace A, consolidate, trace B, consolidate — through the Alibaba Cloud runtime (remote Qwen mutations and Tair storage; the command refuses to run with local adapters, so the reported numbers describe the same system judges use). The per-tier tables report correctness over expected_words (expected recall) and stale-fact leakage over faded_words (faded leakage). The script runs on your machine with QWENTEXT_ENV=cloud, TAIR_HOST/TAIR_USERNAME/ TAIR_PASSWORD, and DASHSCOPE_API_KEY set.

qwentext-export-evolution converts the dataset plus trace files into one Parquet file ready to push to a Hugging Face dataset repository, and --from-parquet converts a published Parquet back into the local dataset and trace files. See docs/design-notes.md for the schema and tier methodology.

qwentext-web hosts both the HTML site and JSON API. Useful endpoints are /api/chat, /api/events, /api/worker/tick, /api/context, /api/world, /api/graph, /api/trace, /api/health, /api/huggingface/files, /api/huggingface/conversations, /api/huggingface/playback, and /api/huggingface/import. Pass user_id to context, world, graph, snapshot, trace, worker, investigation, chat, playback, and import calls to isolate a namespace.

qwentext-container is the Docker entry point for ECS. It reads QWENTEXT_ROLE; with backend,worker it starts the FastAPI site and a background memory worker in one container. Cloud runtime and optional Qwen Code settings are detailed in LOCAL_DEVELOPMENT.md and the Terraform deployment section below.

Demonstration Strategy

The demo should show one concrete memory evolution:

  1. A user or imported agent trace introduces a fact.
  2. A later event contradicts or refines the fact.
  3. The branch inspector shows previous_values and superseded_event_ids after the stale fact is replaced by compact current memory.
  4. The background worker selects that branch.
  5. Qwen emits a structured state delta.
  6. The memory store mutates durable memory state.
  7. Repeated evidence for the corrected value lowers value-history entropy, and the next answer uses the corrected compact context.

This directly addresses efficient storage and retrieval, timely forgetting, and critical recall within a limited context window.

Why This Beats Vector RAG

The demo is not claiming vectors are useless. It shows why vectors alone are a weak long-term memory substrate: a chunk retriever can recall stale evidence when a later event supersedes it, and it tends to grow prompt context as more history arrives. Qwentext mutates compact durable memory instead: the latest preference or fact becomes current_value, stale values move into previous_values with superseded evidence ids, and context retrieval can send only matching branches.

Generate the comparison artifacts (the script runs on your machine; Qwen and Tair are the remote Alibaba Cloud services):

export QWENTEXT_ENV=cloud  # plus DASHSCOPE_API_KEY, TAIR_HOST, TAIR_USERNAME, TAIR_PASSWORD
qwentext-compare-rag --output-dir rag-comparison-output

Open rag-comparison-output/rag_comparison.csv for raw metrics or rag-comparison-output/rag_comparison_overall.png for the seaborn bar chart. In the contradiction case Qwentext recalls the corrected preference from real Qwen-consolidated durable memory in Tair, while the chunk baseline retrieves the earlier stale chunk.

Deployment Notes

Terraform should create only infrastructure:

  • VPC and subnet;
  • security group;
  • ECS instance for the backend and background worker container;
  • optional Alibaba Cloud Container Registry Enterprise Edition instance, namespace, and repository;
  • Tair instance for durable memory state and events;
  • Tair application account for Redis-compatible access;
  • startup environment shared from Terraform to the container;
  • outputs for public endpoint and non-secret Tair connection metadata.

Terraform must not contain real secrets. Use environment variables such as TF_VAR_dashscope_api_key and provider-supported secret management. The repository should include placeholder variables and documentation, not credentials.

Required deployment inputs include:

export ALICLOUD_ACCESS_KEY="..."
export ALICLOUD_SECRET_KEY="..."
export TF_VAR_assume_role_arn="acs:ram::<account-id>:role/QwentextTerraformDeployer"
export TF_VAR_container_image="ghcr.io/<owner>/qwentext:latest"
export TF_VAR_acr_registry_domain="ghcr.io"
export TF_VAR_acr_registry_user="..."
export TF_VAR_acr_registry_pass="..."
export TF_VAR_dashscope_api_key="..."
export TF_VAR_tair_password="..."

ALICLOUD_ACCESS_KEY and ALICLOUD_SECRET_KEY should belong to the caller RAM user or role that is allowed to call sts:AssumeRole; Terraform then assumes TF_VAR_assume_role_arn and receives the temporary STS security token internally. The assumed role needs permissions for VPC, ECS, Tair/KVStore, and ACR only when acr_enabled=true.

The Terraform defaults are sized for a small single-user demo in Singapore: ap-southeast-1, ap-southeast-1a, ecs.t6-c1m1.large, a 20 GiB ECS system disk, 1 Mbps public outbound bandwidth, and tair.rdb.1g. Override those variables only if the selected zone does not offer that exact ECS or Tair class.

terraform/terraform.tfvars.example contains the non-secret shape for these values. Copy its keys into your own untracked .tfvars file or provide them through TF_VAR_ environment variables.

By default, Terraform does not create ACR. Use a private external registry to avoid the ACR Enterprise Edition subscription cost. Good low-cost options are:

  • GitHub Container Registry (ghcr.io): supports private images and repository permission inheritance through GitHub Packages.
  • GitLab Container Registry (registry.gitlab.com): available on the GitLab Free tier for private projects.
  • Docker Hub: Docker Personal includes one private repository.

For the default private-registry path:

podman build -t qwentext:local-test .
podman tag qwentext:local-test ghcr.io/<owner>/qwentext:latest
podman login --username=<registry-username> ghcr.io
podman push ghcr.io/<owner>/qwentext:latest

export TF_VAR_container_image="ghcr.io/<owner>/qwentext:latest"
export TF_VAR_acr_registry_domain="ghcr.io"
terraform -chdir=terraform apply

ECS uses TF_VAR_acr_registry_user and TF_VAR_acr_registry_pass to log in before pulling the private image. The image defaults to qwentext-container, so container_command can stay empty unless startup needs to be overridden.

If you explicitly need Alibaba Cloud ACR Enterprise Edition, set TF_VAR_acr_enabled=true. Enterprise Edition is subscription-based, so the defaults use the smallest profile configured here: Basic, one month, manual renewal, and provider-valid minimum quotas: 5 namespaces and 1000 repositories. Because ECS must pull an image that already exists, use a two-step flow the first time:

terraform -chdir=terraform apply -target=alicloud_cr_ee_instance.image -target=alicloud_cr_ee_namespace.image -target=alicloud_cr_ee_repo.image
terraform -chdir=terraform output acr_instance_endpoints
terraform -chdir=terraform output -raw acr_registry_domain
terraform -chdir=terraform output -raw acr_image_name

podman build -t qwentext:local-test .
podman tag qwentext:local-test <acr-domain>/<namespace>/<repo>:latest
podman login --username=<registry-username> <acr-domain>
podman push <acr-domain>/<namespace>/<repo>:latest

export TF_VAR_acr_enabled=true
export TF_VAR_acr_registry_domain="<acr-domain>"
terraform -chdir=terraform apply

The ECS user-data template writes /etc/qwentext.env with DASHSCOPE_API_KEY, QWENTEXT_ROLE=backend,worker, BACKEND_HOST=0.0.0.0, BACKEND_PORT, and TAIR_HOST, TAIR_PORT, TAIR_USERNAME, TAIR_PASSWORD, and TAIR_SSL. It then runs the configured container image as a qwentext systemd service using that env file.

Remote Tair Demo Mode

For a lower-cost live demo, disable ECS and run the app on your laptop while Terraform provisions only remote Tair and supporting network resources:

export TF_VAR_ecs_enabled=false
export TF_VAR_local_tair_connection_enabled=true
export TF_VAR_local_client_ip="$(curl -s https://ifconfig.me/ip)"
terraform -chdir=terraform apply

Then point the local app at remote Tair and Qwen:

export QWENTEXT_ENV=cloud
export BACKEND_HOST=127.0.0.1
export BACKEND_PORT=8000
export DASHSCOPE_API_KEY="$TF_VAR_dashscope_api_key"
export DASHSCOPE_EMBEDDING_MODEL="text-embedding-v4"
export DASHSCOPE_EMBEDDING_DIMENSIONS=128
# Optional for Singapore/Hong Kong workspaces: set the Model Studio compatible-mode embeddings endpoint.
# export DASHSCOPE_EMBEDDING_ENDPOINT="https://dashscope.aliyuncs.com/compatible-mode/v1/embeddings"
export TAIR_HOST="$(terraform -chdir=terraform output -raw local_tair_host)"
export TAIR_PORT="$(terraform -chdir=terraform output -raw tair_port)"
export TAIR_USERNAME="$(terraform -chdir=terraform output -raw tair_account_name)"
export TAIR_PASSWORD="$TF_VAR_tair_password"
export TAIR_SSL=false
export TAIR_TIMEOUT_SECONDS=5
# Optional: isolate this demo from earlier persisted Tair runs.
export QWENTEXT_TAIR_PREFIX="qwentext-local-$(date +%Y%m%d%H%M)"
qwentext-container

Open http://127.0.0.1:8000. When finished, run terraform -chdir=terraform destroy to remove Tair and networking resources. If you reuse the default prefix and the default UI user demo, Tair will keep older durable memory branches between app restarts; those branches remain visible as stored memory, while the context sent to the answer path is filtered to the branches relevant to the current query when possible. Startup memory loading reads the Tair memory-state key once through /api/snapshot. TAIR_TIMEOUT_SECONDS bounds Tair connect and read calls so a stale external connection shows an error instead of leaving the UI waiting indefinitely.

Terraform tests live in terraform/tests/*.tftest.hcl and use Terraform's native test runner with a mocked Alicloud provider. Run them with:

terraform -chdir=terraform init -backend=false
terraform -chdir=terraform test

The repository .gitignore excludes Terraform state, plans, local .tfvars files, local environment files, Python caches, and common key material. Keep only placeholder examples such as terraform/terraform.tfvars.example under version control.

Useful Terraform outputs are:

  • backend_url;
  • backend_public_ip;
  • acr_instance_id;
  • acr_instance_endpoints;
  • acr_registry_domain;
  • acr_image_name;
  • container_image;
  • tair_instance_id;
  • tair_connection_domain;
  • tair_port;
  • tair_account_name.

backend_url is the public hosted HTML demo and API base URL. Open it in a browser after terraform apply completes.

License

Qwentext is released under the MIT License. See LICENSE.

Contributors

eerwitt

Issues