Qwentext is a hackathon project for Track 1: MemoryAgent in the Global AI Hackathon Series with Qwen Cloud. The project demonstrates long-term memory for Qwen agents without relying on lossy semantic chunk RAG as the main memory layer.
Qwentext is not a bigger prompt and not a vector scrapbook. It is compact durable memory state where newer evidence can supersede stale facts while preserving the audit trail that explains why the current branch is trusted.
The core idea is event-sourced durable memory. User interactions are treated as dataclass-backed events, and background workers consolidate those events into a structured, persistent memory tree. The agent uses only the relevant branches of that memory state at answer time, while the background worker resolves contradictions, asks the configured model to connect related branches, ranks evidence by confidence and importance, and lowers uncertainty in high-entropy memory branches.
- Stale facts are superseded instead of duplicated.
- Current memory remains compact enough to send only relevant branches to Qwen.
- Previous values and superseded event ids remain auditable in branch metadata.
- Background consolidation links related memories and adjusts branch importance.
- Qwen answer generation receives relevant branch context, and chat also exposes the same answer request without memory context for comparison.
- Tair stores durable memory, events, search metadata, vector routing hints, and metrics.
- JSONL traces expose every memory mutation.
Branch-level supersession metadata is visible when a graph node is selected:
The full demo runs offline in under five minutes, no cloud credentials needed:
python -m pip install -e .
qwentext-trace-sample > sample.jsonl
qwentext-local-e2e sample.jsonl --query "travel"
qwentext-web --host 127.0.0.1 --port 8000Open http://127.0.0.1:8000, chat "I prefer hiking now" then "Actually I prefer cycling now", and watch the preferences.travel branch supersede the old value with an audit trail. Deeper mechanism documentation lives in docs/design-notes.md.
The cloud runtime is real adapter code, not configuration stubs:
src/qwentext/memory.py—TairMemoryStoreandRedisTairClientusing TairDoc JSON documents, Tair streams, TairSearch metadata indexes, TairVector routing (TVS.CREATEINDEX/TVS.HSET/TVS.KNNSEARCH), and TairTS metrics.src/qwentext/models.py—DashScopeQwenClient,QwenHostedModel, andQwenAnswerModelcalling Qwen through Alibaba Cloud Model Studio (DashScope).src/qwentext/embeddings.py—QwenEmbeddingModelusing the DashScopetext-embedding-v4OpenAI-compatible endpoint.src/qwentext/runtime.py— the factory that swaps the local adapters for the Alibaba Cloud ones whenQWENTEXT_ENV=cloudorTAIR_HOSTis set.terraform/main.tf— the VPC, ECS, security group, and Tair infrastructure envelope.
Build a deployment-ready demonstration of long-term memory that can:
- preserve user and world facts over long time horizons;
- update old facts when newer evidence supersedes them;
- avoid stuffing raw chat history into the prompt;
- show exactly how each memory mutation happened;
- run locally for fast iteration and deploy to Alibaba Cloud for judging.
The deployed Alibaba Cloud shape is:
The diagram uses Alibaba Cloud Draw.io icon assets for Alibaba Cloud, VPC, ECS, ACR, Tair, and RAM/STS so the deployed products are visible at a glance. Qwen Cloud / Alibaba Cloud Model Studio is consumed as a managed DashScope API, while Qwen Code is an optional background-investigation wrapper inside the worker runtime.
The system has two paths.
The fast path handles the interactive user request.
- The frontend sends a chat message or a selected Teich-style JSONL agent trace to the Python backend.
- The backend loads a compact context slice from the memory store.
- In cloud mode, TairSearch branch metadata and TairVector routing hints are used to select candidate branches before the lexical fallback is considered.
- Qwen receives the user request plus only the relevant durable memory branches through the hosted answer adapter.
- The backend returns the Qwen answer with model name, context branch paths, context size, approximate token count, trace metadata, and degraded-mode status.
- The interaction is appended to the event stream for later consolidation.
For cloud runtime requests, the backend does not wait for memory mutation or
idle investigation. It returns consolidation: queued; the UI polls the worker
and refreshes memory, graph, context, and trace state after consolidation.
Deterministic local and comparison paths retain synchronous processing for
reproducible tests and evaluation.
The slow path is the background memory worker.
- The worker reads pending event records from the message bus.
- It maps each event to affected memory branches.
- It calculates normalized Shannon entropy over each branch's observed value history, then combines it with evidence volume for investigation priority.
- It asks the configured
MemoryModeladapter for a structured state delta or investigation plan. - If the model asks for external evidence and
QWENTEXT_WEB_SEARCH_ENABLED=true, the backend executes the search tool and calls the model again with the search results. - If
QWENTEXT_QWEN_CODE_ENABLED=true, the worker can optionally ask Qwen Code through the Python SDK to inspect project context in plan mode and return the same structured memory-investigation schema. - It applies model-returned branch, relationship, and importance changes to durable memory state through deterministic repository methods.
- It records a JSONL trace showing the reasoning, tool call, and resulting mutation.
The target deployment uses Alibaba Cloud services directly relevant to the hackathon.
| Capability | Service | Purpose |
|---|---|---|
| Model inference | Qwen Cloud / Alibaba Cloud Model Studio | Generates responses and state deltas. |
| Code investigation | Qwen Code SDK | Optionally inspects project context in plan mode and returns structured MemoryInvestigation data. |
| Compute | ECS | Runs the Python backend and background worker. |
| Memory engine | Tair Enterprise Edition | Stores durable memory state, streams, search metadata, vector routing hints, and temporal metrics. |
| Container image | Container Registry Enterprise Edition, or an external private registry | Provides the image pulled by ECS at startup. |
| Deployment identity | RAM and STS AssumeRole | Lets Terraform deploy without committing long-lived credentials. |
| Networking | VPC, VSwitch, security group | Isolates backend and Tair. |
| IaC | Terraform | Creates and destroys the cloud stack. |
The Qwen model itself is not provisioned by Terraform. It is consumed as a managed API. Terraform handles the application infrastructure and passes the Qwen API key into compute through sensitive variables or environment configuration. Do not hard-code API keys, passwords, tokens, or secrets in Terraform files.
The Tair-backed adapter uses the following data model:
| Tair capability | Use |
|---|---|
| TairDoc / JSON document operations | Current consolidated memory state through JSON.SET and JSON.GET. |
| Streams | Interaction events and background queue messages through XADD and XRANGE. |
| TairVector | Real Qwen branch embeddings written with TVS.CREATEINDEX and TVS.HSET for semantic context routing before lexical fallback. |
| TairSearch | Searchable branch metadata hashes indexed with FT.CREATE. |
| TairTS | Temporal memory metrics written with TS.ADD, including event counts and memory branch counts. |
| exHash / hash fields | Branch metadata such as path, confidence, importance, evidence count, and searchable text. |
TairMemoryStore and TairMessageBus also keep Redis-compatible string/list mirrors so a developer can run tests against simpler Redis-compatible endpoints without changing the core MemoryAgentService contract. The Alibaba Cloud judging runtime uses TairMemoryStore, TairMessageBus, QwenHostedModel, QwenAnswerModel, QwenEmbeddingModel, RedisTairClient, and DashScopeQwenClient; they expose the same dataclass-facing interfaces and serialize to JSON-compatible dictionaries only at adapter boundaries. Local deterministic setup is documented in LOCAL_DEVELOPMENT.md.
Cloud retrieval embeds branch text with Alibaba Cloud Model Studio text-embedding-v4 through the OpenAI-compatible embeddings API, creates a TVS.CREATEINDEX <index> 128 FLAT COSINE index by default, stores branch vectors with TVS.HSET, and routes answer context with TVS.KNNSEARCH plus TairSearch keyword candidates. Configure it with DASHSCOPE_EMBEDDING_ENDPOINT, DASHSCOPE_EMBEDDING_MODEL, and DASHSCOPE_EMBEDDING_DIMENSIONS; defaults are https://dashscope.aliyuncs.com/compatible-mode/v1/embeddings, text-embedding-v4, and 128. If vector commands or embedding calls fail, the trace reports degraded vector status and falls back to TairSearch or lexical ranking.
QWENTEXT_TAIR_PREFIX is the deployment-level Tair application prefix. The web UI namespace field is different: it sends the selected value as backend user_id for chat, memory state, graph, context, trace playback, imports, worker ticks, and idle investigation. In Tair this means one running deployment can isolate demo memories under keys such as qwentext:users:demo:... and qwentext:users:review-a:... without restarting or changing the Tair prefix.
Core Python code uses dataclass schemas such as WorldState, MemoryBranch, MemoryLink, MemoryEvent, QueueMessage, MemoryMutation, MemoryLinkMutation, MemoryInvestigation, ContextResponse, SessionEvent, TraceMessageEvent, and ModelChangeEvent; JSON dictionaries are kept at serialization boundaries such as Teich JSONL traces, client payloads, Hugging Face dataset rows, hosted model clients, and store adapters.
MemoryBranch serialization includes both value and current_value, plus previous_values, superseded_event_ids, resolution_reason, resolved_by_model, last_updated_event_id, and resolved_at. When a later mutation changes the same path, old current evidence moves into the supersession audit trail instead of being silently overwritten. Judges can inspect these fields in /api/world, graph nodes, and the UI node inspector.
MemoryAgentService composes the repository, message bus, model adapter, and trace sink:
ingest_trace_jsonl(user_id, text)parses Teich-style JSONL message records intoMemoryEventdataclasses.- Each memory event is appended to the configured
MemoryStore. - A
QueueMessageis published to the configuredMessageBus. process_all(consumer_id)simulates the background worker and asks the configured model forMemoryMutationrecords.- Event mutations are applied deterministically to
WorldStatebranches with event ids recorded as evidence. - Idle investigation asks the configured model for
MemoryInvestigationrecords that may include branch mutations, relationship links, importance deltas, or a web search query. - When Qwen Code is enabled, the configured model is wrapped by
QwenCodeMemoryModel; it usesqwen-code-sdkonly during idle investigation, defaults to permission modeplan, and must return structured memory deltas rather than editing the store directly. - When search is enabled, the backend executes the search tool and sends the returned results back to the model for follow-up integration before applying any external-reference links.
get_context(user_id, query)asks the store for branch candidates first; Tair-backed stores use TairSearch and TairVector routing, while local stores fall back to deterministic lexical ranking.chat(user_id, message)sends the message and compact context to the configuredAnswerModel; it also asks the same adapter for a no-context answer variant so the UI can compare memory impact.- Hosted Qwen calls retry transient 429 and 5xx failures with backoff. If answer generation still fails, the API returns a degraded local answer with the error recorded in the answer trace.
- Ingest, processing, investigation, context, and answer generation requests are logged through the trace sink as JSONL-compatible
AgentTraceEventrecords.
Agent traces are read and written as newline-delimited JSON. This keeps compatibility with Teich-generated Hugging Face datasets and with local text-chat uploads to the backend. Once a JSONL trace is loaded, the backend converts common event groups into dataclasses:
sessionbecomesSessionEvent;- nested
messageevents becomeTraceMessageEventwithTraceMessageandTextContent; model_changebecomesModelChangeEvent;- less common groups such as
turn_context,event_msg,response_item,session_meta,session_info, andthinking_level_changeare preserved asAgentTraceEventpayloads until a narrower schema is needed.
Serialize these dataclasses back to JSONL when returning traces to the client or preparing dataset input. Do not pass raw JSON dictionaries through memory processing when a dataclass schema exists.
The hosted demo is served by qwentext.web with FastAPI and package-local static
assets. It intentionally uses plain HTML, CSS, JavaScript, and canvas so the
memory system remains the focus. Run it locally with:
qwentext-web --host 127.0.0.1 --port 8000The first screen is a three-column app:
-
Chat:
- choose a memory namespace, defaulting to
demo; - type a memory-bearing prompt;
- see the assistant response;
- expand native thinking details to inspect the compact context sent with the request;
- inspect the answer trace showing model name, context branches, context size, approximate token count, provider trace, and degraded-mode status.
- choose a memory namespace, defaulting to
-
Memory context and updates:
- show the context branches used for the current answer;
- show stored durable memory branches;
- show startup cloud health for the memory store and answer adapter;
- show JSONL memory evolution trace records emitted by the backend.
-
Entropy graph:
- show branches as nodes and model-returned relationships as graph edges;
- load existing durable memory branches on first page load, including Tair-backed state from earlier judging or remote Tair runs;
- show loading and error status while startup memory is being fetched;
- select a branch node or stored-branch row to inspect storage metadata, search inclusion, evidence event ids, entropy, and update priority;
- show a bottom status strip for startup loading, background worker activity, selected nodes, and idle memory investigations;
- tick the background worker every 10 seconds, then ask the backend to rotate across stored branches only when no queued messages are available;
- report model-guided branch investigation, relationship linking, importance changes, and external search enrichment when a search adapter is configured;
- color high-entropy branches as unstable;
- animate a sprite-backed background agent walking toward a rest point below the current target branch;
- use the real
MemoryAgentServiceandBackgroundMemoryWorkerpath, backed by local adapters by default.
The chat panel also includes a Hugging Face trace playback flow, collapsed by
default behind a compact top bar so the main judging view stays focused on
manual chat, context, and graph evolution. Expand the bar, select a namespace
first if you want an isolated review run, enter a dataset repository id or URL,
load available .jsonl files through the Hugging Face Hub tree API, select a
session from the chosen file, then load it into a fixed playback tray above the
chat transcript.
Play next sends one trace message at a time through the same memory-event
queue used by chat, so judges can watch context, stored branches, traces, and
the graph evolve incrementally. Import all remains available for the faster
demo path that processes the full selected conversation at once.
The graph layout is topology-driven. It computes connected components from real
memory links, ignores synthetic world contains links, attracts nodes in the
same component toward a shared area, and pushes unconnected components apart.
There are no domain-specific graph clusters for programming, travel, identity,
or other branch names.
The frontend should remain a demonstration surface, not a large application. The backend memory behavior is the product.
Python code is a package under src/qwentext.
Required practices:
- write or update tests before changing implementation behavior;
- use
argparsefor command-line entry points; - prefer dataclasses for Python schemas instead of passing raw dictionaries through core APIs;
- expose focused
qwentext-*CLIs instead of one command with many subcommands; - keep external dependencies low;
- prefer Python standard library modules;
- use
unittestfor Python tests; - keep logic modular and shared through small reusable modules;
- use type hints on method and function declarations;
- use the Python
loggingmodule for command and service output; - never call
print()from code undersrc/; - keep comments minimal and only where they clarify non-obvious behavior;
- avoid broad exception handling and avoid try/except unless there is a concrete recovery path.
.
├── AGENTS.md
├── CLAUDE.md
├── LICENSE
├── README.md
├── pyproject.toml
├── src/
│ └── qwentext/
│ ├── __init__.py
│ ├── cli_entropy.py
│ ├── cli_compare_rag.py
│ ├── cli_local_e2e.py
│ ├── cli_trace_sample.py
│ ├── entropy.py
│ ├── memory.py
│ ├── messaging.py
│ ├── models.py
│ ├── qwen_code.py
│ ├── runtime.py
│ ├── schema.py
│ ├── service.py
│ ├── static/
│ │ ├── app.js
│ │ ├── index.html
│ │ └── styles.css
│ ├── web.py
│ ├── worker.py
│ └── traces.py
├── tests/
│ ├── test_adapters.py
│ ├── test_cli.py
│ ├── test_e2e_local.py
│ ├── test_entropy.py
│ ├── test_logging_policy.py
│ ├── test_memory.py
│ ├── test_qwen_code.py
│ ├── test_runtime.py
│ ├── test_traces.py
│ └── test_web.py
└── terraform/
├── main.tf
├── variables.tf
├── outputs.tf
├── terraform.tfvars.example
├── templates/
│ └── user_data.sh.tftpl
└── tests/
└── qwentext_stack.tftest.hcl
See LOCAL_DEVELOPMENT.md for editable installs, deterministic offline adapters, environment variables, local commands, and guidance on when to validate with Qwen plus Tair during development.
The package exposes focused argparse entry points.
qwentext-entropy '{"a": 3, "b": 1}'
qwentext-synthesize-evolution --pairs 1000 --seed 42 --dataset ds.json --traces-dir traces
qwentext-generate-evolution --pairs 10 --seed 42 --dataset generated.json --traces-dir generated-traces
qwentext-mine-evolution all_conversations.jsonl --dataset ds.json --traces-dir traces
qwentext-compare-rag --output-dir out --limit 3 # small budget-capped smoke run
qwentext-export-evolution --dataset ds.json --traces-dir traces --parquet evolution.parquet
qwentext-compare-rag --output-dir out --evolution-dataset ds.json --traces-dir traces
qwentext-compare-rag --output-dir out --evolution-dataset ds.json --traces-dir traces --evolution-only --limit 3
qwentext-container
qwentext-compare-rag --output-dir rag-comparison-output
qwentext-local-e2e path/to/sample.jsonl --query "travel"
qwentext-trace-sample
qwentext-web --host 127.0.0.1 --port 8000
qwentext-worker --interval-seconds 2
python -m qwentext.cli_entropy '{"a": 3, "b": 1}'
python -m qwentext.cli_local_e2e path/to/sample.jsonl --query "travel"
python -m qwentext.cli_trace_sampleqwentext-local-e2e wires only local adapters. It imports a JSONL trace, publishes memory-event messages, processes the queue with deterministic local rules, and logs the compact context returned for the query. Add --show-agent-trace to log the generated memory evolution trace.
qwentext-compare-rag runs 20+ scripted scenarios through the real Qwentext
cloud runtime (remote Qwen mutations and Tair storage; the command refuses to
run with local adapters) against deterministic in-process baselines named
chunk_rag_top1, chunk_rag_top3, recency_rag, summary_memory,
vector_like_rag, and latest_fact_profile, then writes
rag_comparison.csv, chart PNGs (rag_comparison_overall.png,
rag_comparison_efficiency.png, rag_comparison_context_size.png, and
rag_comparison_tradeoff.png), and
rag_comparison_report.md. Latency remains in the CSV but is not charted
because remote Qwen/Tair and in-process baselines are not comparable. The CSV is
seaborn-ready with columns system,scenario,metric,value; metrics include
answer correctness, contradiction handling, context size, retrieved/stored
item count, measured persistence after a real store reload, measured retrieval
latency, prompt size, stale-fact leakage, current-fact inclusion, and
estimated context-cost proxy, with aggregate ALL rows for charting and a
model column recording which Qwen model produced each Qwentext row. Pass
--evolution-dataset and --traces-dir to also run every memory-evolution
pair through all baselines; the CSV gains ALL_<tier> aggregate rows and the
Markdown report gains per-tier correctness/leakage tables so the synthetic,
manual, and mined tiers can be compared directly. Add --prompt-artifacts to
write per-scenario with-context and without-context prompt files.
Use --evolution-only with --evolution-dataset to exclude the built-in
scripted scenarios. --limit N is applied after this filtering, so a small
synthetic smoke test runs only the first N evolution pairs.
Long runs are resumable: each completed Qwentext scenario is appended to
<output-dir>/qwentext_cache.jsonl (override with --cache) together with
the model that produced it and a timestamp. Stop the process at any point and
re-run the same command — cached scenarios are skipped and only the remainder
costs tokens. The cache is keyed by scenario and model, so switching
DASHSCOPE_MODEL re-evaluates under the new model while keeping the earlier
model's results for comparison. Use --limit N first to smoke-test the
pipeline on a handful of scenarios before committing a large token budget.
QWENTEXT_MODEL_FALLBACKS accepts a comma-separated ordered model list (for
example qwen3.6-flash,qwen3.5-flash,qwen-flash). All Qwen adapters share one
ModelRotation: when the active model returns a quota-exhaustion error code
(DashScope Throttling.*, Arrearage, or OpenAI-compatible
insufficient_quota), every service in the process permanently switches to
the next model in the list and continues, and the answer trace plus compare
cache record the model actually used.
Model Studio API keys are region-scoped. The remote-Tair setup script uses
Singapore endpoints under dashscope-intl.aliyuncs.com. A Beijing key must
instead use the corresponding dashscope.aliyuncs.com endpoints. Mixing a
key and endpoint from different regions returns HTTP 401.
The memory-evolution dataset supports four provenance tiers; every record carries a source
field. qwentext-synthesize-evolution generates thousands of exact-label
synthetic pairs from a seeded slot grammar (the generator knows which facts
it superseded, so labels are ground truth by construction).
qwentext-generate-evolution asks the configured Qwen model to construct
realistic earlier/later conversations plus expected and faded phrase labels.
These records use source=generated; their labels are model-generated and
must not be described as exact construction labels or human ground truth.
Generated records include a neutral current-state query. Expected labels cover
retained, replacement, and newly added durable facts. Faded labels are never
added to that query. Answer-time compact context is capped at 2,000 characters.
Each completed pair is saved immediately so a partially completed run remains
inspectable. qwentext-mine-evolution mines mined silver-label pairs from consecutive
conversations in a local or Hugging Face JSONL trace file using a
deterministic keyword diff. manual records come from human curation on the
/dataset page.
Evolution pairs are replayed by qwentext-compare-rag --evolution-dataset —
trace A, consolidate, trace B, consolidate — through the Alibaba Cloud runtime
(remote Qwen mutations and Tair storage; the command refuses to run with local
adapters, so the reported numbers describe the same system judges use). The
per-tier tables report correctness over expected_words (expected recall) and
stale-fact leakage over faded_words (faded leakage). The script runs on your
machine with QWENTEXT_ENV=cloud, TAIR_HOST/TAIR_USERNAME/
TAIR_PASSWORD, and DASHSCOPE_API_KEY set.
qwentext-export-evolution converts the dataset plus trace files into one
Parquet file ready to push to a Hugging Face dataset repository, and
--from-parquet converts a published Parquet back into the local dataset and
trace files. See docs/design-notes.md for the schema
and tier methodology.
qwentext-web hosts both the HTML site and JSON API. Useful endpoints are
/api/chat, /api/events, /api/worker/tick, /api/context, /api/world,
/api/graph, /api/trace, /api/health, /api/huggingface/files,
/api/huggingface/conversations, /api/huggingface/playback, and
/api/huggingface/import. Pass user_id to
context, world, graph, snapshot, trace, worker, investigation, chat, playback,
and import calls to isolate a namespace.
qwentext-container is the Docker entry point for ECS. It reads
QWENTEXT_ROLE; with backend,worker it starts the FastAPI site and a
background memory worker in one container. Cloud runtime and optional Qwen Code
settings are detailed in LOCAL_DEVELOPMENT.md and the
Terraform deployment section below.
The demo should show one concrete memory evolution:
- A user or imported agent trace introduces a fact.
- A later event contradicts or refines the fact.
- The branch inspector shows
previous_valuesandsuperseded_event_idsafter the stale fact is replaced by compact current memory. - The background worker selects that branch.
- Qwen emits a structured state delta.
- The memory store mutates durable memory state.
- Repeated evidence for the corrected value lowers value-history entropy, and the next answer uses the corrected compact context.
This directly addresses efficient storage and retrieval, timely forgetting, and critical recall within a limited context window.
The demo is not claiming vectors are useless. It shows why vectors alone are a
weak long-term memory substrate: a chunk retriever can recall stale evidence
when a later event supersedes it, and it tends to grow prompt context as more
history arrives. Qwentext mutates compact durable memory instead: the latest
preference or fact becomes current_value, stale values move into
previous_values with superseded evidence ids, and context retrieval can send
only matching branches.
Generate the comparison artifacts (the script runs on your machine; Qwen and Tair are the remote Alibaba Cloud services):
export QWENTEXT_ENV=cloud # plus DASHSCOPE_API_KEY, TAIR_HOST, TAIR_USERNAME, TAIR_PASSWORD
qwentext-compare-rag --output-dir rag-comparison-outputOpen rag-comparison-output/rag_comparison.csv for raw metrics or
rag-comparison-output/rag_comparison_overall.png for the seaborn bar chart.
In the contradiction case Qwentext recalls the corrected preference from real
Qwen-consolidated durable memory in Tair, while the chunk baseline retrieves
the earlier stale chunk.
Terraform should create only infrastructure:
- VPC and subnet;
- security group;
- ECS instance for the backend and background worker container;
- optional Alibaba Cloud Container Registry Enterprise Edition instance, namespace, and repository;
- Tair instance for durable memory state and events;
- Tair application account for Redis-compatible access;
- startup environment shared from Terraform to the container;
- outputs for public endpoint and non-secret Tair connection metadata.
Terraform must not contain real secrets. Use environment variables such as TF_VAR_dashscope_api_key and provider-supported secret management. The repository should include placeholder variables and documentation, not credentials.
Required deployment inputs include:
export ALICLOUD_ACCESS_KEY="..."
export ALICLOUD_SECRET_KEY="..."
export TF_VAR_assume_role_arn="acs:ram::<account-id>:role/QwentextTerraformDeployer"
export TF_VAR_container_image="ghcr.io/<owner>/qwentext:latest"
export TF_VAR_acr_registry_domain="ghcr.io"
export TF_VAR_acr_registry_user="..."
export TF_VAR_acr_registry_pass="..."
export TF_VAR_dashscope_api_key="..."
export TF_VAR_tair_password="..."ALICLOUD_ACCESS_KEY and ALICLOUD_SECRET_KEY should belong to the caller RAM
user or role that is allowed to call sts:AssumeRole; Terraform then assumes
TF_VAR_assume_role_arn and receives the temporary STS security token
internally. The assumed role needs permissions for VPC, ECS, Tair/KVStore, and
ACR only when acr_enabled=true.
The Terraform defaults are sized for a small single-user demo in Singapore:
ap-southeast-1, ap-southeast-1a, ecs.t6-c1m1.large, a 20 GiB ECS system
disk, 1 Mbps public outbound bandwidth, and tair.rdb.1g. Override those
variables only if the selected zone does not offer that exact ECS or Tair class.
terraform/terraform.tfvars.example contains the non-secret shape for these values. Copy its keys into your own untracked .tfvars file or provide them through TF_VAR_ environment variables.
By default, Terraform does not create ACR. Use a private external registry to avoid the ACR Enterprise Edition subscription cost. Good low-cost options are:
- GitHub Container Registry (
ghcr.io): supports private images and repository permission inheritance through GitHub Packages. - GitLab Container Registry (
registry.gitlab.com): available on the GitLab Free tier for private projects. - Docker Hub: Docker Personal includes one private repository.
For the default private-registry path:
podman build -t qwentext:local-test .
podman tag qwentext:local-test ghcr.io/<owner>/qwentext:latest
podman login --username=<registry-username> ghcr.io
podman push ghcr.io/<owner>/qwentext:latest
export TF_VAR_container_image="ghcr.io/<owner>/qwentext:latest"
export TF_VAR_acr_registry_domain="ghcr.io"
terraform -chdir=terraform applyECS uses TF_VAR_acr_registry_user and TF_VAR_acr_registry_pass to log in
before pulling the private image. The image defaults to qwentext-container, so
container_command can stay empty unless startup needs to be overridden.
If you explicitly need Alibaba Cloud ACR Enterprise Edition, set
TF_VAR_acr_enabled=true. Enterprise Edition is subscription-based, so the
defaults use the smallest profile configured here: Basic, one month, manual
renewal, and provider-valid minimum quotas: 5 namespaces and 1000 repositories.
Because ECS must pull an image that already exists, use a two-step flow the
first time:
terraform -chdir=terraform apply -target=alicloud_cr_ee_instance.image -target=alicloud_cr_ee_namespace.image -target=alicloud_cr_ee_repo.image
terraform -chdir=terraform output acr_instance_endpoints
terraform -chdir=terraform output -raw acr_registry_domain
terraform -chdir=terraform output -raw acr_image_name
podman build -t qwentext:local-test .
podman tag qwentext:local-test <acr-domain>/<namespace>/<repo>:latest
podman login --username=<registry-username> <acr-domain>
podman push <acr-domain>/<namespace>/<repo>:latest
export TF_VAR_acr_enabled=true
export TF_VAR_acr_registry_domain="<acr-domain>"
terraform -chdir=terraform applyThe ECS user-data template writes /etc/qwentext.env with DASHSCOPE_API_KEY, QWENTEXT_ROLE=backend,worker, BACKEND_HOST=0.0.0.0, BACKEND_PORT, and TAIR_HOST, TAIR_PORT, TAIR_USERNAME, TAIR_PASSWORD, and TAIR_SSL. It then runs the configured container image as a qwentext systemd service using that env file.
For a lower-cost live demo, disable ECS and run the app on your laptop while Terraform provisions only remote Tair and supporting network resources:
export TF_VAR_ecs_enabled=false
export TF_VAR_local_tair_connection_enabled=true
export TF_VAR_local_client_ip="$(curl -s https://ifconfig.me/ip)"
terraform -chdir=terraform applyThen point the local app at remote Tair and Qwen:
export QWENTEXT_ENV=cloud
export BACKEND_HOST=127.0.0.1
export BACKEND_PORT=8000
export DASHSCOPE_API_KEY="$TF_VAR_dashscope_api_key"
export DASHSCOPE_EMBEDDING_MODEL="text-embedding-v4"
export DASHSCOPE_EMBEDDING_DIMENSIONS=128
# Optional for Singapore/Hong Kong workspaces: set the Model Studio compatible-mode embeddings endpoint.
# export DASHSCOPE_EMBEDDING_ENDPOINT="https://dashscope.aliyuncs.com/compatible-mode/v1/embeddings"
export TAIR_HOST="$(terraform -chdir=terraform output -raw local_tair_host)"
export TAIR_PORT="$(terraform -chdir=terraform output -raw tair_port)"
export TAIR_USERNAME="$(terraform -chdir=terraform output -raw tair_account_name)"
export TAIR_PASSWORD="$TF_VAR_tair_password"
export TAIR_SSL=false
export TAIR_TIMEOUT_SECONDS=5
# Optional: isolate this demo from earlier persisted Tair runs.
export QWENTEXT_TAIR_PREFIX="qwentext-local-$(date +%Y%m%d%H%M)"
qwentext-containerOpen http://127.0.0.1:8000. When finished, run
terraform -chdir=terraform destroy to remove Tair and networking resources.
If you reuse the default prefix and the default UI user demo, Tair will keep
older durable memory branches between app restarts; those branches remain visible
as stored memory, while the context sent to the answer path is filtered to the
branches relevant to the current query when possible.
Startup memory loading reads the Tair memory-state key once through /api/snapshot.
TAIR_TIMEOUT_SECONDS bounds Tair connect and read calls so a stale external
connection shows an error instead of leaving the UI waiting indefinitely.
Terraform tests live in terraform/tests/*.tftest.hcl and use Terraform's native test runner with a mocked Alicloud provider. Run them with:
terraform -chdir=terraform init -backend=false
terraform -chdir=terraform testThe repository .gitignore excludes Terraform state, plans, local .tfvars files, local environment files, Python caches, and common key material. Keep only placeholder examples such as terraform/terraform.tfvars.example under version control.
Useful Terraform outputs are:
backend_url;backend_public_ip;acr_instance_id;acr_instance_endpoints;acr_registry_domain;acr_image_name;container_image;tair_instance_id;tair_connection_domain;tair_port;tair_account_name.
backend_url is the public hosted HTML demo and API base URL. Open it in a
browser after terraform apply completes.
Qwentext is released under the MIT License. See LICENSE.

