A tool-using agent you can run as it is: one runtime behind a Feishu/Lark bot, a Slack bot, a browser, or a command line for your own machine. Each gives the agent a shell sandbox, MCP servers, skills, memories, credentials, scheduled tasks, subagents, a knowledge base and file publishing — and lets it pick up new abilities in conversation, without a redeploy.
It is also a library. If you are building your own agent on Spring Boot 4 and Spring AI, read docs/sdk.md; if you want to change this repository, read docs/contributing.md. At a glance below is the shape of it in one diagram, and docs/architecture.md draws the rest — how a run starts, what it is offered, where state lives, and which module may depend on which. docs/integrations.md indexes every module and its own README, docs/events.md covers the agent watching rather than waiting, and docs/advanced.md what a deployment can turn on that most do not need.
Every property and environment variable is documented in place, with the reason for its default, in
spring-agent-app-feishu/src/main/resources/application.yaml.
That file, not this README, is the configuration reference.
flowchart LR
subgraph TALK[Somebody talks to it]
feishu[a Feishu or Lark chat]
slack[a Slack channel]
browser[a browser]
cli[a terminal]
end
subgraph WATCH[Or nobody does]
hooks[GitHub GitLab Grafana deliveries]
mail[a watched mailbox]
clock[a scheduled task firing]
end
events[correlated into a situation and left to settle]
agent[SpringAgent the one entry point]
model[the model over any OpenAI compatible endpoint]
subgraph PERRUN[Assembled for that one run]
tools[built-in tools]
mcp[MCP servers this user may reach]
skills[that identity skills and memories]
kb[the knowledge base it may read]
end
subgraph OWNED[Owned by an identity never by the process]
home[a home with files credentials and a sandbox]
store[SQLite or MongoDB or Redis]
end
feishu --> agent
slack --> agent
browser --> agent
cli --> agent
hooks --> events
mail --> events
events --> agent
clock --> agent
agent --> model
model --> agent
agent --> tools
agent --> mcp
agent --> skills
agent --> kb
agent --> home
agent --> store
Everything funnels through one type. Whatever started a run — a person on a surface, a webhook nobody
was addressing, a clock — it becomes an AgentRequest handed to SpringAgent, and tool composition,
the model call, MCP lifecycle and cancellation all happen inside. Surfaces follow along as listeners,
including runs they did not start, which is how a scheduled task can report back into a chat.
One qualification the picture leaves out: exactly one chat surface belongs on an application's
classpath, which is why Feishu and Slack are separate applications rather than one server with
both. The browser is the exception, and is what
spring-agent-app-web-feishu pairs with a chat. The
reasoning, and four more diagrams — how a run starts, what it is offered, where state lives, which
module may depend on which — are in docs/architecture.md.
The agent is not a fixed feature list. Almost everything it can do arrives as a tool, and the registries that decide what tools exist are themselves tools — so the agent extends itself, in conversation:
| Ask it to… | …and it calls | …which gives the next run |
|---|---|---|
| use an MCP server you name | AddMcpServer |
every tool that server offers |
| learn a procedure | WriteSkillFile |
a skill, loaded on demand |
| remember something about you | the memory tools | notes it reads back before replying |
| hold a token for later | SetCredential |
the secret in its sandbox, never in a prompt |
| do this every Monday | CreateScheduledTask |
a run that fires on its own, within app.scheduling.sweep-interval of the moment, and still fired if the agent was down when it came round |
| every 10 minutes until it is fixed, or 10 times | CreateScheduledTask with maxRuns, StopThisScheduledTask |
a task that counts its own runs, and ends itself when what it watched for happens |
| take on something long | StartSubagent |
a second run doing the work, reporting back |
| write this down for the team | IndexKnowledge |
an answer that consults it from then on |
Nothing is registered up front. The tool set is assembled once per request out of the built-in tools, the MCP servers that user can reach, that user's skills and whatever the deployment configured — so a set that grew a minute ago is offered to the next turn. Because it is open-ended, both applications turn on tool search, which retrieves the few tools a turn actually needs instead of sending the model all of them.
Sometimes the open-ended set is the wrong thing. Ask what do we do about a failing canary and the agent may search the web, read a file and shell out — a better answer in general, and the wrong one when what you wanted was what this team wrote down, in a form you can go and check.
Start the message with /kb and that turn is answered out of the knowledge base and the agent's
memories alone, with no web, no shell, no files and no MCP servers:
/kb what do we do about a failing canary
/mini is the other one: it answers out of the model alone — no tools, no knowledge base, no
memories — for a quick aside you do not want in the thread. It leaves no trace, so the next question
cannot refer back to it. /one-off and /one_off mean the same.
If you find yourself typing one of them every time, ask the agent to make it your default — "from
now on answer me out of the knowledge base" — and it will. /full is then how you ask for an
ordinary run for one message, since a word in the message always beats your default.
/knowledge-base and /knowledge_base mean the same and case does not matter. The word may sit
anywhere in the message, not only at the front — so it still works in a group chat, where the bot
has to be mentioned first, and it reads naturally at the end of a thought:
what do we do when a deployment failed? /kb, tell me something
It works on every surface: Feishu, Slack, the browser and the command line. A slash word the agent does not recognise is left alone and reaches the model as ordinary text.
Every run carries a user id, and that id — not the process — owns the state. Under
app.storage.location each user has a home with memories/, skills/, workspace/ and
artifacts/, and the filesystem tools are confined to that root, so one user's agent cannot read
another's files. The same line runs through everything else:
- MCP servers are registered by their owner.
ShareMcpServergrants use to another person, a group chat, or everyone, while editing, removing and re-sharing stay with the owner — the recipient never sees the URL or headers. Servers the deployment configures for everybody underspring.ai.mcp.client.*are listed alongside them and belong to nobody. On the browser surface they are also a page — Customize → MCP servers — which is where a URL and a token are worth typing rather than dictating to a model that will echo them into a transcript. Adding one connects to it and lists what it offers before anything is stored, so a wrong URL is caught where it was typed; the stored token is never sent back to the page, and a server can be turned off without being forgotten. - Skills are folders with a
SKILL.mdin the user's own skills directory. The agent writes and deletes them on request; paths outside that directory are refused. On the browser surface they are also a page of their own — Customize → Skills lists what you and your company have, opens one into a file tree beside the file being read, and lets you write a skill by hand, drop a.zipof one in, download one back out again, or correct a line of an existing one without asking the model to do it for you. After a turn that has cost a great many tool calls the agent also offers one unprompted — it finishes the answer, then asks whether to keep the method it worked out, and writes the skill only if you say yes. Turn the offer off withSKILLS_OFFER_AFTER_EXPENSIVE_RUNS=false, or move the bar withSKILLS_TOOL_CALL_THRESHOLD. - Memories are files the agent writes and reads back itself, in each of the homes a request
reaches: your own, the group chat's, and the company's. It names the scope when it saves, so a
convention that binds a chat is remembered by the chat rather than by whoever happened to be
typing. A shared memory can only be written from a group chat, where the write happens in front
of the people it affects — in a one-to-one chat the agent reads the company's memory but saves to
your own, unless you are listed in
AI_ADMINS. On the browser surface they are also a page — Customize → Memories — which lists what the agent has learnt with the front matter of each memory read onto its card, and opens one to correct or forget it. That matters more than it does for a skill: a skill is something you wrote, while a memory is what the agent concluded about you, and the only other way to fix one is to ask the thing that wrote it to unwrite it. Writing the company's from there followsAI_NON_ADMIN_TENANT_WRITES, the same switch that governs company skills and knowledge. - Credentials are per-user: a Kubernetes Secret mounted into that user's sandbox, or an
encrypted row, so a token reaches a shell as an environment variable and never a prompt. On
Kubernetes an operator can also share Secrets they provisioned themselves — with a group, a
tenant, or one named person — by labelling them to match a selector under
app.ai.tools.shell.kubernetes.credentials.shared; a credential the user set for themselves still wins the name. - The sandbox shell is a Pod or container per user, with its own slice of the volume, torn down when idle and rebuilt on the next command.
- The knowledge base is scoped to a person, a group chat or the whole tenant, and a run only ever searches what its own identity may read.
- The chat model itself, where the deployment allows it: a user can register their own OpenAI-compatible endpoints and have their conversations answered by one of them, leaving everybody else on the application's. See Bring your own model.
A message from a group chat also reaches the group's home and the group's knowledge, which is how a
team shares skills and notes without sharing anything private. Nobody needs an administrator to set
any of this up — they ask the agent, and it registers it for them. On the command line the same
machinery serves the one person at the keyboard, out of ~/.spring-agent.
Every run goes through the application's model unless the person asking has chosen another. Off
unless USER_MODELS_ENCRYPTION_KEY is set — the tokens people register are bearer credentials for
somebody else's paid endpoint, and the only alternative to storing them sealed is storing them in
the clear, so a deployment that cannot do the first does not offer the feature at all:
export USER_MODELS_ENCRYPTION_KEY=$(openssl rand -base64 32)Keep it out of the database and out of version control. Rotating it does not re-seal what is already stored: those rows stop being readable and say so, rather than quietly behaving as though nobody had registered anything.
With it set, a user can ask the agent — add my Kimi endpoint, what models do I have, switch me back to the default — through the AddChatModel, ListChatModels, UseChatModel and
DeleteChatModel tools. An endpoint is connection-tested before it is stored: if it cannot be
reached, the token is refused or the model name is unknown, nothing is saved and the reason comes
back. Tokens are never shown again, to anyone, including the person who set them.
There is also a way in that does not involve the agent, and it is the important one. A model that
stops answering would otherwise break the very run needed to undo it, so /config never touches
the LLM:
| Surface | How |
|---|---|
| Feishu | Send /config. A card opens with a dropdown of what you could be on and fields for adding an endpoint. |
| Slack | Type /config. A modal opens, private to you, so the API token never enters channel history. The command has to be declared on the Slack app — see below. If nobody did, send /config with a leading space instead: Slack sends that verbatim rather than looking for a command, and the same form arrives as a message. |
| Command line | /config lists your models, /config <name> switches, /config default returns to the built-in one, and /config <name> <effort> or /config default <effort> sets how hard it thinks. |
| Browser | The Model tab under Customize, drawn only where the encryption key is set. It lists what you have registered, switches between them and back to the built-in one, and opens each as the form that registered it. |
An endpoint that needs more than a bearer token — a gateway wanting a tenant or routing header —
takes extra headers, typed as Name: value per line on the Feishu card and the browser form, or
handed to AddChatModel as a map. They are sealed the same way the token is, and like the token
they are never shown again: a listing says which headers are set, never what is in them.
The dropdown also lists what the application's own endpoint reports it can serve, so choosing among
the models the deployment already pays for needs no token of your own. That listing is best-effort:
an endpoint that does not answer GET /models simply shows the one built-in entry. What it does
answer with is usually more than chat models — the embedding model this deployment uses, a
reranker, a speech or image model — and those are filtered out by name, since GET /models says
nothing about what a model is for. Where the endpoint serves more chat models than a card can hold
the list is cut short. Either way, fill in the Model field alone, leaving name, base URL and
token empty, to name any model the endpoint serves directly.
The same form carries the reasoning effort, chosen from a list rather than typed — none, minimal,
low, medium, high, xhigh, max, as the OpenAI API spells them — plus two entries that are
not values:
- whatever this deployment is set to, which is
OPENAI_REASONING_EFFORTand what every model answers with until somebody chooses otherwise; - not sent at all, for a gateway that rejects
reasoning_effortoutright rather than ignoring it.noneis a real value the newer models act on, so it is not a way of leaving the parameter out.
It applies to whichever model the form leaves you on, the built-in ones included, and is remembered per model: switching away and back keeps it. The effort shown on the form is the one in force, and the thinking panel on a Feishu reply reports the effort that run was actually made with rather than the deployment's.
An effort is part of what gets connection-tested when an endpoint is registered, so a gateway that refuses one says so before anything is stored.
The embedding model is deliberately not configurable this way. The knowledge base is shared between users and its collections are built with one embedding model, so letting one person change theirs would invalidate vectors that are not theirs.
Other knobs, all optional: USER_MODELS_MAX_PER_USER (default 10), USER_MODELS_CACHE_SIZE
(default 50 live endpoints) and USER_MODELS_PROBE_TIMEOUT (default 30s).
Given app.events.enabled, the agent watches instead of waiting: alerts and code-hosting webhooks
arrive at /events/webhooks/<source>, mail arrives in a watched mailbox, group chat messages it was
not addressed in arrive through the chat integration. Related events are correlated into one
situation and left to settle — a thousand alerts from one outage become one run, not a thousand —
and only then is the agent woken to decide whether it has anything worth saying. Silence is a normal
answer.
flowchart LR
a1[a thousand alerts from one outage]
a2[an issue opened and then commented on]
a3[a mail thread]
intake[EventIntakes each intake isolated]
sit[one situation correlated by key and debounced]
triage[a triage run under its own identity]
out[an opinion on the chat or nothing at all]
a1 --> intake
a2 --> intake
a3 --> intake
intake --> sit
sit --> triage
triage --> out
Payload text is written by whoever caused the event, so a triage run treats it as evidence, never routing and never instructions, and runs as the agent rather than as any person.
GitHub, GitLab, Grafana and a mailbox ship as sources. Each authenticates its own deliveries, and a source nobody configured a secret for refuses everything, so the endpoint is safe to expose but useless until somebody sets it up. What the agent should actually do about a source's events is not a setting but a playbook: documents you write into the knowledge base and edit like any other, without a deployment.
It is all off by default, and there is more to decide here than anywhere else in the agent — who a source runs as, who it will listen to, and what any of that is worth against text a stranger wrote. docs/events.md is the whole of it.
Five applications, one runtime. Pick the surface you want; each has its own page with the setup steps, its variables and what it carries.
spring-agent-app-feishu |
A Feishu/Lark bot | The published image; carries every optional module |
spring-agent-app-slack |
A Slack bot | The same server, Slack instead |
spring-agent-app-webui |
A browser | The runtime with everything a run does made visible, and no chat platform |
spring-agent-app-web-feishu |
Both at once | So a conversation can be handed between a Feishu chat and the browser |
spring-agent-app-cli |
Your own terminal | SQLite under ~/.spring-agent, and a shell on your own machine |
The quickest of them:
docker run --env-file .env -p 8080:8080 ghcr.io/kezhenxu94/spring-agent:latestIt needs to be told where the models are and which models to ask for, and there are four ways to say
it. Either the six OpenAI-compatible variables — OPENAI_BASE_URL, OPENAI_API_KEY, OPENAI_MODEL,
EMBEDDING_BASE_URL, EMBEDDING_API_KEY, EMBEDDING_MODEL — which is what any gateway or
self-hosted server takes; or, on Alibaba Cloud DashScope, DASHSCOPE_API_KEY plus the two model
names, DASHSCOPE_CHAT_MODEL and DASHSCOPE_EMBEDDING_MODEL, since one credential covers every
DashScope endpoint but no endpoint has a default model. Add DASHSCOPE_BASE_URL — a host, with no
path — for the international endpoint or a Model Studio workspace.
Or, on Google Gemini, GEMINI_API_KEY plus GEMINI_CHAT_MODEL and GEMINI_EMBEDDING_MODEL, with
CHAT_MODEL_PROVIDER and EMBEDDING_MODEL_PROVIDER set to google-genai. Gemini is reachable
through its OpenAI-compatible endpoint with the first set too; naming the provider is what gets its
thinking levels, its own embeddings, and image models that edit from a reference image. The provider
switches are per kind, so the sets mix — a Gemini chat model over DashScope embeddings is one line
from each.
Or, on Anthropic, CHAT_MODEL_PROVIDER=anthropic with ANTHROPIC_API_KEY and
ANTHROPIC_CHAT_MODEL. Claude can also be served by a Google Cloud project that holds the
entitlement: set ANTHROPIC_BACKEND=vertex with ANTHROPIC_VERTEX_PROJECT and
ANTHROPIC_VERTEX_LOCATION instead of a key, and the credential comes from gcloud or from the
workload identity the deployment already runs under. Anthropic serves no embeddings, so this set is
always paired with one of the others for EMBEDDING_MODEL_PROVIDER — which the per-kind switches
make ordinary rather than a workaround.
Nothing starts without one of those sets: the application says which variable is missing rather than failing on the first run. The embedding model is needed even if you index nothing, since tool search is built by embedding tool descriptions. Each application's page lists what else it needs.
Three switches decide what a deployment actually is, and they mean the same thing in every application:
| Property (env var) | Values | Default |
|---|---|---|
app.persistence.type (PERSISTENCE_TYPE) |
jpa (SQLite, no server needed), mongodb, redis |
jpa |
app.ai.tools.shell.type (TOOLS_SHELL_TYPE) |
none, kubernetes, docker, local |
none |
spring.ai.model.image (IMAGE_MODEL_PROVIDER) |
none, openai, dashscope, google-genai |
none |
The shell defaults to none because it runs commands the model wrote. Turn it on deliberately, and
prefer kubernetes or docker, which give each user a disposable sandbox, over local, which does
not.
Image generation defaults to none for a related reason: it is a paid third-party API the agent
would start calling on the model's say-so, and the two providers' image APIs are genuinely different
rather than one endpoint with two hostnames. With none there is no GenerateImage tool at all,
which is better than a tool that always fails. spring.ai.model.image is Spring AI's own property,
not one of this project's — the third switch borrows a namespace rather than adding one. docker-compose.yaml has a compose profile per value of both switches, so the containers and the
application's own choice cannot drift apart:
PERSISTENCE_TYPE=redis VECTORSTORE_TYPE=milvus \
COMPOSE_PROFILES=$PERSISTENCE_TYPE,$VECTORSTORE_TYPE docker compose up # backends onlyAdd app to COMPOSE_PROFILES to run everything in containers. The default pair needs no server at
all.
Every module in the repository has a README of its own saying what it is, what it needs and what to know before changing it — docs/integrations.md is the index, and the place where what they all have in common is written down.
make # ./gradlew build
make test # needs a running Docker daemon: the tests start MongoDB and Redis via Testcontainers
make lint # spotlessApplyJava bytecode targets 21, built with a GraalVM 25 toolchain because native-image ships with it.
CI builds and tests every push to main and every pull request; run make before pushing anyway. See
docs/contributing.md for the rest.