Describe it in plain English. Get a production-ready Splunk app in minutes.
Building a Splunk app is hard for a reason most people don't appreciate until they try. It's not the Python. It's not even the SPL. It's the gap between knowing what you want the app to do and knowing how Splunk wants you to express it — conf files, restmap routes, AppInspect rules, packaging conventions, agent wiring, schema grounding against your specific indexes. A senior Splunk engineer moves through that gap in a couple of days. Everyone else stalls. The people who feel the operational pain — the ops engineer watching 4xx errors spike, the security analyst wanting an enrichment workflow — are usually not the people who can build the app that fixes it.
SplunkForge closes that gap. Not with a scaffolder that produces a starting point you still have to finish, but with something that produces a finished app. Describe what you need in one sentence. Get back a deployed, verified, working Splunk app — schema-grounded, logic-engineered, AppInspect-clean. In about four minutes.
Built for the Splunk Agentic Ops Hackathon 2026 — Platform & Developer Experience track.
User Prompt (natural language)
│
▼
Intent Classifier ────────── LLM analyzes intent and produces a structured App Spec
│
▼
Expert Expansion ───────────── Expands vague intent into a complete spec
│ Forces minimum artifact counts (8+ panels, 3+ subagents, 5+ tools)
│ Covers 6 operational dimensions: current state, trend, breakdown,
│ comparison, anomaly, action
│
▼
MCP Schema Grounder ────────── Live connection to your Splunk instance
│ splunk_get_indexes Discovers real indexes, sourcetypes, field names
│ splunk_get_metadata No placeholders — grounded SPL from day one
│ saia_generate_spl (if available)
│
▼
┌─ Spec Review Gate ──────────── Human review: adjust index, sourcetype, capabilities
│
▼
Logic Planner (SSE stream) ─── Agentic multi-step planner designs:
│ Archetype A1: Spike detection + AI triage agent
│ Archetype A2: Data enrichment custom command
│ Archetype A3: Scheduled intelligence report
│
▼
┌─ Logic Review Gate ─────────── Human review: SPL, agent prompts, tool definitions
│
▼
App Assembler ──────────────── Parallel generators produce all required files:
│ props.conf · transforms.conf · commands.conf · savedsearches.conf
│ Dashboard Studio JSON · REST handlers · alert actions
│ splunklib.ai agents · tool registry · hosted model config
│
▼
AppInspect Fix Loop (≤3x) ─── Runs AppInspect CLI, auto-fixes failures with LLM
│
▼
✅ .spl Package ──────────────── Ready to install
| Feature | How SplunkForge Uses It |
|---|---|
| Splunk MCP Server | Schema grounding — splunk_get_indexes and splunk_get_metadata discover real indexes, sourcetypes, and field names from your live Splunk instance |
AI for Splunk Apps (splunklib.ai) |
Generated apps include full agentic backends — supervisor/specialist topology, tool registries, structured outputs |
| Hosted Models | Generated apps are wired to use Splunk-hosted models by default; configurable to Anthropic/OpenAI/Groq/Cerebras |
| AppInspect | Automated validation + LLM-powered fix loop — generated packages pass AppInspect before you see them |
saia_generate_spl (if available) |
When a Splunk AI token is configured, saved search SPL is generated by the AI Assistant using live schema context for higher accuracy |
SplunkForge is a full-stack application built in about a week, with a Python/FastAPI backend and a Next.js frontend communicating over Server-Sent Events for real-time streaming.
The backend is a pipeline of independent, composable stages with Pydantic v2 schemas at every boundary. Each stage produces a typed intermediate (AppSpec → ExpansionResult → LogicPlan → GeneratedApp) consumed by the next — making individual stages debuggable and iterable in isolation.
-
Intent Classifier (
core/classifier.py) — Takes a natural language prompt and calls an LLM to produce a structuredAppSpec: app id, label, capabilities, target index, sourcetype, and whether AI is requested. -
Expert Expansion — The differentiator. Runs before the Logic Plan stage and forces six operational dimensions and hard minimum artifact counts, preventing the model from taking the easy path. This is what gets from one-panel scaffolding to genuinely useful apps.
-
MCP Schema Grounder (
mcp/grounding.py) — Connects to the Splunk MCP Server and callssplunk_get_indexesandsplunk_get_metadatato enrich the spec with real field names and schema-accurate SPL. If a Splunk AI token is configured,saia_generate_splalso runs. Grounding degrades gracefully — the UI surfacesok / unconfigured / failedstatus explicitly so users always know whether their SPL is schema-grounded or templated. -
Logic Planner (
core/logic.py) — A four-stage agentic pipeline that streams each logic item to the frontend over SSE as it's generated. Detects which archetype applies (A1/A2/A3) and designs the agent topology, tool definitions, SPL queries, system prompts, and structured output schemas. -
Generators (
generators/) — One Python module per file type. TheAppAssemblercalls each applicable generator in parallel and collects output into a directory tree. Generators use Jinja2 templates for conf files and Python boilerplate, with logic from theLogicPlaninjected at render time. -
AppInspect Fix Loop (
validators/) — Runs the AppInspect CLI as a subprocess, parses the JSON report, and passes failing checks to an LLM fix agent. Repeats up to 3 times. ThePackagerthen tarballs the directory into a.splfile. -
Settings API (
api/routes/settings.py) — Lets the UI configure Splunk credentials and LLM keys at runtime, persisting to.envand updating in-memory config without a restart. Pydantic v2BaseSettingsmodels are frozen by default, so the fix wasobject.__setattr__(settings, attr, val)to bypass the frozen check — a deliberate workaround for a specific constraint.
The frontend is a three-panel workspace: prompt input on the left, pipeline visualization and Logic Plan review in the center, generated file tree on the right. The Logic Plan view is the centerpiece — it makes the AI's reasoning visible and editable rather than hiding it in a black box.
-
Chat + SSE orchestration (
lib/simulator.ts) — TheSimulatorclass drives the full pipeline: POSTs to/api/generate/draft, reads theAppSpec, then opens SSE streams for logic planning and app assembly. Each event updates the Zustand store, which re-renders the relevant tab. -
Phase-gated UI — Two human review gates. The Spec Review Gate lets users edit the
AppSpecbefore logic planning begins. The Logic Review Gate shows each generated logic item (SPL queries, agent prompts, tool definitions) with inline editing before assembly starts. -
Workspace tabs —
SpecReviewTab,LogicPlanPanel,LogicReviewTab,FileExplorer,PipelineGraph, andAppInspectTabeach subscribe to slices of the Zustand store and update in real time.
splunklib.ai's Agent is fully async — you await it inside an async with block. Splunk's SCPv2 custom search command protocol calls generate() synchronously. These don't naturally connect, and nothing in the docs covers using both at once. The bridge is one line:
def generate(self): # sync — SCPv2 requires this
yield from asyncio.run(self._run())
async def _run(self): # async — splunklib.ai requires this
async with Agent(...) as agent:
response = await agent.invoke_with_data(...)
yield recordasyncio.run() spins up a fresh event loop, runs the coroutine to completion, and returns control to the sync caller. yield from drains the results back into the sync generator. We found this by reading the splunklib.ai source — it's not documented anywhere.
Vague prompts producing thin apps
The first version generated technically correct but operationally useless output — one-panel dashboards, single-tool agents. The expert expansion stage, with hard artifact minimums across six operational dimensions, was the fix. Forcing the model to ask "what would a senior engineer build here?" before writing anything is qualitatively different from interpreting the prompt literally.
splunklib.ai is alpha
The library is on the develop branch and the SCPv2 streaming command protocol needed manual async-to-sync bridging that isn't documented. We worked it out by reading the source and packaging the wrapper into our generators.
MCP query timeouts
Schema discovery against large indexes hit MCP's timeout. Fixed by chunking into smaller targeted queries and caching grounded schema per session.
AppInspect compliance
Generated apps initially failed checks on subtle rules — [triggers] reload names must not include .conf, [ui] label must be 5–80 characters without "Splunk For", local/ directories must never be shipped. We hardened the generators to produce passing structures by construction rather than by post-hoc fixing.
Proving the AI part actually works
A code generator that produces unrun code is not impressive. The hardest engineering decision was committing to the verification loop — automating deploy, search execution, agent invocation, and AppInspect into a single smoke test that surfaces real evidence in the UI. A generated app that hasn't been run is a prototype. One that has been deployed, executed, and validated is a product.
SSE streaming edge cases
Errors in SSE events were initially swallowed silently by a catch block that caught everything. Fixing it required distinguishing SyntaxError (benign, from partial chunks) from application errors, and guaranteeing isProcessing resets even if the stream ends unexpectedly.
AI code generation gets dramatically better when you make the model think like a domain expert before it writes anything. Literal prompt interpretation produces thin output. The expert expansion step — asking "what would a senior engineer build here?" — produces output that's qualitatively different: not just more code, but better-reasoned code.
The verification loop changes how the project is perceived. A generated app that hasn't been run is a prototype; one that has been deployed, executed, and validated against real data is a product.
The intermediate representation matters more than the final output. The Logic Plan being plain English and reviewable is what makes SplunkForge a tool engineers actually want to use, not just a black box.
- Expand from three app archetypes to the full set: security enrichment pipelines, observability dashboards, custom SPL commands, ML-enabled apps.
- Conversational refinement — edit the Logic Plan through dialogue rather than direct manipulation.
- Multi-app composition — apps that extend or depend on other apps.
- Splunk Cloud deployment alongside Splunk Enterprise.
- A shareable library of Logic Plans that users can fork and remix — not as code, but as intent.
- Python 3.11+
- Node.js 18+ (for the web UI)
- A running Splunk instance with token auth enabled
- An LLM API key (Anthropic, OpenAI, Groq, or Cerebras)
- AppInspect CLI:
pip install splunk-appinspect
git clone https://github.com/JyothsnaAshok/SplunkForge.git
cd SplunkForge
# Backend
cd splunkforge
pip install -e ".[dev]"
cp .env.example .env
# Edit .env with your Splunk credentials and LLM API key
# Frontend
cd ../web
npm install# Terminal 1 — backend (from splunkforge/)
uvicorn splunkforge.api.server:app --reload --port 8000
# Terminal 2 — frontend (from web/)
npm run devOpen http://localhost:3000.
Click the ⚙ Settings button in the top-right to configure Splunk credentials and your LLM API key — no manual .env editing required.
| Variable | Description | Required |
|---|---|---|
ANTHROPIC_API_KEY |
Anthropic API key | One of these |
OPENAI_API_KEY |
OpenAI API key | One of these |
GROQ_API_KEY |
Groq API key | One of these |
CEREBRAS_API_KEY |
Cerebras API key | One of these |
SPLUNK_AI_TOKEN |
Splunk AI (SAIA) token | One of these |
SPLUNK_HOST |
Splunk hostname/IP | Yes |
SPLUNK_PORT |
Splunk management port (default: 8089) | No |
SPLUNK_TOKEN |
Splunk auth token for MCP grounding | Yes |
SPLUNK_MCP_URL |
MCP Server URL | No |
See splunkforge/.env.example for the full list.
web/ Next.js 15 frontend
app/ App Router pages
components/
chat/ Conversational input bar
layout/ TopBar, settings modal
workspace/ Spec, logic, pipeline, AppInspect, file explorer tabs
lib/
store.ts Zustand state
simulator.ts SSE stream orchestration
types.ts Shared TypeScript types
splunkforge/ Python backend
splunkforge/
api/
routes/ FastAPI route handlers (generate, spec, settings, package)
sse.py SSE event helpers
core/
classifier.py Intent → App Spec (LLM)
spec.py Pydantic App Spec models
logic.py Logic planner (SSE streaming)
assembler.py Generator orchestration
mcp/
grounding.py Splunk MCP client + schema collection
generators/ One generator per file type
validators/
appinspect.py AppInspect subprocess + report parser
fix_agent.py LLM-powered fix agent
templates/ Jinja2 templates for all generated files
tests/ 88 passing tests
- Intent classification → structured App Spec
- Expert expansion with minimum artifact counts
- Live MCP schema grounding (indexes, sourcetypes, fields, AI-generated SPL)
- Spec review gate with inline editing
- Agentic logic planner with SSE streaming (A1/A2/A3 archetypes)
- Logic review gate with per-component editing
- Full generator suite (conf files, dashboards, commands, REST, alerts, AI agents)
- async-to-sync bridge for SCPv2 + splunklib.ai
- AppInspect auto-fix loop (≤3 iterations)
-
.splpackage export - Settings UI (no manual
.envediting) - MCP connection status + grounding failure surfacing
- 88 passing tests
MIT — see LICENSE.
SplunkForge · Built for the Splunk Agentic Ops Hackathon 2026 · Platform & Developer Experience Track
