Generate educational, illustrative images from the command line using OpenAI
(gpt-image-2, …) or Google Gemini image models. One command turns a prompt into a
saved PNG, with style presets and reusable templates tuned for learning material.
It works anywhere you want images on disk. If you keep an Obsidian vault
(or any Markdown notes), gen-image can also print and copy a ready-to-paste ![[wikilink]] for
the file, but that integration is entirely optional.
There are plenty of ways to call an image model from a terminal. They cluster into three groups, and gen-image deliberately sits in the gap between them:
- Thin OpenAI image CLIs (dallecli,
openai-cli-art, the official
openaiCLI) are raw passthroughs: prompt in, PNG out. Single provider, no opinion about what you're making, no cost tracking, no reuse. - Obsidian plugins (AI Assistant, obsidian-ai-images) are convenient but GUI-locked inside the editor, not scriptable, and have no budget discipline.
- Mega multi-provider LLM CLIs (aichat,
llm) do everything, so image generation is an afterthought bolted onto a chat pipeline with no domain tuning.
gen-image's job is narrower and sharper: opinionated, cost-aware, scriptable image generation tuned for material that teaches. It trades breadth for taste.
| Thin OpenAI CLIs | Obsidian plugins | aichat / llm |
gen-image | |
|---|---|---|---|---|
| Purpose-tuned style presets | ✗ raw | ✗ raw | ✗ raw | ✓ educational / mnemonic / diagram / first-person / manga / blueprint |
| Reusable templates (variable substitution) | ✗ | ✗ | ✗ | ✓ |
| Cost discipline (estimate, budget cap, history, stats, dry-run) | mostly ✗ | ✗ | partial | ✓ first-class |
| Scriptable CLI | ✓ | ✗ GUI-locked | ✓ | ✓ |
| Multi-provider, image-focused | ✗ OpenAI only | mixed | ✓ kitchen-sink | ✓ Gemini + OpenAI, curated |
| PKM output without lock-in | ✗ | ✗ lock-in | ✗ | ✓ optional wikilink |
The three things nothing else combines:
- Style presets are baked-in prompt engineering, not passthrough.
--style educational-cartoonencodes a tested prompt for memorable, text-free explanatory art. SeeGALLERY/for the same concept rendered across every preset. - Cost is a feature. Per-image estimates, monthly budget caps, a history log,
--stats, and--dry-runmean you can run it daily without dreading the bill. - CLI, not plugin. Runs in cron, scripts, and any editor; the Obsidian wikilink is convenience, never a requirement.
The honest tradeoff: the moat is taste, not technology. A determined user could paste a good prompt into any thin CLI and approximate the presets. The value is in the curation, the cost habits, and the learning-focused defaults being there by default. If you want raw breadth or in-editor clicking, the tools above serve you better.
- Two providers: Google Gemini (default) and OpenAI, selectable via config.
- Style presets built for explanation:
educational-cartoon,mnemonic,diagram-alternative,first-person,manga-strip,vintage-blueprint,custom. - Templates with variable substitution for repeatable prompts (e.g. vocab mnemonics).
- Cost-aware: per-image cost estimates, monthly budget limits, generation history.
- Interactive mode, a TOML config, and
--dry-runfor zero-cost previews.
- Python 3.11+
uv- An API key for at least one provider: Google AI Studio or OpenAI
git clone https://github.com/grepinsight/gen-image.git
cd gen-image
uv sync
uv tool install . # exposes the `gen-image` command on your PATH
gen-image --helpSet a key for your provider. Gemini is the default:
export GOOGLE_API_KEY="..." # or GEMINI_API_KEY
# to use OpenAI instead:
export OPENAI_API_KEY="sk-..."
gen-image --edit-config # set provider = "openai"A shell export only lives for that session, so cron jobs and scripts won't see it. Put the key
in a durable, gitignored file at ~/.config/gen-image/.env (chmod 600) and gen-image loads it
automatically on every run:
mkdir -p ~/.config/gen-image
printf 'GEMINI_API_KEY=%s\n' "your-key-here" > ~/.config/gen-image/.env
chmod 600 ~/.config/gen-image/.envThe loader accepts KEY=value, an optional export prefix, and quoted values, and skips
blank/comment lines. An exported environment variable always wins over the file. Recognized
keys: GEMINI_API_KEY, GOOGLE_API_KEY, OPENAI_API_KEY.
gen-image "A cartoon librarian locking a book in a safe" --output ownership.pnggen-image "Technical concept..." --style educational-cartoon --output img.png # default
gen-image "Visual wordplay..." --style mnemonic --output vocab.png
gen-image "System architecture..." --style diagram-alternative --output diagram.png
gen-image "Late-night trading desk, a screen reading -50%..." --style first-person --output scene.png
gen-image --list-stylesgen-image --list-templates
gen-image --show-template mnemonic-vocab
gen-image --template mnemonic-vocab \
--var word="語りかける" \
--var romaji="katarikakeru" \
--var meaning="to address, to speak to" \
--var concept="words as birds flying to the ear" \
--output vocab.png
gen-image --edit-template my-custom-template # create/edit in $EDITORBuilt-in templates include educational-metaphor, mnemonic-vocab,
visual-comparison, and three tuned for transcript/podcast-derived visuals:
spectrum-triad— a three-point spectrum (e.g. fragile/robust/antifragile), one figure per column.metaphor-mapping— render an abstract concept as its central metaphor, with annotated metaphor→meaning callouts.thought-experiment— two diverging outcomes under a shared setup banner (e.g. ensemble vs time average).
Run gen-image --show-template <name> to see each one's required --var slots.
gen-image "..." --output img.png --quality hd # higher quality, costs more
gen-image "..." --output img.png --size 1792x1024 # landscape (or 1024x1792 portrait)
gen-image "..." --output img.png --show-prompt # preview the full prompt
gen-image "..." --output img.png --dry-run # no API call, no costgen-image --interactive # guided prompts
gen-image --show-config # view config
gen-image --edit-config # edit config in $EDITOR
gen-image --list-models # list available models per provider
gen-image --history # generation history
gen-image --history-search rust # search history
gen-image --stats # cost statistics- Models:
gemini-3.1-flash-image-preview(default),gemini-3-pro-image-preview,gemini-2.5-flash-image - Cost: ~$0.039 per image
- Key:
GEMINI_API_KEYorGOOGLE_API_KEY
- Models:
gpt-image-2,gpt-image-1.5,gpt-image-1,gpt-image-1-mini - Cost: $0.04 (standard) / $0.08 (HD)
- Key:
OPENAI_API_KEY
Note
gen-image calls paid third-party APIs (OpenAI / Google) with your own key and is billed to your account. These are official provider SDKs; gen-image is an independent wrapper and is not affiliated with OpenAI or Google. Model names and pricing change over time, so check each provider's current docs.
✅ Image generated successfully
File: attachments/ownership.png
Size: 1024x1024
Cost: $0.04
Wikilink:
![[ownership.png]]
✓ Copied to clipboard
- If the output file already exists, gen-image auto-increments (
file-1.png,file-2.png, …). - The
![[wikilink]]is convenience for Obsidian/Markdown users and can be disabled in config.
There's a companion agent skill that drives this CLI from inside Claude Code (and other
skills-compatible agents):
gen-image in the insight marketplace.
It's a thin orchestration layer on top of the CLI: you ask in natural language ("make me 5 images showing the OAuth flow"), and the skill picks a sensible style, writes good prompts, generates batches in parallel, verifies the outputs, and hands back ready-to-paste image links. The CLI does the generation, cost tracking, and provider work; the skill does the prompting and orchestration.
# Claude Code
/plugin marketplace add grepinsight/insight-agent-marketplace
/plugin install gen-image@insight
# Any skills-compatible agent
npx skills add grepinsight/insight-agent-marketplace -s gen-image -g -yThe skill requires this gen-image CLI on your PATH (see Install above); installing
the skill does not install the CLI.
just test # pytest suite
just lint # ruffDocumentation site (MkDocs):
uv tool install mkdocs --with mkdocs-material
mkdocs serve # http://127.0.0.1:8000MIT.