The best email tool for LLMs.
email-cli turns .eml files and RFC 822 / RFC 5322 messages into stable,
agent-friendly JSON and prompt-ready text. It decodes the email machinery that
LLMs should not have to reason about: MIME boundaries, encoded headers,
charsets, quoted-printable bodies, base64 attachments, nested messages, reply
chains, and quoted-reply clutter.
It is deterministic, read-only, and built on top of
stalwartlabs/mail-parser.
Raw email is hostile input for LLMs. A single message can contain multiple body
alternatives, malformed headers, inline images, forwarded .eml attachments,
quoted replies, duplicated headers, suspicious filenames, and enough MIME syntax
to bury the actual conversation.
email-cli gives agents a cleaner contract:
- normalized message fields with source provenance
- decoded body text plus body-fragment metadata
- explicit quote handling for prompt rendering
- attachment metadata without inlining large binary payloads
- stable MIME-path part IDs for extraction
- reconstructed threads from supplied
.emlfiles - structured diagnostics instead of brittle stderr parsing
- JSON that works naturally with
jq
It does not summarize, classify, redact, send, mutate, index, or fetch mail. It turns email into dependable data so another tool, script, or model can do the next step.
From this repository:
cargo install --path .For development:
cargo run -- message.emlParse one message as JSON:
email-cli message.emlRender a message as prompt-ready text:
email-cli message.eml --format text --quotes collapseRead from stdin:
cat message.eml | email-cli
email-cli - < message.emlWhen run with no arguments from an interactive terminal, email-cli prints help
instead of waiting for stdin.
email-cli [FILE|-] \
[--format json|text] \
[--html] \
[--max-body-bytes N] \
[--quotes keep|collapse|drop] \
[--headers standard|all]
email-cli thread FILE... \
[--format json|text] \
[--html] \
[--max-body-bytes N] \
[--quotes keep|collapse|drop] \
[--headers standard|all] \
[--subject-fallback]
email-cli messages FILE... \
[--format json|ndjson] \
[--html] \
[--max-body-bytes N] \
[--quotes keep|collapse|drop] \
[--headers standard|all]
email-cli extract FILE --part <MIME_PATH_ID> [-o OUT]The default output is JSON with a project-owned schema, not a direct dump of
mail-parser internals. Every message includes:
schema_versionsourcepath, byte size, and SHA-256- normalized message fields
- threading hints
- decoded body text
- body alternatives and HTML availability metadata
- fresh/quoted body fragments
- ordered headers, including raw header text
- MIME parts and attachment metadata
- nested
message/rfc822summaries - structured diagnostics
Example:
email-cli message.eml | jq '.message.subject, .body.text, .parts'Output is token-lean by default. Headers are limited to a standard
identity/threading/content set, with the omitted count recorded in
headers_omitted; pass --headers all to include everything (Received chains,
DKIM/ARC signatures, X-* headers). The body alternative that produced
body.text (or body.html with --html) is not repeated inside
body.alternatives: its entry keeps the part metadata and hash, sets its
content field to null, and marks same_as: "body.text". body.text_part_id
names the MIME part the body text came from.
Message IDs are normalized without surrounding angle brackets. Dates are
normalized to UTC RFC3339 in message.date, while the original Date header is
kept in message.date_original.
Email is usually read as a conversation, not isolated records. email-cli thread
reconstructs threads across the .eml files you supply:
email-cli thread inbox/*.eml | jq '.threads[].timeline'Threading uses Message-ID, In-Reply-To, and References. It prefers
ID-based links, falls back from unresolved In-Reply-To to the last
References entry, and records missing parents or duplicate IDs as diagnostics.
Subject-only grouping is available, but intentionally opt-in:
email-cli thread inbox/*.eml --subject-fallbackAttached message/rfc822 parts are extractable messages, but they are not
silently added to a thread. Thread membership only comes from files explicitly
supplied to the command.
Default JSON includes attachment metadata, not attachment contents:
{
"part_id": "1.2",
"kind": "attachment",
"content_type": "application/pdf",
"filename": "invoice.pdf",
"safe_filename": "invoice.pdf",
"decoded_size_bytes": 12345,
"decoded_sha256": "...",
"extractable": true
}Use extract to write decoded bytes:
email-cli extract message.eml --part 1.2 -o invoice.pdfPart IDs are MIME-path IDs, not filenames or array indexes. This matters because filenames can be absent, duplicated, malicious, or unsafe as filesystem paths.
Text output is designed for feeding models directly:
email-cli thread *.eml --format text --quotes dropQuote handling:
--quotes keep: preserve quoted content--quotes collapse: replace quoted runs with deterministic markers--quotes drop: omit quoted runs from text output
JSON output always preserves fragment metadata, even when text rendering drops or collapses quoted content.
Use messages for scripts and pipelines:
email-cli messages *.eml --format json
email-cli messages *.eml --format ndjsonBatch commands keep going after per-file read or parse failures and include
structured diagnostics in the output. If every supplied input fails,
email-cli exits nonzero.
- Deterministic output: same input content and arguments produce the same data.
- Read-only operation: no mail is sent, modified, deleted, indexed, or fetched.
- Provenance first: output points back to source files, hashes, headers, body fragments, and MIME part IDs.
- Agent ergonomics: schema-versioned JSON, stable diagnostic codes, and predictable command shapes.
- Honest scope: email decoding and structure, not email understanding.
Run the test suite:
cargo testRun lint checks:
cargo clippy --all-targets --all-features -- -D warningsFormat code:
cargo fmt