Roasbeef/jevlar

Typed, correct-by-construction Jev decisions for Go

★ 0Forks 0GoGitHub ↗Compare

README

jevlar

CI Go Reference

Typed Jev decisions for Go. The name is Jev + kevlar: the types are there to stop the invalid request or the misread answer before it goes anywhere.

Jev answers three kinds of question about a piece of content: a noul (a yes/no statement, answered with a probability), a choice (pick one of N labels, with a probability for each), and a score (rate against an ordered rubric, answered with an expected level). All three can be mixed in one request. jevlar makes each of those shapes a distinct Go type, so a choice can't be read as a score, a label you never asked for can't come back as an answer, and the selected alternative arrives as your domain type rather than a string you then have to switch on and hope you spelled right.

The library is stdlib only. It has one HTTP client with no retries, no credential discovery, and no logging; the pure request/response layer is exported for anyone who'd rather bring their own transport.

Install

go get github.com/roasbeef/jevlar

Requires Go 1.27 or newer: Client.Evaluate and Batch.Map are generic methods, which landed in 1.27.

A mixed batch, end to end

This is the same program that's compiled as Example in example_test.go, minus the fixture server. Set JEV_API_KEY and run examples/triage to see it hit the real service.

package main

import (
	"context"
	"fmt"
	"os"

	"github.com/roasbeef/jevlar"
)

// Queue is the application's own routing decision. Jev only ever sees the
// labels; the library hands back one of these values.
type Queue int

const (
	Billing Queue = iota
	Support
	Sales
)

// Triage is what we want to know about an inbound message.
type Triage struct {
	Route   jevlar.ChoiceAnswer[Queue]
	Urgency jevlar.ScoreAnswer
	Spam    jevlar.Probability
}

func main() {
	client := jevlar.NewClient(os.Getenv("JEV_API_KEY"))

	// Building a choice binds each label to a Go value. The alternative
	// count and label uniqueness are checked here, before any request
	// exists.
	route, err := jevlar.Choice(
		jevlar.Text("Which team should handle this?"),
		jevlar.Alt("billing", Billing).Describe(
			jevlar.Text("Charges, refunds, and invoices"),
		),
		jevlar.Alt("support", Support),
		jevlar.Alt("sales", Sales),
	)
	if err != nil {
		panic(err)
	}

	urgency, err := jevlar.Score(jevlar.Text("How urgent is this?"),
		jevlar.Text("Can wait"),
		jevlar.Text("Needs attention this week"),
		jevlar.Text("Needs attention today"),
	)
	if err != nil {
		panic(err)
	}

	// Three independent questions become one typed answer. The names are
	// the keys Jev uses in its reply; they never leak past this call.
	batch := jevlar.Map3(
		jevlar.Named("route", route),
		jevlar.Named("urgency", urgency),
		jevlar.Named("spam", jevlar.Noul(jevlar.Text("Is this spam?"))),
		func(r jevlar.ChoiceAnswer[Queue], u jevlar.ScoreAnswer,
			s jevlar.Probability) Triage {

			return Triage{Route: r, Urgency: u, Spam: s}
		},
	)

	eval, err := client.Evaluate(context.Background(),
		jevlar.Text("I was charged twice for last month. Please help."),
		batch,
	)
	if err != nil {
		panic(err)
	}

	// The switch is over our own type, so the compiler sees every case.
	switch eval.Answers.Route.Selected {
	case Billing:
		fmt.Println("route: billing")
	case Support:
		fmt.Println("route: support")
	case Sales:
		fmt.Println("route: sales")
	}
	fmt.Println("urgency level:", eval.Answers.Urgency.Nearest())
	fmt.Println("spam:", eval.Answers.Spam)
}

The three question types

Every constructor returns a Question[A], where A is the answer type. The question carries the only decoder that can produce that answer, so the compiler tracks which kind of answer you'll get back from each one.

Noul is a yes/no statement. jevlar.Noul(instructions) returns a Question[Probability]: the probability that the statement holds. There's no confidence field and no threshold; the library never turns a probability into a bool on your behalf, since you're the one who knows what a false positive costs. NoulWithCriteria(instructions, yes, no) lets you spell out what counts as each side (either may be nil).

Choice picks one of N labelled alternatives. jevlar.Choice(instructions, alts...) returns a Question[ChoiceAnswer[T]], with T being whatever type you put in the alternatives. An Alternative[T] is a label, your value, and an optional description; Alt(label, value) and .Describe(content) build one. The answer's Selected is your T, Label is the wire label (handy for logs), Confidence is Jev's own confidence, and Probabilities lists every alternative in the order you gave them, each paired with its probability. The constructor rejects zero alternatives, more than 255, an empty label, or a duplicate label, since any of those would be silently mangled by the JSON object the API expects.

Score rates content against an ordered rubric. jevlar.Score(instructions, levels...) returns a Question[ScoreAnswer]. Levels go from lowest to highest, and a level's position is its score, starting at zero. The answer's Value is the probability-weighted expected level (so 1.7 is perfectly normal), Nearest() rounds it to an integer level, and Legend and Probabilities are slices indexed by level, always exactly one entry per level you asked for. Rubrics must have two to ten levels.

Instructions, state, descriptions, and rubric levels are all Content. Only Text, Object, and Array implement it, which is exactly the set the API accepts at the top level; there's no way to send a bare number or null. jevlar.Marshal(v) turns a struct into Content (and refuses scalars), for when the state is structured rather than free text. On the answer side the same trick applies: Question[A] is constrained to the closed Answer interface (Probability, ChoiceAnswer[T], ScoreAnswer), so there's no Question[string] to accidentally build.

Batches

A Batch[A] is a set of named, independent questions whose answers combine into one value of type A. Named(name, q) lifts a single question into a batch. From there:

// Two to four heterogeneous questions into a struct.
jevlar.Map2(a, b, func(A, B) C) Batch[C]
jevlar.Map3(a, b, c, func(A, B, C) D) Batch[D]
jevlar.Map4(a, b, c, d, func(A, B, C, D) E) Batch[E]

// Any number of same-typed questions into a slice, in input order.
jevlar.All(batches...) Batch[[]A]

// Post-process an answer without touching the questions.
b.Map(func(A) B) Batch[B]

All is the one to reach for when the number of questions isn't known at compile time, e.g. scoring every passage in a document:

passages := make([]jevlar.Batch[jevlar.ScoreAnswer], len(docs))
for i, doc := range docs {
	passages[i] = jevlar.Named(fmt.Sprintf("p%d", i), relevance(doc))
}
eval, err := client.Evaluate(ctx, jevlar.Text(query),
	jevlar.All(passages...))
// eval.Answers is a []jevlar.ScoreAnswer, one per doc, in order.

Questions in a batch are independent: nothing in one question can see another's answer. A decision that depends on an earlier answer is a second request. Duplicate or empty names are caught when the request is built, and the response is checked to contain exactly the names that were asked for, nothing more and nothing less.

Client and transport

NewClient(apiKey, opts...) takes WithHTTPClient, WithBaseURL, and WithModel. The default model is jevlar.Latest (jev-latest), which is a moving alias; EvaluateWith pins a versioned name for a single call, and client.Models(ctx) lists what your account can use.

Every call is exactly one HTTP attempt. A non-200 status comes back as an *HTTPError with the status, headers, and raw body (422s carry a JSON list of what the service didn't like). HTTPError.Retryable() mirrors the official SDKs' policy (408, 429, 5xx) and RetryAfter() reads retry-after-ms or Retry-After for you, but the retry loop, the backoff, and the deadline are yours. A 200 that doesn't match the request's contract (a missing answer, a label that wasn't offered, a probability of 1.2) is an error wrapping ErrInvalidResponse.

If you'd rather not use the built-in client, EvaluationRequest(model, state, batch) and ModelsRequest() hand you a *Request[A] with Method(), Path(), and Body(). Do the exchange however you like, attach the bearer token yourself, and pass the 200 body to req.Decode. The decoder is bound to the request, so a response can't be decoded against the wrong batch. client.Do(ctx, req) is the same path with the built-in transport.

What's checked where

Whatever the type system can rule out, it does: answer type per question, domain type per choice, content shape, probability range. What it can't (alternative and level counts, label and name uniqueness, a blank model, a nil state) is checked once when the question or request is built, and never reaches the wire. On the way back, every answer's type tag, every label, every probability key and value, the score bounds, the legend, and the token counts are validated against the request that produced them. The distribution sum is allowed a tolerance of 0.01 per entry, since each value may be rounded independently. The library validates structure and bounds, not the truth of the model's judgement; a typed answer can still be wrong.

docs/API.md records the upstream contract this is built against and where the upstream documents disagree with each other.

Development

go test -race ./...
gofmt -l .
go vet ./...

The test suite covers wire encoding for all three question types, request building errors, malformed and inconsistent responses, HTTP failures and retry hints, and property-based round trips of generated choice and score questions via rapid. No API key is needed to run it.

License

MIT.

Contributors

Roasbeef

Issues