HenryNdubuaku/vector

Programming language for ML on XLA with JAX-level speed; compiles for CPU, TPUs/Nvidia GPUs, AMD GPUs, Apple GPUs, etc.

★ 15Forks 6RustGitHub ↗Compare
deep-learningfunctional-programmingjaxmachine-learningmathematical-programmingprogramming-languagepythonpytorchxla

README

Vector

Logo

A programming language for machine learning, compiled to CPUs, GPUs and TPUs through XLA.

ML in Python relies on special libraries with C/C++ backend; PyTorch, NumPy, JAX, Pandas. JAX is particularly fast, thanks to the XLA compiler which compiles for CPU, TPUs/Nvidia GPUs, AMD GPUs, Apple GPUs, etc.

Vector brings JAX-level speed across the entire program with Pythonic-syntax and functional paradigm, your entire code run on accelerators, not just the training loop.

  • 200 full-batch gradient-descent steps of a 1→1024→1024→1 tanh network on 2048 points of sin(x), f32.
  • JAX runs a jitted fori_loop; PyTorch runs both its standard eager loop and a torch.compiled step.
  • Timings are the median of 5 runs after one warm-up, excluding compilation.
Device Vector JAX PyTorch (eager) PyTorch (compiled)
Apple M5 Max GPU 0.27s — 0.32s 0.30s
NVIDIA RTX 4000 Ada 0.09s 0.09s 0.29s 0.25s
Google TPU v6e 0.01s 0.01s — —
Apple M5 Max CPU (ARM) 1.51s 1.62s 2.11s 2.15s

Overview

The tour below is one program: train a network to fit sin(x):

n = 16000
hidden_size = 1024
learning_rate = 0.03
epochs = 30
batch_size = 32
batches = 500

inputs = reshape(linspace(-pi, pi, n), n, 1)
targets = sin(inputs)

Vector modules are analogous to PyTorch nn.Module:

module Mlp(hidden):
  l1 = Linear(1, hidden)
  l2 = Linear(hidden, 1)

  forward(self, x):
    self.l2(tanh(self.l1(x)))

  loss(self, inputs, targets):
    mse(self(inputs), targets)

model = Mlp(hidden_size)

Training is whole-model arithmetic and the loop compiles to a single XLA op. Each epoch shuffles and slices minibatches, like a DataLoader with shuffle=True:

function train_epoch(model, inputs, targets, lr, n, batch, batches):
  perm = permutation(n)
  xs = inputs[perm]
  ts = targets[perm]
  for step in 0..batches:
    x = xs[step * batch : step * batch + batch]
    t = ts[step * batch : step * batch + batch]
    model = model - lr * grad(model.loss, x, t)
  model

for epoch in 0..epochs:
  model = train_epoch(model, inputs, targets, learning_rate, n, batch_size, batches)
  print(model.loss(inputs, targets))

Weights save as safetensors, tensors as numpy .npy:

save(model, "mlp.safetensors")
model = load("mlp.safetensors")

eval_inputs = reshape(linspace(-pi, pi, 9), 9, 1)
eval_targets = sin(eval_inputs)
print(model(eval_inputs))
print(eval_targets)

save(model(eval_inputs), "predictions.npy")
print(load("predictions.npy") - eval_targets)

The trained forward pass exports as StableHLO:

export(model, "mlp.mlir", eval_inputs)

You can serve the exported model over http:

vector serve mlp.mlir 8080
curl -d '{"inputs": [[[-3.14], [-2.36], [-1.57], [-0.79], [0.0], [0.79], [1.57], [2.36], [3.14]]]}' http://127.0.0.1:8080/

A table is a record of columns, saved and loaded as .csv, like pandas:

save({x: inputs, sin: targets, mlp: model(inputs)}, "predictions.csv")
table = load("predictions.csv")
print(mean(table.mlp - table.sin))

Plots are matplotlib-style, rendered as .svg:

plot(inputs, targets, "sin")
plot(inputs, model(inputs), "mlp")
title("sin approximation")
savefig("sin.svg")

Image processing capabilites are natively shipped:

grid = sin(linspace(-pi, pi, 64))
surface = 0.5 + 0.5 * matmul(reshape(grid, 64, 1), reshape(grid, 1, 64))
save(resize(surface, 32, 32), "surface.png")
imshow(load("surface.png"))
title("sin(x) * sin(y)")
savefig("surface.svg")

Audio is a record {samples, rate}:

tone = sin(linspace(0.0, 1382.3, 4000))
save({samples: tone * 0.5, rate: 8000.0}, "tone.wav")

Get Started

1. Install on any machine with a CPU, Nvidia GPU, TPU, and AMD GPU:

curl -fsSL https://raw.githubusercontent.com/HenryNdubuaku/vector/main/install.sh | sh && . "$HOME/.cargo/env"

2. Check the machine: trains a small model on the CPU and the accelerator; if anything is missing, vector prints the exact commands to fix it:

vector test

3. Run the tour: paste the overview cells into a file filename.vec:

vector filename.vec

Programs run on the machine's GPU or TPU automatically; add --cpu to force the CPU:

vector filename.vec --cpu

4. Read more: docs/reference.md covers the whole language; the examples train a GPT on Shakespeare and a vision transformer on MNIST, each in about a hundred lines.

Roadmap

When Goal
July 2026 Finish syntax and behaviour
August 2026 Parity with Python ML libs
September 2026 Large-scale distributed ML
October 2026 Vector libraries
November 2026 Release v1

Contributing

  • todo.md is the official list, pick anything from it.
  • The core items are ordered; the ecosystem tracks (notebooks, editor support, packages, docker) are self-contained and make good first contributions.
  • Follow the intuitive and minimalist coding established in the codebase.
  • cargo test must stay green: the golden tests are the merge gate, and the docs coverage test keeps the reference honest.

Contributors

HenryNdubuakukar-m

Issues