Micovec/rex

★ 0Forks 0RustGitHub ↗Compare

README

rex

A RAPTOR journey planner for public transit, written in Rust, built to run in two places at once: on a backend server and inside an Android app.

The routing core has no dependencies beyond std, so the same code that answers queries on a server cross-compiles to an .so and answers them on a phone with no network access. Both paths go through the same facade and produce the same JSON, so an offline answer and an online answer are the same answer.

                    ┌───────────┐
                    │ GTFS feed │
                    └─────┬─────┘
                          │  rex-build (once per feed release)
                    ┌─────▼─────┐
                    │ city.rex  │  binary timetable, loads in ~100 ms
                    └─────┬─────┘
              ┌───────────┴───────────┐
        ┌─────▼──────┐         ┌──────▼──────┐
        │ rex-server │         │  Android    │
        │   (axum)   │         │  (JNI/.so)  │
        └────────────┘         └─────────────┘

Crates

Crate What it is
rex-core The timetable model and the RAPTOR search. No dependencies except optional serde/bincode. This is the library.
rex-gtfs Imports GTFS feeds (directory or .zip) into a timetable.
rex-api Engine: loosely-typed requests in, journeys out. Shared by the server and Android so both behave identically.
rex-jni JNI bindings — the .so the Android app loads.
rex-server HTTP server (rex-server) and feed converter (rex-build, one or many feeds).
android/rex Kotlin library module wrapping the .so.

Quick start

# Convert a feed once. Output loads in milliseconds; GTFS import does not.
curl -O https://data.pid.cz/PID_GTFS.zip
cargo run --release --bin rex-build -- PID_GTFS.zip -o prague.rex

# Serve it.
cargo run --release --bin rex-server -- prague.rex --bind 127.0.0.1:8080

curl 'localhost:8080/plan?from=U123Z1&to=50.0875,14.4213&date=2026-09-02&time=08:30'

Measure it on your own feed:

cargo run --release --example bench -p rex-api -- prague.rex 2000

Merging several feeds

# Names come from the filenames; write name=path to choose your own.
rex-build pid=prague.zip jmk=south-moravia.zip -o merged.rex

Every id is namespaced by its feed's name, so feeds cannot corrupt each other — without that, two feeds that both number a stop 1 would silently become one, and the second feed's trips would call at the first's platforms. What joins the feeds back up is footpath generation: where two feeds describe the same physical stop, they sit at the same coordinates and a short walk appears between them. Do not disable --walk-radius when merging — nothing else connects the feeds, and rex-build warns if you do. Place grouping hides the seam in search too, so one "Hlavní nádraží" covers both feeds' platforms.

The calendar spans the union of the feeds' service windows.

If the feeds genuinely draw stop ids from one register — as Czech regional feeds are supposed to — --shared-stop-ids merges equal ids into one stop instead, which is better than two stops a metre apart joined by an invented footpath. That claim is verified, not trusted: if the repeated ids turn out to sit somewhere else entirely, the build fails rather than splicing unrelated places together.

$ rex-build --shared-stop-ids pid=prague.zip jmk=south-moravia.zip -o merged.rex
error: feed "jmk" was loaded as sharing a stop register, but 723 of its 723
repeated ids describe a different place (U1622Z1 is "Strážní" here and
"Černošice,Centrum Vráž" there, 186 km apart). These feeds only look alike;
load them without shared stop ids so each keeps its own.

Overlapping coverage is not deduplicated: if two feeds both describe the same regional bus, it appears twice and journeys along it are duplicated. Prefer feeds that partition the network.

Using the library

use rex_core::{Date, Query, Router, Time};

let timetable = rex_core::codec::load("prague.rex")?;
let mut router = Router::new();   // reusable scratch space; keep one per thread

let query = Query::new(from_stop, to_stop, Date::new(2026, 9, 2), Time::from_hms(8, 30, 0))
    .with_max_transfers(3);

for journey in router.search(&timetable, &query)?.journeys {
    println!("{}", journey.describe(&timetable));
}

search returns the Pareto set over arrival time and number of transfers — the direct-but-slow option and the fast-but-two-changes one, because which is better is the passenger's call, not the router's. Journeys come back fastest first; later entries arrive later but use fewer changes. (Sorting the other way round is tempting and wrong: on a real network the zero-change option can be a night bus sixteen hours later.)

Router::search_range answers a profile query instead: every distinct journey you could take by leaving within a window. It costs one full search per journey returned, plus a few to confirm them — so a profile query is roughly an order of magnitude dearer than a plain one, and that is where a slow response usually comes from. search_range_parallel spreads each of those searches over the machine's threads. Give it max_journeys when you only mean to show a few — it then stops as soon as those few are final, so widening the window costs nothing, and the results are exactly the ones a full sweep would have put first.

Prague, Anděl -> Airport, 09:00, at most 2 changes:

  rangeMinutes=60,  maxResults=5     14 searches    50 ms
  rangeMinutes=240, maxResults=5     14 searches    57 ms     same 5 journeys
  rangeMinutes=60,  no limit         48 searches   205 ms    all 29 journeys

Timetable is immutable and Sync; share one &Timetable across threads and give each thread its own Router.

Threads

There are two ways to use more than one core, and they are not interchangeable.

Across queries — share an Arc<Timetable>, give each thread its own Router, and plan different journeys at once. Nothing is synchronised, so this scales with cores. It is what a server should do, and what rex-server does by default.

Within one query — Router::search_parallel(tt, query) splits each round's route scans across every CPU thread the machine has, as the paper suggests. It takes no thread count: available_parallelism knows better than a constant, and rounds with too little work to divide fall back to one thread on their own. Use it when a single query's latency is what matters and the cores are otherwise idle — a phone planning one journey. On the Prague feed, 12 cores:

threads   median   p99
      1  10.7 ms  16.0 ms
      2  10.8 ms  16.2 ms     no better: coordination costs what it saves
      4   8.1 ms  12.6 ms
      8   6.5 ms  10.3 ms
     12   6.1 ms   9.6 ms     1.7x

1.7×, not 12×, and that is about the ceiling. Within a round the scans are independent, but the rest of a round is not: copying the previous round's labels forward (24 bytes per stop, ~4.7 MB per query here), relaxing footpaths, and applying the results are all serial. Below four threads the coordination costs more than it saves.

The parallel search returns exactly what the sequential one does — the same journeys boarding the same trips, not merely equally good ones. Scans do not write labels; they propose them, and the proposals are applied in queue order, which is the order a sequential round would have applied them in. A test asserts that over hundreds of random networks at 2, 3 and 8 threads.

Engine::with_parallel_queries(bool) chooses between the two, and is on by default. Router::search_parallel_on(tt, query, threads) pins the count, for benchmarks and tests that need to vary it.

Which to prefer depends on load, and the crossover is sharper than it sounds. Serving the same profile query on 12 cores:

concurrent requests      --parallel-queries        --sequential-queries
        1              40 req/s,  24 ms            27 req/s,  37 ms
        4              88 req/s,  45 ms            75 req/s,  51 ms
       12              76 req/s, 156 ms  (varies)  89 req/s, 131 ms
       32              71 req/s, 448 ms            84 req/s, 364 ms

Parallel is clearly better while requests do not overlap — two thirds the latency and half again the throughput. But each request spawns a thread per core, so twelve concurrent requests ask for 144 threads on 12 cores; past saturation it loses on throughput and latency, and its throughput becomes erratic (76 req/s on one run, 47 on the next) where sequential holds 89 req/s exactly. Peak throughput is much the same either way — the difference is where each reaches it and how it behaves beyond.

So: parallel by default, --sequential-queries under real traffic. The principled fix would be to size each query's thread count by how many requests are in flight; that is not implemented.

Turn the whole thing off with --no-default-features (feature parallel); it uses only std::thread, so it costs no dependency.

Android

cargo install cargo-ndk
rustup target add aarch64-linux-android armv7-linux-androideabi x86_64-linux-android
export ANDROID_NDK_HOME=~/Android/Sdk/ndk/27.0.12077973

scripts/build-android.sh              # writes android/rex/src/main/jniLibs/*/librex_jni.so

Add the module to your app's settings.gradle.kts:

include(":rex")
project(":rex").projectDir = file("../rex/android/rex")

Then drop prague.rex into src/main/assets and:

val engine = RexEngine.fromAsset(context, "prague.rex")

// Autocomplete a departure box; debounce and call off the main thread.
val suggestions = withContext(Dispatchers.Default) {
    engine.suggest("palmov", near = location.lat to location.lon)
}
// -> [PlaceSuggestion(name="Palmovka", stopCount=11, modes=[metro, tram, bus], ...)]

val response = withContext(Dispatchers.Default) {
    engine.plan(
        PlanRequest(
            from = Place.named(suggestions.first().name),
            to = Place.coordinate(50.0875, 14.4213),
            date = "2026-09-02",
            time = "08:30",
            maxTransfers = 3,
        )
    )
}

response.journeys.forEach { journey ->
    val lines = journey.rides.mapNotNull { it.route?.name }.joinToString(" → ")
    Log.i("rex", "${journey.departure}–${journey.arrival}  $lines")
}

RexEngine owns native memory: keep one for the life of the process and close() it when done. Queries are thread-safe and run in parallel, but on a large timetable they are not instant — call from a background dispatcher.

Errors that the request caused (unknown stop, date outside the timetable) arrive as RexException with a stable code; they are not exceptions in the native layer, so no try/catch is needed around ordinary "no journey found" cases — that is simply an empty journeys list.

The .so is built without the GTFS importer by default, since phones load a prebuilt .rex. Pass --features gtfs to cargo ndk if the app must import feeds itself.

HTTP API

Endpoint
GET /health liveness
GET /info what the loaded timetable covers
GET /places/suggest?q=&lat=&lon=&limit= autocomplete for a search box
GET /stops/search?q=&limit= individual platforms
GET /stops/near?lat=&lon=&radius=&limit= nearest stops
POST /plan body is a PlanRequest
GET /plan?from=&to=&date=&time=&… same, for links and testing

rex-server --parallel-queries spreads each query over every CPU thread. Off by default: serving requests concurrently already uses the cores, and gets more throughput out of them.

from and to accept a place name, a stop id, or a lat,lon pair. Optional parameters: maxTransfers, maxWalkSeconds, walkSpeedMps, modes (bus,tram,rail,…), rangeMinutes (profile query), maxResults, includeIntermediateStops.

Request errors return HTTP 400 with {"code": "...", "message": "..."}.

Autocomplete

A search box wants places, not the stops a feed actually contains. Prague's Palmovka is eleven stops — two metro platforms plus tram, bus and trolleybus stops spread over 331 m — and listing them individually is useless. rex-build groups stops by name into places, and suggest searches those:

typed: 'palmov'
  Palmovka                    11 stops  32 lines  tram,subway,bus,trolleybus
  Divadlo pod Palmovkou        1 stops  15 lines  tram,bus

Names are folded at build time, so search matches what people actually type on a phone: mustek finds Můstek, andel finds Anděl, and hl nadr finds Hlavní nádraží by matching the start of each word in any order. Without folding, all three of those return nothing.

Ranking is: how well the name matched, then how much service calls there, then distance from lat/lon if you pass them. That last part matters more than it sounds — typing nadrazi in Braník surfaces Nádraží Braník, a small station that never ranks nationally:

typed: 'nadrazi'
  no location    Nádraží Veleslavín, Nádraží Holešovice, Nádraží Vysočany
  at Veleslavín  Nádraží Veleslavín (0.5 km), Nádraží Podbaba (3.5 km), ...
  at Braník      Nádraží Braník (0.4 km), Nádraží Modřany (2.7 km), ...

Each suggestion carries a stable id, so a selection goes straight back in as {"from": {"placeId": "U529S1"}}. All eleven platforms then become search origins and the router picks whichever gets you there soonest — better than guessing a platform, and what multi-origin search is for. {"placeName": "Palmovka"} also works for links and hand-written requests.

Searching is a linear scan over the ~8,500 places of the Prague feed with no allocation: 0.8 ms, comfortably inside a keystroke, so there is no index to build or keep warm. Debounce on the client anyway and call it off the main thread.

Stops group into a place when they share a name and are within place_radius_m (default 1 km) of each other, transitively. Both halves matter:

  • Name alone merges the 89 Prague names that different villages share — Osek names two places 133 km apart, Chrášťany four.
  • A tight radius splits real interchanges. Palmovka is 331 m across and Hlavní nádraží 556 m, and neither has a single pair of stops close together; they hold together because the linking is transitive.

Measured across the feed, 1 km separates every reused name while keeping every genuine interchange whole, and produces no place wider than 1.1 km — the transitive rule does not run away. Places whose name is shared are flagged nameIsAmbiguous so a client can show the distance to tell them apart.

Times and dates

All times are feed-local and naive: 2026-09-02T08:14:00, no offset, no Z. Transit feeds are published in local time, and inventing a UTC offset would be inventing precision that is not in the data. The caller always supplies the service date; the library never reads a clock or a timezone database, which is also what keeps rex-core dependency-free and its results reproducible.

GTFS times past midnight (25:10:00) are preserved, and journeys crossing midnight work in both directions: the router probes the previous, current and next service day for every trip lookup. A journey found after midnight is reported against the following calendar date.

What it models, and what it does not

Supported: multiple candidate origins and destinations with access/egress walking, footpath transfers with per-stop change times, service calendars and exceptions, trips running past midnight, mode filters, transfer limits, and profile queries over a departure window.

Not supported: real-time delays, fares, wheelchair and bicycle constraints, and arrive-by queries. Adding arrive-by means a backward search, which is a real piece of work rather than a flag.

Notes on the implementation

A few decisions that are not obvious from the code:

Trips are grouped into "RAPTOR routes" by stop pattern, then split so that trips within one never overtake each other. RAPTOR's route scan assumes the earliest departure implies the earliest arrival; an express sharing a stop sequence with a local breaks that, so the builder puts them in separate routes.

Each service day is scanned separately. The non-overtaking property holds within a service day but not across days — tomorrow's 06:37 express leaves after tonight's late run and still arrives first. Tracking one current trip per probed day and taking the best label restores the invariant. This is subtle enough that tests/raptor.rs cross-checks the router against an exhaustive fixpoint search on several hundred random networks.

Most feeds need generated footpaths. Without them RAPTOR cannot change between a bus stop and the tram stop across the road, which matters more than any amount of tuning. rex-gtfs links platforms of the same station and generates footpaths within a walking radius; a transfer_type=3 row still wins.

Transfers naming a specific trip or route are ignored. All 11,234 of the Prague feed's transfer_type=1 rows carry a from_trip_id and to_trip_id: they are guaranteed connections between two particular runs, not footpaths. Applying them to the stop pair in general — at zero seconds, as transfer_type=1 implies — hands every journey a free platform change and understates travel times. Those platform changes fall back to real walking.

Walking legs are merged. The footpath sweep finds a walk as a chain of hops, and timed transfers contribute zero-second ones; a journey ending "walk 178 s, walk 0 s" is noise. Consecutive walking legs collapse into one, preserving total walking time.

The JNI boundary is JSON strings. Marshalling a journey graph field by field across JNI would be a lot of unsafe code maintained in two languages; a few kilobytes of JSON parses in well under a millisecond and lets both sides evolve independently.

Performance

Prague PID (https://data.pid.cz/PID_GTFS.zip) — the whole city plus the Central Bohemia region: 18,017 stops, 7,795 RAPTOR routes, 79,143 trips, 1.6M stop times. Release build, one core:

GTFS import (46 MB zip)   0.7 s
load .rex                  75 ms
median query              8.2 ms    (5.2 ms across 12 threads)
p99 query                12.2 ms
memory                     29 MB    (21.8 MB on disk)

92% of random stop pairs are reachable within 4 transfers. Roughly three quarters of a query is the route scan: ~237,000 route-stop visits across ~15,000 route scans and 4.9 rounds. Two things keep that number honest — the scan reads one dense ready array per stop rather than chasing a label, a parent and a 104-byte Stop for a two-byte change time, and each service day is dismissed with two comparisons against the route's departure window instead of a binary search per stop.

The rest is dominated by how generously walking is modelled, which is a modelling choice rather than a tuning one:

walking budget between vehicles     query    pairs reachable
  900 s (default)                  8.1 ms          92%
  300 s                            7.8 ms          92%
  120 s                            6.4 ms          86%

footpath radius at build time
  400 m (default)                  8.1 ms          92%
  250 m  (rex-build --walk-radius) 7.2 ms          92%
``` Queries are
single-threaded and independent, so a server scales across cores by running them
in parallel, and 29 MB is small enough to ship to a phone.

Reproduce with `cargo run --release --example bench -p rex-api -- prague.rex`.

## Watch out for the feed's calendar

`rex-build` reports which part of the calendar actually has service, because the
declared date range is usually much wider than the real one:

service: 14 of 365 days at full timetable (2026-09-02..2026-09-15), busiest 2026-09-03 with 50433 trips warning: the calendar spans 365 days but only 14 carry a full timetable


That is the Prague feed being normal, not broken: publishers commit full service
a fortnight ahead and leave a thin regional skeleton for the rest of the year.
Query a date outside the dense window and you will get few journeys or none —
which looks exactly like a broken router if nobody told you. `Timetable::
service_density` and `busiest_service_date` expose the same information
programmatically.

## Development

```bash
cargo test --workspace
cargo clippy --workspace --all-targets

License

MIT OR Apache-2.0.

Contributors

Micovec

Issues