multi-tenant ci for rust workloads

Precision,
not horsepower

Generic runners treat a build as a black box — shell scripts inside a fast VM. Lathe reads rustc, LLVM, macro expansion, monomorphization and the linker, then removes the work that should not be there.

curl -sSL https://lathe.run/install.sh | shcargo lathe init
.github/workflows/ci.ymlyaml
jobs:  test:    runs-on: [lathe-ubuntu-16core]    steps:      - uses: actions/checkout@v4      - uses: lathe/cache@v1      - run: cargo nextest run
one line changed
median build
42.1s−71%
297-crate workspace, n=1204
cold start
0ms
warm pool, all tiers
cache restore
94.2%
reflink, 30-day window
verified against
  • rustc 1.70 → nightly
  • cargo-nextest
  • sccache
  • mold
  • cranelift
  • cross
  • wasm-pack
  • x86_64
  • aarch64

Stock toolchains from static.rust-lang.org, pinned by rust-toolchain.toml. No fork, no patched compiler.

measured, not claimed

Same workspace. Same commit.

Three real Rust workspaces, run on a generic managed runner and on Lathe with nothing else changed. Every row is a p50 across the full sample, and the methodology is under the table.

297 crates · 1.4M lines · proof assistant

tarski @ 9f2c14erustc 1.83.016 vCPUn=1204
stagegeneric runnerlathedelta
Runner acquisition31s0ms−100%
Cache restore, 18.4 GB94s0.21s−99%
Cold build, full workspace2m 27s42.1s−71%
Incremental, 3 crates changed48.2s6.4s−87%
Link step, 1.2 GB binary22.8s8.1s−64%
cargo nextest, 4,912 tests3m 04s1m 11s−61%
Full pipeline, wall clock7m 48s2m 08s−73%

p50 of 1204 runs, 2026-07-01 → 2026-08-14. Generic runner: 16-core x86_64, 64 GB, sccache enabled, actions/cache warm. Full methodology and raw data at lathe.run/bench.

four places the time goes

We win by understanding the compiler

Not by renting bigger machines. Every one of these changes what gets executed, not what it runs on. Roughly 70% of the available performance is left on the table by treating a Rust build as a shell script.

94.2%restore rate

Keyed on what rustc computes, not on a hash of your lockfile

Generic caching hashes Cargo.lock and throws the whole target directory away when anything moves. We key on the fingerprint rustc itself computes for each compilation unit.

ci.ymlyaml
- uses: lathe/cache@v1  with:    scope: workspace    fingerprint: rustc

Bumping one dependency invalidates the crates that actually changed and nothing downstream of them that did not.

migration

Change one line

Lathe runners register as self-hosted runners against your existing GitHub Actions workflows. No new YAML dialect, no build wrapper, no rewrite.

  1. 01

    Install the app

    Grant the GitHub App access to the repositories you want to move. Read access to workflow files, write access to check runs. Nothing else.

  2. 02

    Point one workflow at a runner

    Change runs-on. The first build populates the cache; the second is the one worth measuring.

  3. 03

    Read the schedule

    cargo lathe explain prints where the time went — per crate, per codegen unit, per link. Keep the output whether or not you keep the runner.

.github/workflows/ci.ymldiff
  jobs:    test:-     runs-on: ubuntu-latest+     runs-on: [lathe-ubuntu-16core]      steps:        - uses: actions/checkout@v4+       - uses: lathe/cache@v1        - run: cargo nextest run
two lines added · one removed
observability

Every build, dimensioned

A generic runner gives you one duration and a wall of log text. Lathe reports the schedule it actually executed: which crate was the long pole, which units were restored, and where the critical path went.

passedrun_8f2c14e
tarskimain@9f2c14e297 crates2m 08s
critical path42.1s
restored281/ 297
recompiled16crates
cpu-seconds604s
stageschedulewall
acquire runnerwarm pool0ms
reflink restore18.4 GB0.21s
tarski-syntaxrecompiled3.8s
tarski-corecritical path · 7 CGUs12.1s
tarski-kernelrecompiled9.4s
tarski-tacticsrecompiled5.2s
link tarskihigh-clock tier8.1s
nextest, 4,912 tests12 shards1m 11s

Every stage carries the unit it was measured in. The schedule is available as JSON at /api/v1/runs/{id}/schedule and as a check-run annotation on the pull request.

the honest version

Three ways to run Rust CI

Self-hosted runners solve the hardware problem and hand you an operations problem. Generic managed runners solve neither. This is the whole comparison, including the parts that do not favour us.

Lathe compared with a generic managed runner and a self-hosted fleet
LatheGeneric managed runnerSelf-hosted fleet
Execution
Cache keyed on rustc fingerprintyeslockfile hashlockfile hash
Cache restore pathreflink, same hostnetwork tarballlocal disk
Codegen-unit schedulingyesnono
Linker offload to high-clock tieryesnobuild it yourself
Cold start, p500ms31s0ms
Operation
Crate-level build observabilityyeslog textlog text
Concurrency ceilingnoneplan-cappedyour fleet size
Runner patching and upkeepourstheirsyours
Cache sizing and evictionmanagedmanual, 10 GB capyours
Isolation modelmicroVM per runmicroVM per rundepends
Cost
Billing granularityper secondper minute, rounded upper instance-hour
Idle capacity paid fornonenonepaid hourly
Engineer time to operatenonenone0.2–1 FTE

Self-hosted is the right answer for some teams, and we will say so on a call. If your workspace is small enough that compile time is not the constraint, a generic runner is cheaper and you should keep it.

objections

The questions an engineer asks first

Is this a fork of rustc?

No. Toolchains come from static.rust-lang.org and are pinned by your rust-toolchain.toml. We schedule and cache around the compiler; we do not patch it. Anything that builds on your machine builds here, and the binary is identical.

What has to change in my workflow?

One line: runs-on. Lathe runners register as self-hosted runners against your existing GitHub Actions workflows. The cache action is optional and replaces actions/cache — if you skip it you still get the runner and the scheduler, just not the reflink restore.

How is the cache keyed?

On the fingerprint rustc computes for each compilation unit — the same input rustc uses to decide whether to recompile. Not a hash of Cargo.lock. Bumping one dependency invalidates the crates whose fingerprint actually changed and leaves the rest reflinked.

Where does the cache live?

On the same NVMe as the runner, in the same rack. A warm restore is a copy-on-write reflink and crosses no network boundary. Cache is per-organisation, content-addressed, and encrypted at rest.

Nightly, custom targets, cross-compilation?

Any channel including nightly and dated nightlies, any target in the standard set, cross-compilation via cross. Custom targets with a JSON spec work if the spec is in the repository. MSRV pinning is respected as-is.

Private registries and git dependencies?

Both. Registry credentials and deploy keys are injected per run, scoped to that run, and never written to the cache. A private registry behind a VPC needs a peering connection; that is an enterprise conversation.

What is the isolation model?

One microVM per run, destroyed at the end of it. No shared kernel, no shared filesystem, no reuse of an instance between tenants. Cache blocks are content-addressed and namespaced per organisation.

What happens if we leave?

Change runs-on back. Your build is stock cargo and never depended on us. lathe cache export writes a plain target directory you can keep. There is no proprietary artefact format and nothing to migrate off.

Removes what should not be there

Point one workflow at a Lathe runner. If the median build does not drop, keep the data and leave.