Precision,
not horsepower
Generic runners treat a build as a black box — shell scripts inside a fast VM. Lathe reads rustc, LLVM, macro expansion, monomorphization and the linker, then removes the work that should not be there.
curl -sSL https://lathe.run/install.sh | shcargo lathe initjobs: test: runs-on: [lathe-ubuntu-16core] steps: - uses: actions/checkout@v4 - uses: lathe/cache@v1 - run: cargo nextest run- rustc 1.70 → nightly
- cargo-nextest
- sccache
- mold
- cranelift
- cross
- wasm-pack
- x86_64
- aarch64
Stock toolchains from static.rust-lang.org, pinned by rust-toolchain.toml. No fork, no patched compiler.
Same workspace. Same commit.
Three real Rust workspaces, run on a generic managed runner and on Lathe with nothing else changed. Every row is a p50 across the full sample, and the methodology is under the table.
297 crates · 1.4M lines · proof assistant
| stage | generic runner | lathe | delta |
|---|---|---|---|
| Runner acquisition | 31s | 0ms | −100% |
| Cache restore, 18.4 GB | 94s | 0.21s | −99% |
| Cold build, full workspace | 2m 27s | 42.1s | −71% |
| Incremental, 3 crates changed | 48.2s | 6.4s | −87% |
| Link step, 1.2 GB binary | 22.8s | 8.1s | −64% |
| cargo nextest, 4,912 tests | 3m 04s | 1m 11s | −61% |
| Full pipeline, wall clock | 7m 48s | 2m 08s | −73% |
p50 of 1204 runs, 2026-07-01 → 2026-08-14. Generic runner: 16-core x86_64, 64 GB, sccache enabled, actions/cache warm. Full methodology and raw data at lathe.run/bench.
118 crates · 240k lines · async service
| stage | generic runner | lathe | delta |
|---|---|---|---|
| Runner acquisition | 29s | 0ms | −100% |
| Cache restore, 5.2 GB | 27s | 0.08s | −99% |
| Cold build, full workspace | 1m 04s | 24.7s | −61% |
| Incremental, 2 crates changed | 21.3s | 3.9s | −82% |
| Link step, 340 MB binary | 6.2s | 2.8s | −55% |
| cargo nextest, 1,847 tests | 58s | 26s | −55% |
| Full pipeline, wall clock | 3m 25s | 57.5s | −72% |
p50 of 880 runs, 2026-07-01 → 2026-08-14. Generic runner: 8-core x86_64, 32 GB, sccache enabled, actions/cache warm. Full methodology and raw data at lathe.run/bench.
842 crates · 6.1M lines · monorepo, LTO release
| stage | generic runner | lathe | delta |
|---|---|---|---|
| Runner acquisition | 34s | 0ms | −100% |
| Cache restore, 61.8 GB | 5m 12s | 0.74s | −99% |
| Cold build, full workspace | 11m 36s | 2m 51s | −75% |
| Incremental, 9 crates changed | 2m 42s | 19.4s | −88% |
| Link step, 3.8 GB binary, LTO | 1m 58s | 41.2s | −65% |
| cargo nextest, 22,104 tests | 14m 20s | 3m 48s | −73% |
| Full pipeline, wall clock | 36m 22s | 8m 20s | −77% |
p50 of 412 runs, 2026-07-01 → 2026-08-14. Generic runner: 32-core x86_64, 128 GB, sccache enabled, actions/cache warm. Full methodology and raw data at lathe.run/bench.
We win by understanding the compiler
Not by renting bigger machines. Every one of these changes what gets executed, not what it runs on. Roughly 70% of the available performance is left on the table by treating a Rust build as a shell script.
Keyed on what rustc computes, not on a hash of your lockfile
Generic caching hashes Cargo.lock and throws the whole target directory away when anything moves. We key on the fingerprint rustc itself computes for each compilation unit.
- uses: lathe/cache@v1 with: scope: workspace fingerprint: rustcBumping one dependency invalidates the crates that actually changed and nothing downstream of them that did not.
The target directory exists before the runner finishes booting
Restores are copy-on-write reflinks on the same filesystem, not a tarball pulled across a network. Nothing is decompressed and nothing is copied byte for byte.
$ lathe cache statsrestored 18.4 GB in 211msreflinked 281/297 cratesmissed 16 crates (source changed)Cache and compute sit in the same rack. A restore is a metadata operation.
Codegen units split by instantiation graph, not by module boundary
A generic-heavy crate compiles as one long pole no matter how many cores you rent. We partition its codegen units by where generics are actually instantiated.
$ cargo lathe test --explain-scheduletarski-core split 1 -> 7 CGUscritical path 41.2s -> 12.1sThe critical path shrinks instead of moving to a different machine.
Linking runs on a machine sized for linking
The link step is single-threaded, memory-hungry and I/O bound — the opposite of what a compile fleet is sized for. It runs on high-clock hardware with local NVMe while the fleet moves on to the next crate.
runs-on: [lathe-ubuntu-16core]lathe: link-tier: high-clockYou are not paying 32 cores to watch 31 of them idle.
Change one line
Lathe runners register as self-hosted runners against your existing GitHub Actions workflows. No new YAML dialect, no build wrapper, no rewrite.
- 01
Install the app
Grant the GitHub App access to the repositories you want to move. Read access to workflow files, write access to check runs. Nothing else.
- 02
Point one workflow at a runner
Change runs-on. The first build populates the cache; the second is the one worth measuring.
- 03
Read the schedule
cargo lathe explain prints where the time went — per crate, per codegen unit, per link. Keep the output whether or not you keep the runner.
jobs: test:- runs-on: ubuntu-latest+ runs-on: [lathe-ubuntu-16core] steps: - uses: actions/checkout@v4+ - uses: lathe/cache@v1 - run: cargo nextest runEvery build, dimensioned
A generic runner gives you one duration and a wall of log text. Lathe reports the schedule it actually executed: which crate was the long pole, which units were restored, and where the critical path went.
Every stage carries the unit it was measured in. The schedule is available as JSON at /api/v1/runs/{id}/schedule and as a check-run annotation on the pull request.
Three ways to run Rust CI
Self-hosted runners solve the hardware problem and hand you an operations problem. Generic managed runners solve neither. This is the whole comparison, including the parts that do not favour us.
| Lathe | Generic managed runner | Self-hosted fleet | |
|---|---|---|---|
| Execution | |||
| Cache keyed on rustc fingerprint | yes | lockfile hash | lockfile hash |
| Cache restore path | reflink, same host | network tarball | local disk |
| Codegen-unit scheduling | yes | no | no |
| Linker offload to high-clock tier | yes | no | build it yourself |
| Cold start, p50 | 0ms | 31s | 0ms |
| Operation | |||
| Crate-level build observability | yes | log text | log text |
| Concurrency ceiling | none | plan-capped | your fleet size |
| Runner patching and upkeep | ours | theirs | yours |
| Cache sizing and eviction | managed | manual, 10 GB cap | yours |
| Isolation model | microVM per run | microVM per run | depends |
| Cost | |||
| Billing granularity | per second | per minute, rounded up | per instance-hour |
| Idle capacity paid for | none | none | paid hourly |
| Engineer time to operate | none | none | 0.2–1 FTE |
Self-hosted is the right answer for some teams, and we will say so on a call. If your workspace is small enough that compile time is not the constraint, a generic runner is cheaper and you should keep it.
The questions an engineer asks first
Is this a fork of rustc?
No. Toolchains come from static.rust-lang.org and are pinned by your rust-toolchain.toml. We schedule and cache around the compiler; we do not patch it. Anything that builds on your machine builds here, and the binary is identical.
What has to change in my workflow?
One line: runs-on. Lathe runners register as self-hosted runners against your existing GitHub Actions workflows. The cache action is optional and replaces actions/cache — if you skip it you still get the runner and the scheduler, just not the reflink restore.
How is the cache keyed?
On the fingerprint rustc computes for each compilation unit — the same input rustc uses to decide whether to recompile. Not a hash of Cargo.lock. Bumping one dependency invalidates the crates whose fingerprint actually changed and leaves the rest reflinked.
Where does the cache live?
On the same NVMe as the runner, in the same rack. A warm restore is a copy-on-write reflink and crosses no network boundary. Cache is per-organisation, content-addressed, and encrypted at rest.
Nightly, custom targets, cross-compilation?
Any channel including nightly and dated nightlies, any target in the standard set, cross-compilation via cross. Custom targets with a JSON spec work if the spec is in the repository. MSRV pinning is respected as-is.
Private registries and git dependencies?
Both. Registry credentials and deploy keys are injected per run, scoped to that run, and never written to the cache. A private registry behind a VPC needs a peering connection; that is an enterprise conversation.
What is the isolation model?
One microVM per run, destroyed at the end of it. No shared kernel, no shared filesystem, no reuse of an instance between tenants. Cache blocks are content-addressed and namespaced per organisation.
What happens if we leave?
Change runs-on back. Your build is stock cargo and never depended on us. lathe cache export writes a plain target directory you can keep. There is no proprietary artefact format and nothing to migrate off.
Removes what should not be there
Point one workflow at a Lathe runner. If the median build does not drop, keep the data and leave.