← Back to homeci-cd

Ephemeral CI runners on AWS, and the caching that makes them worth it

← All writing

We inherited a pipeline last year that cost more per month than the production infrastructure it deployed. Not because the builds were slow — because there were 340 of them a day across eight repos, all on 4-vCPU hosted runners, all reinstalling the same dependencies from scratch.

The fix was two changes that people usually treat as one project: move the compute to machines we control and shut them down when idle, and then put back the warmth that ephemerality takes away. Do only the first and you'll trade a predictable bill for slow builds. Do only the second and you'll have fast builds on expensive always-on hardware. This post is about doing both, with the numbers and the traps.

Two axes, not one

Every runner decision sits on a grid:

  • Who owns the machine — the CI vendor, or your AWS account.
  • How warm it starts — from a bare OS image, or from something pre-stocked with your toolchain and dependency caches.

Hosted runners are vendor-owned and lukewarm: GitHub bakes in Node, Python, Docker and a large tool cache, which is why they feel fine out of the box. The moment you self-host you lose that baked image, and the naive migration — Ubuntu AMI plus a bootstrap script — makes builds slower even though the hardware is faster. That surprise is the single most common reason teams roll back.

The shapes that work on AWS

ApproachStart-upCost profileBest for
Vendor-hosted~10s~$0.008/min per 2-vCPU minuteLow volume, OSS, no VPC access needed
EC2 ASG, long-lived0s (already up)Pay 24/7Teams with steady all-day load
EC2 per job (ephemeral, on-demand)60–120sPer-second billing, ~$0.35/hr for m6a.2xlarge in ap-south-1Most product teams
EC2 per job on Spot60–150sRoughly 60–70% off on-demandRetryable builds
ECS Fargate / EKS (ARC)20–45sPer-vCPU-second, no AMI to manageContainer-native pipelines
Lambda~1sSub-cent per buildTiny, short jobs only (15-min ceiling, no Docker)

For GitHub, we default to philips-labs/terraform-aws-github-runner for EC2-per-job, or Actions Runner Controller when the client already runs EKS. The EC2 path looks roughly like this:

module "runners" {
  source  = "philips-labs/github-runner/aws"
  version = "~> 5.10"

  aws_region = "ap-south-1"
  vpc_id     = data.aws_vpc.main.id
  subnet_ids = data.aws_subnets.private.ids

  runner_extra_labels = ["self-hosted", "linux", "x64", "lum-2xl"]
  instance_types      = ["m6a.2xlarge", "m6i.2xlarge", "m5a.2xlarge"]
  instance_target_capacity_type = "spot"

  ami_filter = { name = ["lum-runner-ubuntu-22.04-*"] }  # our baked AMI

  enable_ephemeral_runners = true
  idle_config = [{
    cron      = "* 3-12 * * 1-5"   # 08:30–18:00 Colombo
    timeZone  = "Etc/UTC"
    idleCount = 2
  }]
}

Two lines there do most of the work. enable_ephemeral_runners guarantees one job per instance — no state leaking between builds, which is the security argument for this whole exercise. idle_config keeps a small warm pool during working hours so the first push after standup doesn't wait two minutes for an instance to boot.

Jenkins gets there with the EC2 Fleet plugin against a Spot-enabled ASG. The knobs that matter:

minSize = 0
maxSize = 12
numExecutors = 1          // one build per node; keeps caches clean
idleMinutes = 5           // billing is per-second, but AMI boot isn't free
cloudStatusIntervalSec = 10
maxTotalUses = 1          // ephemeral

numExecutors = 1 costs you a little density and buys you reproducibility. We'll take that trade every time; parallel executors sharing a workspace directory is how you get a flaky pipeline nobody can debug.

Cold start is the tax on ephemerality

A fresh instance has no npm cache, no Docker layers, no Gradle artefacts. Left unaddressed, you pay that tax on every single build. Four layers of cache, outermost first:

1. The AMI. Bake it with Packer weekly: OS patches, the runner agent, Docker, Node via a version manager, and — critically — a docker pull of your two or three base images. Our runner AMI takes 11 minutes to build and removes about 70 seconds from every job that uses it.

2. Dependency caches in S3. On self-hosted runners, actions/cache still talks to GitHub's cache service over the internet, and you pay NAT Gateway egress for the privilege. Point it at your own bucket instead:

- uses: tespkg/actions-cache@v1
  with:
    bucket: lum-ci-cache-aps1
    key: pnpm-${{ runner.os }}-${{ hashFiles('pnpm-lock.yaml') }}
    restore-keys: |
      pnpm-${{ runner.os }}-
    path: |
        ~/.local/share/pnpm/store
        .next/cache

The restore-keys prefix is the part teams skip. Without it a single dependency bump is a total cache miss; with it you restore last week's store and fetch only the delta. And put an S3 Gateway VPC Endpoint on the private subnets — it's free, and it takes cache traffic off the NAT Gateway entirely. On one pipeline that alone was ~160 GB/month of NAT processing, about USD 7, plus meaningfully faster restores.

3. Docker layer cache. Never rely on the daemon's local cache on an ephemeral host. Use buildx with a remote cache backend:

docker buildx build \
  --cache-from type=registry,ref=$ECR/app:buildcache \
  --cache-to   type=registry,ref=$ECR/app:buildcache,mode=max \
  --push -t $ECR/app:$GIT_SHA .

Pair it with cache mounts inside the Dockerfile so package managers survive layer invalidation:

RUN --mount=type=cache,target=/root/.local/share/pnpm/store \
    pnpm install --frozen-lockfile

Add an ECR interface endpoint if your build pushes large images; the data transfer adds up.

4. Framework caches. For Next.js, .next/cache is the difference between a 40-second and a 4-minute build on a large app. Cache it keyed on the lockfile and restore it loosely — it's a heuristic cache, a stale one is still useful. Same logic for Turborepo/Nx remote caching, which we'd rather back with S3 than a third-party service.

One rule: cache the package manager's store, not node_modules. A 900 MB node_modules tarball takes longer to compress, upload and expand than a clean install from a warm pnpm store with hard links.

What it did to one pipeline

A Next.js monorepo, three apps, 340 builds/day at peak, migrated from 4-vCPU hosted runners to Spot m6a.2xlarge ephemeral runners with the cache stack above.

BeforeAfter
Median build (install → test → image push)11m 20s4m 05s
p95 build19m7m 10s
Queue wait, business hours40s12s (warm pool)
Compute cost / month~USD 980~USD 210
S3 + NAT + endpoint cost—~USD 26
Terraform to maintain0 lines~180 lines

Most of the speed came from caching, not from bigger instances. Most of the savings came from Spot and per-second billing. They're separable wins, and worth measuring separately so you know which one regressed when something regresses.

The honest downsides

  • You now run a fleet. Someone has to patch the AMI, rotate the GitHub App key, watch for runners that register and never take a job. Budget half a day a month.
  • Spot interruptions are real. With 3–4 instance types and ephemeral runners, we see under 1% of jobs killed mid-build in ap-south-1, but your deploy job should be idempotent and your CI should auto-retry. Keep release jobs on on-demand.
  • Cache poisoning is a genuine risk. A cache written by a PR branch and read by main is a supply-chain hole. Scope cache keys by ref for untrusted contributions, and never run pull_request_target workflows on runners with production IAM.
  • Cold starts never fully go away. A 3 a.m. hotfix will wait 90 seconds for an instance. Acceptable, but tell the team.
  • Debugging is harder. The machine that failed is gone. Ship runner logs to CloudWatch and keep a debug label pointing at a long-lived instance for the bad days.

Where we'd start on Monday

Don't migrate everything. Pick the noisiest repo and do it in this order, measuring after each step:

  1. Add proper cache keys with restore-keys on the existing hosted runners. Free, usually 30–50% of the total speed win.
  2. Move Docker builds to buildx with a registry cache.
  3. Stand up ephemeral EC2 runners with a stock AMI, on-demand, for that one repo.
  4. Bake an AMI. Switch to Spot with a mixed instance policy.
  5. Add the S3 gateway endpoint and move dependency caches into your own bucket.
  6. Only then roll the pattern out to the rest of the org.

The ordering matters because steps 1 and 2 are reversible and immediately visible, and they make the case for the infrastructure work that follows. If your build is slow because nobody set restore-keys, a fleet of Spot instances is an expensive way to avoid a two-line YAML change.

Enjoyed the read? We build this stuff for clients too.

Start a project