← Back to blog

·Mergestorm Team·Engineering

Surge public beta: on-demand CI runners that scale to zero

Surge runs your GitHub Actions jobs on fresh cloud VMs dedicated to your repo, created when CI needs them and deleted when it goes idle. Minutes are included with Scale, Pro and Maelstrom plans. Here is how it works and how to set it up.

Pixel bar graph of Surge runners rising and dropping to zero on black, with the Surge mark and the line On-demand CI runners that scale to zero
SurgeGitHub ActionsSelf-hosted runnersCI runnersScale to zeroMergestorm

Mergestorm already reviews your pull requests (Vortex on every push, Tempest for a deep pass), patches clear findings (Cyclone), and lands stacks through the merge queue. Today it also runs the CI those steps wait on. Surge is in public beta.

Surge gives a private GitHub repository on-demand, self-hosted GitHub Actions runners. Each runner is a fresh cloud VM dedicated to your repo. Mergestorm creates it when CI needs it and deletes it after about 10 minutes without jobs. The pool scales to zero, and no Surge minutes are counted while it is idle.

Surge minutes are included with Scale (300 per period), Pro (1,000 per seat) and Maelstrom (3,000 per seat) plans; see Billing. Surge is a public beta: capacity is limited, and your jobs fall back to GitHub-hosted runners whenever Surge cannot take them.

Animation: Surge runners start when CI queues, finish the jobs, and scale back to zero

Why we built it

Two things about CI runners bothered us.

Idle runners cost money. A self-hosted runner that is always on bills around the clock, including nights, weekends, and the hours between pushes. Stopping the VM does not always fix that: at many providers, Hetzner included, a stopped server is still billed. The only runner that costs nothing is one that does not exist.

GitHub-hosted runners are small. The standard ubuntu-latest runner for a private repository has 2 vCPUs. A monorepo test suite with a few workspaces spends much of its time waiting for those two cores.

Surge answers both: bigger machines than the default runner, and none at all when there is no work.

How it works

You add one line to the jobs you want on Surge:

runs-on: ${{ vars.CI_SURGE_RUNNER || 'ubuntu-latest' }}

CI_SURGE_RUNNER is a repository variable that Mergestorm manages. Mergestorm sets it to your pool's runner label only while at least one Surge runner is online, and deletes it when the last runner goes away. When the variable is set, the job runs on Surge. When it is unset, the expression falls back to ubuntu-latest and the job runs on a GitHub-hosted runner exactly as it did before.

That fallback is the safety net for the whole product. If Surge is off, still starting, out of minutes, or the beta pool is full before your first runner is up, your jobs simply run on GitHub-hosted runners as usual. A workflow never waits on a pool that is not there.

The rest happens on our side:

  1. Demand. In Auto mode, Surge watches your repository's workflow runs and starts runners when jobs queue. For now, Auto needs the Mergestorm Cyclone GitHub App connected on the same repository, because that is how workflow events reach us. In Manual mode, you pick a runner count and Surge brings that many up; Manual works with the Surge app alone.
  2. Boot. Each runner is a new VM. It registers with GitHub as an ephemeral runner for a single job, runs it, wipes its workspace, and asks for the next one.
  3. Scale to zero. When a pool sees no jobs for its idle timeout (10 minutes by default, in either mode), Mergestorm turns routing off first, waits until nothing is busy or queued, and then deletes the servers. It never stops them, it deletes them.

A separate reaper process on another machine sweeps for leftover servers every 5 minutes and deletes anything past its maximum age. The rules for when routing may be on and when a server may be deleted are written down as a TLA+ model and checked, so a runner is never pulled out from under a job that is running on it.

What we measured on our own CI

Our own CI workflow moved to Surge on September 29. Because the runs-on fallback sends jobs to GitHub-hosted runners whenever Surge is off, the same jobs ran on both kinds of runner over the same two days, which gives us a like-for-like comparison.

We took the 200 most recent CI workflow runs (September 29 17:15 UTC to September 30 20:04 UTC) and every successful job among our four heavy jobs (Core CI, Frontend CI, and two workspace test shards), grouped by runner label:

JobsMedian run timeMedian wait before start
GitHub-hosted ubuntu-latest1384m 35s2s
Surge2172m 10s1m 55s

Once a job starts, it finishes in less than half the time on Surge. Run time is the job's started_at to completed_at from the GitHub Actions API.

The wait column is the honest part. A GitHub-hosted runner is already warm; a Surge runner that scaled to zero has to boot first, and that wait is most of the gap. End to end (wait plus run), the median job took 4m 7s on Surge versus 4m 37s on GitHub-hosted runners. The speed-up is real inside the job, and the cold start is the next thing we are working on (see below).

Set it up

You need a private repository (public and internal repositories are not supported in the beta), admin access to it on GitHub, and a Scale, Pro or Maelstrom plan with an active or trialing subscription. Free and Starter accounts have no Surge minutes, so Surge will not start runners for them.

  1. Install the app. Install the Mergestorm Surge GitHub App on the repository. It needs this so Mergestorm can register runners and manage the routing variable.

  2. Create a pool. In Mergestorm, open Dashboard › Surge. Enter owner/repo and select Create pool. Mergestorm checks that the app is installed, that the repository is private, and that you are an admin on it.

  3. Route your jobs. In your workflow, change runs-on on the jobs you want to move:

    jobs:
      test:
        runs-on: ${{ vars.CI_SURGE_RUNNER || 'ubuntu-latest' }}
        steps:
          - uses: actions/checkout@v4
          - uses: actions/setup-node@v4
            with:
              node-version: 22
          - run: npm ci
          - run: npm test
    

    Start with the jobs that spend the most time on CPU: test suites, builds, type checks. Small jobs that finish in seconds are fine where they are.

  4. Turn it on. Back on the Surge tab, pick Auto to have runners start when CI needs them and stop after the idle timeout, or Manual to start a fixed number of runners now. Auto needs the Mergestorm Cyclone GitHub App connected on the same repository for now; Manual works with the Surge app alone. Stop Surge turns everything off.

  5. Push. The next run picks up Surge as soon as a runner is online. The Surge tab shows each runner, its state, and the minutes it has used. Minutes used and left this period show on the Surge tab and on the Billing page.

The full reference is in the Surge docs.

Surge runners are fresh Ubuntu 24.04 VMs, not a copy of GitHub's ubuntu-latest image, so they do not carry every tool GitHub preinstalls. Use the setup-* actions for toolchains, as most workflows already do. Beta runners have no Docker and no sudo, so keep jobs that use service containers, container actions, docker build, or apt-get install on ubuntu-latest for now.

Isolation and security

  • One repository per pool, never shared. A pool belongs to one repository. Its VMs are never shared across customers or repositories.
  • No way in. Each VM is created with no SSH key, and its firewall allows no inbound traffic. Outbound traffic is limited to DNS, HTTP, and HTTPS.
  • No GitHub credential on the VM. The VM boots with a single-use token, trades it for a session token that can only request runner registrations for its own pool, and gets a fresh just-in-time registration for every job. The job itself runs as an unprivileged user.
  • Short-lived machines. The workspace is wiped after each job, idle servers are deleted, and every server is deleted once it reaches its maximum age, whatever it is doing.

Beta details and limits

  • Minutes come with paid plans. Surge minutes are included with Scale (300 per period), Pro (1,000 per seat) and Maelstrom (3,000 per seat) plans; see Billing. Each allowance is per credit period, while the subscription is active or trialing; Pro and Maelstrom count up to 20 seats. When a period's minutes run out, routing stops and your jobs run on GitHub-hosted runners until the next period.
  • Beta capacity is limited. When the beta pool is full, new runners wait for room. A pool that already has runners online keeps using them, and a pool with none online sends its jobs to GitHub-hosted runners as usual, thanks to the runs-on fallback.
  • Auto needs Cyclone for now. Auto mode needs the Mergestorm Cyclone GitHub App connected on the same repository. Manual mode works with the Surge app alone.
  • No Docker and no sudo on beta runners yet. Jobs that need service containers, container actions, docker build, or root keep running on ubuntu-latest.
  • Private repositories only, one pool per repository, and you need to be a repository admin to create it.
  • Runners run in the US (Ashburn, Virginia).
  • Cold starts. A pool at zero needs a couple of minutes to bring its first runner up, as the numbers above show.

What's next

These are on the roadmap. They are not shipped yet.

  • A regional dependency cache, so a runner that just booted starts with your packages warm instead of downloading them again.
  • More runners per VM, so one machine can take several jobs at once and a burst of jobs waits on fewer boots.

Try it

Open Dashboard › Surge and create a pool for a private repository. Change one runs-on line, push, and watch the runners come up and go back to zero. The Surge docs cover every setting. Surge minutes are included with Scale (300 per period), Pro (1,000 per seat) and Maelstrom (3,000 per seat) plans; see Billing.

Mergestorm now covers the whole path from a push to a merge: review, fixes, the merge queue, and the CI compute underneath them.