Sprint Demo · July 24 – August 21, 2026

TechOps

Infrastructure · Legacy retirement

The warehouse is off Rundeck

Every warehouse job now runs on GitHub Actions, on a schedule you can read in a pull request.

  • Each schedule enabled here was disabled in Rundeck, one for one — no job and no report ever ran on both.
  • Every job is testable before it ships. Run it from a branch against dev, then merge. Rundeck's cadences and scripts only ever existed inside Rundeck.
  • Real alerting. Datadog watches every warehouse pipeline for failure, rather than Slack messages that need to be manually reconciled.
  • Runtime upgrades are back on the table — we took wanderu/reports from Python 3.9 to 3.14.
  • Jenkins no longer builds mammoth's Docker images, and its secrets moved from plain-text files to AWS Secrets Manager.
60 on live schedules (was 16)
74 warehouse jobs defined in code
gh-scheduler dashboard filtered to wanderu/mammoth: a status banner reading 60 scheduler cron, 0 gh cron, 14 manual, 0 late/stalled, filtered from 116, above the mammoth group of 74 jobs in Rundeck's own sections — Adyen, Cloudflare, Consolidated, Core Schema — each row on-time with its cron and next due time

All 74 mammoth jobs, in Rundeck's own sections. 60 on live crons, 0 late or stalled.

Reliability · Scheduled jobs

The scheduler now catches its own failures

One dashboard and one alert path for every scheduled job at Wanderu.

  • A job can opt in to one automatic retry before anyone is paged. Most stalls are just a reclaimed spot instance — those now recover quietly and read as Retried, not Failed.
  • Alerts scale to each job's cadence and land in the owning team's Slack channel.
  • The dashboard got usable: sortable, resizable columns and grouping that spans every repo.
Why it matters On August 10 Cloudflare silently stopped firing our every-minute cron for 71 minutes. We caught it — and the scheduler can now be driven from outside Cloudflare when the platform can't do it itself.
116 scheduled jobs, one dashboard
87 of them on scheduler crons
Two Slack messages from gh-scheduler-alerts: at 12:15 AM, a job dispatched but never started, queued 43 minutes, explaining that a stuck run holds its concurrency group and listing the owning team; at 1:00 AM, a green Queue cleared message for the same job

The alert explains the cause and names the owner — then posts its own all-clear.

Infrastructure · CI

GitHub Actions runner improvements — runs-on v3

Our self-hosted GitHub Actions runners moved from runs-on v2 to v3, with no cutover window.

  • Critical jobs can now opt out of spot entirely — every shared workflow takes a spot flag, and prod deploys default to on-demand. The seven longest mammoth jobs went first, picked by measured runtime.
  • The spot circuit breaker is tripped half as long — while it is, every job in the account pays for on-demand.
  • Two failure modes we used to see have not come back — being unable to schedule a runner, and failing to mint its credentials. v3 retries a registration name conflict instead of failing the job, and stops a stale app installation from breaking token refresh.
  • Machines that boot but never run their job halved — v3 keeps processing webhooks through a malformed payload instead of choking on it.
0 failures to schedule a runner, in 5,552 launches
−53% time with spot disabled fleet-wide
−49% machines that boot but never run their job
21.9s median wait for a runner, newly measurable

2,053 jobs a day, 16% more per hour than v2 — so these rates improved while the load went up. Normalized per 1,000 launches.

Security · Credentials and access

Snowflake password auth is all but gone

Snowflake pushed its password deadline out a month at the last minute. Our accounts were already moved; one shared credential is still to be cleaned up.

  • Donemammoth Kafka consumers
  • DoneNightly booking report
  • DoneModel-training reads
  • DoneSlack lookup commands
  • Up nextDrop the last shared credential
  • The legacy webhook box is gone. Slack's /ticket, /carrier and /station commands moved to a Cloudflare Worker — retiring two EC2 instances, a legacy load balancer and our last dockerhub.wanderu.com dependency with them.
  • Every script runs as its own account with only the permissions that script needs — not one shared credential sitting behind all of them.
  • RSA key pairs, encrypted in AWS Secrets Manager, in place of passwords living in config.
  • S3 ownership got simpler: cross-account writes now land owned by the bucket that holds them, so per-object ACLs stop being something anyone has to reason about.

Infrastructure · Across the stack

Smaller wins

  • Our Kubernetes lint step went from 196s to 14s. It had been compiling the linter from source on every single run; it installs a prebuilt binary now, and the config tool is cached instead of rebuilt.
  • Dev synthetic checks stopped silently dropping results — the machine running them was driving a headless browser under a 1 GB ceiling and getting killed about fifteen times a day.
  • Dev and pilot stopped disappearing. dev.wanderu.dev served nothing for 20 hours after a load balancer was replaced under it; those environments now publish a record that survives the swap.
  • Deploys stopped racing themselves — the cleanup job waits for in-flight post-deploy tests, and those tests now report which commit they actually ran against.

Vendors · GitHub

GitHub went down twice this sprint

August 6 took out Actions. August 17 took out most of GitHub for 7 hours 47 minutes — the platform we now run our CI and our scheduled jobs on.

What GitHub says

  • A core component in their Central US data centre failed to scale as traffic hit a new peak, and the pressure cascaded into authentication failures across github.com, Actions, the APIs, pull requests and Copilot.
  • They point at their own growth: monthly commits nearly doubled between April and August, 1.4 billion to 2.9 billion.
  • Their fix list: three million more CPU cores, 120 petabytes of storage, Azure taken from 12% to 58% of load since May, and retry budgets so one failure stops taking the rest with it.

Where we land

  • We are happy with the featureset and the price. The outages are the problem, not the product.
  • We are asking whether every egg belongs in this basket — and for now the answer is that it does. Nothing else in the space offers a comparable featureset.
  • Leaving would not buy us immunity anyway. Plenty of the upstream tools we depend on are themselves hosted on GitHub.