Sprint Demo · July 24 – August 21, 2026
TechOps
Infrastructure · Legacy retirement
The warehouse is off Rundeck
Every warehouse job now runs on GitHub Actions, on a schedule you can read in a pull request.
- Each schedule enabled here was disabled in Rundeck, one for one — no job and no report ever ran on both.
- Every job is testable before it ships. Run it from a branch against dev, then merge. Rundeck's cadences and scripts only ever existed inside Rundeck.
- Real alerting. Datadog watches every warehouse pipeline for failure, rather than Slack messages that need to be manually reconciled.
- Runtime upgrades are back on the table — we took
wanderu/reports from Python 3.9 to 3.14.
- Jenkins no longer builds mammoth's Docker images, and its secrets moved from plain-text files to AWS Secrets Manager.
Reliability · Scheduled jobs
The scheduler now catches its own failures
One dashboard and one alert path for every scheduled job at Wanderu.
- A job can opt in to one automatic retry before anyone is paged. Most stalls are just a reclaimed spot instance — those now recover quietly and read as Retried, not Failed.
- Alerts scale to each job's cadence and land in the owning team's Slack channel.
- The dashboard got usable: sortable, resizable columns and grouping that spans every repo.
Why it matters
On August 10 Cloudflare silently stopped firing our every-minute cron for 71 minutes. We caught it — and the scheduler can now be driven from outside Cloudflare when the platform can't do it itself.
Infrastructure · CI
GitHub Actions runner improvements — runs-on v3
Our self-hosted GitHub Actions runners moved from runs-on v2 to v3, with no cutover window.
- Critical jobs can now opt out of spot entirely — every shared workflow takes a
spot flag, and prod deploys default to on-demand. The seven longest mammoth jobs went first, picked by measured runtime.
- The spot circuit breaker is tripped half as long — while it is, every job in the account pays for on-demand.
- Two failure modes we used to see have not come back — being unable to schedule a runner, and failing to mint its credentials. v3 retries a registration name conflict instead of failing the job, and stops a stale app installation from breaking token refresh.
- Machines that boot but never run their job halved — v3 keeps processing webhooks through a malformed payload instead of choking on it.
Security · Credentials and access
Snowflake password auth is all but gone
Snowflake pushed its password deadline out a month at the last minute. Our accounts were already moved; one shared credential is still to be cleaned up.
- Donemammoth Kafka consumers
- DoneNightly booking report
- DoneModel-training reads
- DoneSlack lookup commands
- Up nextDrop the last shared credential
- The legacy webhook box is gone. Slack's
/ticket, /carrier and /station commands moved to a Cloudflare Worker — retiring two EC2 instances, a legacy load balancer and our last dockerhub.wanderu.com dependency with them.
- Every script runs as its own account with only the permissions that script needs — not one shared credential sitting behind all of them.
- RSA key pairs, encrypted in AWS Secrets Manager, in place of passwords living in config.
- S3 ownership got simpler: cross-account writes now land owned by the bucket that holds them, so per-object ACLs stop being something anyone has to reason about.
Infrastructure · Across the stack
Smaller wins
- Our Kubernetes lint step went from 196s to 14s. It had been compiling the linter from source on every single run; it installs a prebuilt binary now, and the config tool is cached instead of rebuilt.
- Dev synthetic checks stopped silently dropping results — the machine running them was driving a headless browser under a 1 GB ceiling and getting killed about fifteen times a day.
- Dev and pilot stopped disappearing.
dev.wanderu.dev served nothing for 20 hours after a load balancer was replaced under it; those environments now publish a record that survives the swap.
- Deploys stopped racing themselves — the cleanup job waits for in-flight post-deploy tests, and those tests now report which commit they actually ran against.
Vendors · GitHub
GitHub went down twice this sprint
August 6 took out Actions. August 17 took out most of GitHub for 7 hours 47 minutes — the platform we now run our CI and our scheduled jobs on.
What GitHub says
- A core component in their Central US data centre failed to scale as traffic hit a new peak, and the pressure cascaded into authentication failures across github.com, Actions, the APIs, pull requests and Copilot.
- They point at their own growth: monthly commits nearly doubled between April and August, 1.4 billion to 2.9 billion.
- Their fix list: three million more CPU cores, 120 petabytes of storage, Azure taken from 12% to 58% of load since May, and retry budgets so one failure stops taking the rest with it.
Where we land
- We are happy with the featureset and the price. The outages are the problem, not the product.
- We are asking whether every egg belongs in this basket — and for now the answer is that it does. Nothing else in the space offers a comparable featureset.
- Leaving would not buy us immunity anyway. Plenty of the upstream tools we depend on are themselves hosted on GitHub.