← Wanderu TechOps

Ship It — Continuous Deployment Adoption

Every application currently running in the wanderu-prod EKS cluster, and how it reaches production. Ship It is the reusable CD pipeline in wanderu/actions: on merge to the default branch it builds once, deploys dev and pilot in parallel, gates prod on post-deploy tests passing in both, then deploys and tests prod. This is where each service sits relative to that. Deploy and merge counts cover the 90 days ending 2026-09-04; tiers, workflow presence and prod test counts were re-checked on 2026-09-17.

27
app repos in prod EKS
11
fully continuous — deploys on merge, tests run, both companion workflows present
1
ship on merge but diverge from the Ship It contract
1
prod gate that reports success while executing zero tests
14
no Ship It pipeline
60%
of merges reached prod at all — 40% of them automatically
Continuous — deploys prod on merge, tests actually run, both companion workflows present
Degraded — ships on merge, but the pipeline is red, the gate is empty, or the implementation diverges from the Ship It contract
Manual — Ship It exists, no merge trigger
Tests only — post-deploy suite, no Ship It
None — manual eks-deploy only
Status Repo Prod workloads Prod deploys Merges Deploy rate ship-it.yaml On merge Prod tests run rollback-deployment reset-ship-it Adoption issue
Continuous canopy canopy 287 318 90% Yes Yes 47 / 288 Yes Yes —
Continuous next-prototypes next-prototypes 68 75 90% Yes Yes 24§ Yes Yes —
Continuous ui-react ui-react 61 63 96% Yes Yes 140 / 148 Yes Yes —
Continuous nexus nexus 30 268 11% Yes Yes 55 Yes Yes —
Continuous experiment-router experiment-router 26 29 89% Yes Yes 24 / 26 Yes Yes —
Continuous reservation-service reservation-service 19 14 100% capped Yes Yes 8 Yes Yes —
Continuous datadog-k8s datadog-operator, datadog-cluster-agent, synthetics-private-location 16 12 100% capped Yes (custom deploy) Yes 8 Yes Yes —
Continuous google-partner google-partner, google-partner-sandbox-main 10 10 100% Yes Yes 23 Yes Yes —
Continuous route route 1 2 50% Yes Yes 2 Yes Yes —
Continuous payment-service payment-service, payment-service-consumer 14 23 60% Yes Yes 22 / 28‡ Yes Yes —
Continuous pricing-service pricing-service 1 19 5% Yes Yes 22 Yes Yes —
Degraded messaging messaging 8 10 80% Yes Yes 6 Yes Yes #72
Manual wapi wapi 34 66 51% Yes No 0 / 16 Yes Yes #395
Tests only orders orders, orders-consumer 37 105 35% No No 31 No No #364
None pservs cachescheduler, papi, tripcache 26 27 96% No No 0 No No #582
None wtix wtix-server, wtix-generator, wtix-processor 14 21 66% No No 0 Yes No #606
None cortex cortex 10 23 43% No No 0 No No #134
None wboard wboard 6 11 54% No No 0 Yes No #83
None storage storage 5 11 45% No No 0 Yes No #82
None user user 2 3 66% No No 0 No No #30
None wanderlist-sorting wanderlist-sorting 2 3 66% No No 0 No No #88
None mcp-services mcp-services 1 4 25% No No 0 No No #161
None snowplow-k8s snowplow-collector, snowplow-enricher, snowplow-sinks 1 2 50% No No 0 Yes No #21
None snowflake-kafka-connector-k8s snowflake-kafka-connector-k8s 1 3 33% No No 0 Yes No #12
None google-tag-manager-k8s gtm-preview, gtm-tagging 1 3 33% No No 0 No No #8
None kafka-ui-k8s kafka-ui 1 3 33% No No 0 No No #9
None cloudflared-k8s cloudflared 0 0 — No No 0 No No #5

Prod tests run is tests that actually executed against prod in the most recent successful prod run, over tests the suite defines where the two differ. A bare 0 means no post-deploy suite exists at all.

Why wapi shows 0 executed, and why three other rows no longer do. Until 2026-09-04 a PR-state guard in the shared run-post-deploy-tests.yaml skipped the prod leg for every merged commit while still reporting success, which zeroed messaging, orders and pricing-service. wanderu/actions#360 narrowed the guard to transient review deploys, and the prod logs confirm it: orders executed 31 tests on 2026-09-14 and messaging 6 on 2026-09-04, both through the standalone workflow after manual deploys. pricing-service moved onto Ship It, which runs tests inline. wapi is the one remaining zero: its suite self-disables because no credentials reach the runner, tracked in wanderu/wapi#546.

Deploy rate is the Prod deploys column over the Merges column, so it can be checked against its neighbours. Under true CD every merge produces a prod deploy, so it reads as how much of the merge stream reached production at all, by any route. It says nothing about how: pservs is 96% entirely by hand, and nexus is 11% while auto-deploying every merge, because its 268 merges landed mostly before the trigger was switched on. The ratio is capped at 100%: a repo can reach prod more often than it merges — redeploys, rollbacks and manual dispatches all count — so reservation-service (19 deploys over 14 merges) and datadog-k8s (16 over 12) would compute above 100 and are shown as 100% capped. Those two cells are the only ones that do not equal the plain division. Across the fleet the ratio is 60% of 1,128 merges; the share that reached prod through a merge-triggered Ship It run rather than by hand is 40%.

§ next-prototypes runs 24 gating tests on prod plus one prod-only live BOS→NYC search E2E that is advisory — it is wrapped in if … else ::warning::, so a failure warns but does not gate promotion or trigger rollback.

‡ payment-service defines 28 post-deploy tests and skips 6 on prod by design: the card-issuing lifecycle and the payment lifecycle cases (3DS challenge, auth, capture, refund, void, declined auth) would move real money against live Adyen, so they run on dev and pilot only. The 22 that execute on prod are the intended gate, not a shortfall.

Adoption issue is the per-repo Ship It enablement ticket, tracked as a child of wanderu/actions#277. Repos already on continuous deploy show —.

What needs attention

1. nexus closed its last gap and is now continuous

nexus#877 enabled the push: trigger on 2026-09-04 and nexus#896 put pilot back on the path on 2026-09-08, gated behind its facts sync. Every merge since has shipped dev,pilot,prod green, and prod runs 55 post-deploy tests, all passing, because Ship It executes them inline.

The last divergence was the rollback filename: nexus#988 renamed rollback.yaml to rollback-deployment.yaml on 2026-09-17, so Ship It's Slack failure link resolves, and made a live prod rollback set SHIP_IT_HALTED — previously the next merge to main would have redeployed a commit a human had just backed out. nexus is now Continuous, and the prod run triggered by that merge shipped green with its 55 tests. nexus#348 is still open and can be closed.

The workflow is deliberately not the canonical wanderu/actions caller — it can also roll back the Payload migration batch that shipped with a deploy, which the shared workflow does not know about — so it declares neither target_env nor deployment_name. Ship It's automatic rollback-on-test-failure would need an input shim first, and stays off, as it does in every other Continuous repo.

2. canopy's prod gate does not cover the revenue path

47 of 288 collected tests run against prod. The dev-only gate (!!process.env.CI && DEPLOY_ENV !== 'dev') excludes every @integration checkout spec — payment, insurance, confirmation, processing, livecheck — plus the a11y live-API tests and edge.spec.ts. Prod exercises smoke, metadata, image-optimization and part of ticket-lookup across 4 browser projects. The exclusion is deliberate, keeping dev-scenario order ids out of prod, but the effect is that canopy auto-promotes 287 times a quarter without post-deploy coverage of checkout. The executed count also fell from 128 to 47 between 2026-09-02 and 2026-09-03 and has held at 47 across every prod run since, so the gate narrowed further without the tier changing.

3. messaging is red and its gate is empty

Five of the last six push-triggered runs failed. The last prod run that actually executed tests was 2026-06-16 — the only successful Ship It run in the repo's history — and the post-deploy-test.yaml prod runs between then and 2026-09-04 all passed while running zero tests, skipped by the PR-state guard. Since the guard fix the standalone run does execute its 6 tests, but only after manual eks-deploy, which is how prod stays current; the Ship It pipeline itself has not gone green since June. Its suite is also health-check-only (6 cases), below the Ship It coverage bar.

4. The companion-workflow gap is closed

Ship It treats both companions as required: it hardcodes gh workflow run rollback-deployment.yaml for auto-rollback and points the Slack failure link at that filename, and reset-ship-it.yaml is the only supported way to clear the SHIP_IT_HALTED breaker afterwards.

Four repos were auto-shipping to prod without a usable recovery path and have since been fixed — route#57 and datadog-k8s#35 added rollback-deployment.yaml, and experiment-router#128 added reset-ship-it.yaml. datadog-k8s needed a repo-specific rollback rather than the canonical caller: it deploys three Helm releases in the datadog namespace under a ClusterAdmin role, none of them named after the repo, so the shared workflow would have looked for a release that does not exist. nexus#988 renamed the last non-conforming one. Every repo that ships on merge now has both companions under the filenames Ship It resolves.

5. Thin gates that still auto-promote

route gates prod on 2 tests — one /rest2/trips 200 check and one health check. It shipped through Ship It again on 2026-09-04 once its rollback workflow landed, so the gate is live, not dormant. datadog-k8s's 8 tests assert only that deployments, the node daemonset and the private location are ready, with no assertion that Datadog data is actually flowing. Both were exercised on 2026-09-04. reservation-service's 8 are more substantive than the filename suggests — 2 health plus a 6-test reservation lifecycle including create, retrieve, invalid-UUID and fatal-close.

6. pricing-service and payment-service adopted; orders is the cheapest one left

pricing-service#292 and payment-service#456 both merged on 2026-09-16 with ship-it.yaml, both companions and a canonical test/post-deploy/ suite, and each first push shipped dev, pilot and prod green: 22 prod tests for pricing-service, 22 of 28 for payment-service with the money-moving cases skipped by design. Both are now Continuous.

orders is the remaining easy win: its 31 smoke tests now execute on prod after every manual deploy, so it already clears the hardest prerequisite and lacks only ship-it.yaml and the two companions. 106 merges produced 37 prod deploys, all by hand. Minor bug while in there: orders' deploy artifact reports base_url=orders.prod.wanderu.com, but the workflow's own resolve_url maps prod to orders.wanderu.com.

7. Deployment events are missing from Datadog for eight services

cortex, experiment-router, messaging, mcp-services, next-prototypes, snowflake-kafka-connector-k8s, google-tag-manager-k8s and the datadog-k8s workloads emit no env:prod DORA deployment events over 180 days, despite deploying. Change-tracking and deploy-correlated alerting are blind for those services. Counts in this report come from GitHub Actions for that reason.

8. Epic coverage, now closed out

wanderu/actions#277 now tracks 22 per-repo enablement issues, covering every app in this report. wanderlist-sorting (#88) and kafka-ui-k8s (#9) were missing and have been filed — the epic had recorded wanderlist-sorting as unfileable because its issues were disabled, which is no longer true. messaging (#72) was closed as adopted on 2026-06-16 and has been reopened: the run an hour after it closed is still the only successful Ship It run in that repo's history.

Two items in the epic remain stale but harmless: its "already have Ship It" reference list names five repos where thirteen now have a ship-it.yaml and eleven are fully continuous, and it covers filebeat-k8s and wordpress-k8s, which do not run in the prod EKS cluster and are outside this report's scope.

Method

App inventory: every Deployment in the wanderu-prod EKS cluster, mapped to its source repo by container image. Cluster add-ons not built from a Wanderu app repo (karpenter, coredns, aws-load-balancer-controller, external-dns, metrics-server, efs-csi-controller) are excluded. There are no StatefulSets or CronJobs in the cluster.

Prod deploys: successful GitHub Actions runs over the 90 days ending 2026-09-04 whose run name targets prod — Ship It runs including prod, plus eks-deploy runs to prod. Counted per repo, so a repo owning three workloads counts one deploy per run. Cross-checked against Datadog DORA deployment events, which agreed within 1% where both sources had data.

Merges: pull requests merged into each repo's default branch in the same window. Seven repos still default to master: pservs, cortex, storage, user, wanderlist-sorting, wapi, wtix.

Tier: Continuous requires all four of — ship-it.yaml present, triggering on merge to the default branch, prod post-deploy tests that actually execute, and both companion workflows (rollback-deployment.yaml and reset-ship-it.yaml) present. A pipeline that diverges from that contract is Degraded however green its deploys look, because Ship It's auto-rollback dispatch and circuit-breaker reset both resolve those exact filenames.

Deploy rate: the Prod deploys column divided by the Merges column, capped at 100% and labelled where the cap bites, so the figure is otherwise derivable from the two beside it. It measures whether merges reach production at all, not whether they get there automatically — the tier badge carries that judgement. An earlier revision of this report divided by merge-triggered Ship It runs only, which made the column impossible to check against its own row.

Prod tests: read from the most recent successful prod post-deploy job log, counting tests that actually executed and excluding skips. Static counts of defined cases are shown only as the denominator. This distinction is the whole point of the report: static counting would have credited nexus with 45 prod tests, orders with 31 and wapi with 16, when all three run zero. It also cuts the other way for Playwright repos that run one spec across several browser projects — canopy collects 288 and runs 47, because most @integration specs are gated to dev.