Ship It — Continuous Deployment Adoption
Every application currently running in the wanderu-prod
EKS cluster, and how it reaches production. Ship It is the reusable CD
pipeline in wanderu/actions: on merge to the default
branch it builds once, deploys dev and pilot in parallel, gates prod on
post-deploy tests passing in both, then deploys and tests prod. This is
where each service sits relative to that. Deploy and merge counts cover
the 90 days ending 2026-09-04; tiers, workflow presence and prod test
counts were re-checked on 2026-09-17.
eks-deploy only
| Status | Repo | Prod workloads | Prod deploys | Merges | Deploy rate | ship-it.yaml | On merge | Prod tests run | rollback- |
reset- |
Adoption issue |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Continuous | canopy | canopy | 287 | 318 | 90% | Yes | Yes | 47 / 288 | Yes | Yes | — |
| Continuous | next-prototypes | next-prototypes | 68 | 75 | 90% | Yes | Yes | 24§ | Yes | Yes | — |
| Continuous | ui-react | ui-react | 61 | 63 | 96% | Yes | Yes | 140 / 148 | Yes | Yes | — |
| Continuous | nexus | nexus | 30 | 268 | 11% | Yes | Yes | 55 | Yes | Yes | — |
| Continuous | experiment-router | experiment-router | 26 | 29 | 89% | Yes | Yes | 24 / 26 | Yes | Yes | — |
| Continuous | reservation-service | reservation-service | 19 | 14 | 100% capped | Yes | Yes | 8 | Yes | Yes | — |
| Continuous | datadog-k8s | datadog-operator, datadog-cluster-agent, synthetics-private-location | 16 | 12 | 100% capped | Yes (custom deploy) | Yes | 8 | Yes | Yes | — |
| Continuous | google-partner | google-partner, google-partner-sandbox-main | 10 | 10 | 100% | Yes | Yes | 23 | Yes | Yes | — |
| Continuous | route | route | 1 | 2 | 50% | Yes | Yes | 2 | Yes | Yes | — |
| Continuous | payment-service | payment-service, payment-service-consumer | 14 | 23 | 60% | Yes | Yes | 22 / 28‡ | Yes | Yes | — |
| Continuous | pricing-service | pricing-service | 1 | 19 | 5% | Yes | Yes | 22 | Yes | Yes | — |
| Degraded | messaging | messaging | 8 | 10 | 80% | Yes | Yes | 6 | Yes | Yes | #72 |
| Manual | wapi | wapi | 34 | 66 | 51% | Yes | No | 0 / 16 | Yes | Yes | #395 |
| Tests only | orders | orders, orders-consumer | 37 | 105 | 35% | No | No | 31 | No | No | #364 |
| None | pservs | cachescheduler, papi, tripcache | 26 | 27 | 96% | No | No | 0 | No | No | #582 |
| None | wtix | wtix-server, wtix-generator, wtix-processor | 14 | 21 | 66% | No | No | 0 | Yes | No | #606 |
| None | cortex | cortex | 10 | 23 | 43% | No | No | 0 | No | No | #134 |
| None | wboard | wboard | 6 | 11 | 54% | No | No | 0 | Yes | No | #83 |
| None | storage | storage | 5 | 11 | 45% | No | No | 0 | Yes | No | #82 |
| None | user | user | 2 | 3 | 66% | No | No | 0 | No | No | #30 |
| None | wanderlist-sorting | wanderlist-sorting | 2 | 3 | 66% | No | No | 0 | No | No | #88 |
| None | mcp-services | mcp-services | 1 | 4 | 25% | No | No | 0 | No | No | #161 |
| None | snowplow-k8s | snowplow-collector, snowplow-enricher, snowplow-sinks | 1 | 2 | 50% | No | No | 0 | Yes | No | #21 |
| None | snowflake-kafka-connector-k8s | snowflake-kafka-connector-k8s | 1 | 3 | 33% | No | No | 0 | Yes | No | #12 |
| None | google-tag-manager-k8s | gtm-preview, gtm-tagging | 1 | 3 | 33% | No | No | 0 | No | No | #8 |
| None | kafka-ui-k8s | kafka-ui | 1 | 3 | 33% | No | No | 0 | No | No | #9 |
| None | cloudflared-k8s | cloudflared | 0 | 0 | — | No | No | 0 | No | No | #5 |
Prod tests run is tests that actually executed against prod in the most recent successful prod run, over tests the suite defines where the two differ. A bare 0 means no post-deploy suite exists at all.
Why wapi shows 0 executed, and why three other rows no longer do.
Until 2026-09-04 a PR-state guard in the shared run-post-deploy-tests.yaml
skipped the prod leg for every merged commit while still reporting success, which
zeroed messaging, orders and pricing-service.
wanderu/actions#360
narrowed the guard to transient review deploys, and the prod logs confirm it:
orders executed 31 tests on 2026-09-14 and messaging 6 on 2026-09-04, both through
the standalone workflow after manual deploys. pricing-service moved onto Ship It,
which runs tests inline. wapi is the one remaining zero: its suite self-disables
because no credentials reach the runner, tracked in
wanderu/wapi#546.
Deploy rate is the Prod deploys column over the Merges column, so it can be checked against its neighbours. Under true CD every merge produces a prod deploy, so it reads as how much of the merge stream reached production at all, by any route. It says nothing about how: pservs is 96% entirely by hand, and nexus is 11% while auto-deploying every merge, because its 268 merges landed mostly before the trigger was switched on. The ratio is capped at 100%: a repo can reach prod more often than it merges — redeploys, rollbacks and manual dispatches all count — so reservation-service (19 deploys over 14 merges) and datadog-k8s (16 over 12) would compute above 100 and are shown as 100% capped. Those two cells are the only ones that do not equal the plain division. Across the fleet the ratio is 60% of 1,128 merges; the share that reached prod through a merge-triggered Ship It run rather than by hand is 40%.
§ next-prototypes runs 24 gating tests on prod plus one prod-only live BOS→NYC search E2E that is advisory — it is wrapped in if … else ::warning::, so a failure warns but does not gate promotion or trigger rollback.
‡ payment-service defines 28 post-deploy tests and skips 6 on prod by design: the card-issuing lifecycle and the payment lifecycle cases (3DS challenge, auth, capture, refund, void, declined auth) would move real money against live Adyen, so they run on dev and pilot only. The 22 that execute on prod are the intended gate, not a shortfall.
Adoption issue is the per-repo Ship It enablement ticket, tracked as a child of wanderu/actions#277. Repos already on continuous deploy show —.
What needs attention
1. nexus closed its last gap and is now continuous
nexus#877
enabled the push: trigger on 2026-09-04 and
nexus#896
put pilot back on the path on 2026-09-08, gated behind its facts sync. Every
merge since has shipped dev,pilot,prod green, and prod runs
55 post-deploy tests, all passing, because Ship It executes
them inline.
The last divergence was the rollback filename:
nexus#988
renamed rollback.yaml to rollback-deployment.yaml on
2026-09-17, so Ship It's Slack failure link resolves, and made a live prod
rollback set SHIP_IT_HALTED — previously the next merge to
main would have redeployed a commit a human had just backed out.
nexus is now
Continuous,
and the prod run triggered by that merge shipped green with its 55 tests.
nexus#348
is still open and can be closed.
The workflow is deliberately not the canonical wanderu/actions
caller — it can also roll back the Payload migration batch that shipped with a
deploy, which the shared workflow does not know about — so it declares neither
target_env nor deployment_name. Ship It's automatic
rollback-on-test-failure would need an input shim first, and stays off, as it
does in every other Continuous repo.
2. canopy's prod gate does not cover the revenue path
47 of 288 collected tests run against prod. The dev-only gate
(!!process.env.CI && DEPLOY_ENV !== 'dev') excludes every
@integration checkout spec — payment, insurance, confirmation,
processing, livecheck — plus the a11y live-API tests and
edge.spec.ts. Prod exercises smoke, metadata, image-optimization
and part of ticket-lookup across 4 browser projects. The exclusion is
deliberate, keeping dev-scenario order ids out of prod, but the effect is that
canopy auto-promotes 287 times a quarter without post-deploy coverage of
checkout. The executed count also fell from 128 to 47 between 2026-09-02
and 2026-09-03 and has held at 47 across every prod run since, so the
gate narrowed further without the tier changing.
3. messaging is red and its gate is empty
Five of the last six push-triggered runs failed. The last prod run that
actually executed tests was 2026-06-16 — the only successful Ship It run in
the repo's history — and the post-deploy-test.yaml prod runs
between then and 2026-09-04 all passed while running zero tests, skipped by
the PR-state guard. Since the guard fix the standalone run does execute its
6 tests, but only after manual eks-deploy, which is how prod stays
current; the Ship It pipeline itself has not gone green since June. Its suite is also health-check-only
(6 cases), below the Ship It coverage bar.
4. The companion-workflow gap is closed
Ship It treats both companions as required: it hardcodes
gh workflow run rollback-deployment.yaml for auto-rollback and
points the Slack failure link at that filename, and
reset-ship-it.yaml is the only supported way to clear the
SHIP_IT_HALTED breaker afterwards.
Four repos were auto-shipping to prod without a usable recovery path and have
since been fixed —
route#57
and
datadog-k8s#35
added rollback-deployment.yaml, and
experiment-router#128
added reset-ship-it.yaml. datadog-k8s needed a repo-specific
rollback rather than the canonical caller: it deploys three Helm releases in
the datadog namespace under a ClusterAdmin role, none of them
named after the repo, so the shared workflow would have looked for a release
that does not exist.
nexus#988
renamed the last non-conforming one. Every repo that ships on merge now has
both companions under the filenames Ship It resolves.
5. Thin gates that still auto-promote
route gates prod on 2 tests — one /rest2/trips
200 check and one health check. It shipped through Ship It again on 2026-09-04
once its rollback workflow landed, so the gate is live, not dormant.
datadog-k8s's 8 tests assert only that deployments, the node
daemonset and the private location are ready, with no assertion that Datadog
data is actually flowing. Both were exercised on 2026-09-04.
reservation-service's 8 are more substantive than the filename suggests —
2 health plus a 6-test reservation lifecycle including create, retrieve,
invalid-UUID and fatal-close.
6. pricing-service and payment-service adopted; orders is the cheapest one left
pricing-service#292
and
payment-service#456
both merged on 2026-09-16 with ship-it.yaml, both companions and a
canonical test/post-deploy/ suite, and each first push shipped
dev, pilot and prod green: 22 prod tests for pricing-service, 22 of 28 for
payment-service with the money-moving cases skipped by design. Both are now
Continuous.
orders is the remaining easy win: its 31 smoke tests now
execute on prod after every manual deploy, so it already clears the hardest
prerequisite and lacks only ship-it.yaml and the two companions.
106 merges produced 37 prod deploys, all by hand. Minor bug while in there:
orders' deploy artifact reports base_url=orders.prod.wanderu.com,
but the workflow's own resolve_url maps prod to
orders.wanderu.com.
7. Deployment events are missing from Datadog for eight services
cortex, experiment-router, messaging,
mcp-services, next-prototypes,
snowflake-kafka-connector-k8s, google-tag-manager-k8s
and the datadog-k8s workloads emit no env:prod DORA
deployment events over 180 days, despite deploying. Change-tracking and
deploy-correlated alerting are blind for those services. Counts in this report
come from GitHub Actions for that reason.
8. Epic coverage, now closed out
wanderu/actions#277 now tracks 22 per-repo enablement issues, covering every app in this report. wanderlist-sorting (#88) and kafka-ui-k8s (#9) were missing and have been filed — the epic had recorded wanderlist-sorting as unfileable because its issues were disabled, which is no longer true. messaging (#72) was closed as adopted on 2026-06-16 and has been reopened: the run an hour after it closed is still the only successful Ship It run in that repo's history.
Two items in the epic remain stale but harmless: its "already have
Ship It" reference list names five repos where thirteen now have a
ship-it.yaml and eleven are fully continuous, and it covers
filebeat-k8s and
wordpress-k8s, which do not run in the prod EKS cluster
and are outside this report's scope.
Method
App inventory: every Deployment in the
wanderu-prod EKS cluster, mapped to its source repo by container
image. Cluster add-ons not built from a Wanderu app repo (karpenter, coredns,
aws-load-balancer-controller, external-dns, metrics-server, efs-csi-controller)
are excluded. There are no StatefulSets or CronJobs in the cluster.
Prod deploys: successful GitHub Actions
runs over the 90 days ending 2026-09-04 whose run name targets prod — Ship It
runs including prod, plus eks-deploy runs to prod.
Counted per repo, so a repo owning three workloads counts one deploy per run.
Cross-checked against Datadog DORA deployment events, which agreed within 1%
where both sources had data.
Merges: pull requests merged into each
repo's default branch in the same window. Seven repos still default to
master: pservs, cortex, storage, user, wanderlist-sorting, wapi, wtix.
Tier: Continuous requires all
four of — ship-it.yaml present, triggering on merge to the default
branch, prod post-deploy tests that actually execute, and both companion
workflows (rollback-deployment.yaml and
reset-ship-it.yaml) present. A pipeline that diverges from that contract is
Degraded however green its deploys look, because Ship It's
auto-rollback dispatch and circuit-breaker reset both resolve those exact
filenames.
Deploy rate: the Prod deploys column divided by the Merges column, capped at 100% and labelled where the cap bites, so the figure is otherwise derivable from the two beside it. It measures whether merges reach production at all, not whether they get there automatically — the tier badge carries that judgement. An earlier revision of this report divided by merge-triggered Ship It runs only, which made the column impossible to check against its own row.
Prod tests: read from the most recent
successful prod post-deploy job log, counting tests that actually executed and
excluding skips. Static counts of defined cases are shown only as the
denominator. This distinction is the whole point of the report: static counting
would have credited nexus with 45 prod tests, orders with 31 and wapi with 16,
when all three run zero. It also cuts the other way for Playwright repos that
run one spec across several browser projects — canopy collects 288 and runs 47,
because most @integration specs are gated to dev.