Sprint Demo · September 4 – 18

TechOps

Data · MongoDB to DocumentDB

wapi reads from DocumentDB in production

The largest reader of our trip and reference data is off self-managed MongoDB. Two readers remain, then the old cluster can go.

  • Donewapi — exchange rates, carriers, cities, stations and trip search all read from DocumentDB
  • Up nextroute — one geospatial query to rewrite, then a connection-string change
  • Up nextgoogle-partner — one aggregation stage to remove, then a connection-string change
  • ThenWriters drop the MongoDB side and the legacy cluster shuts down
  • wapi's DocumentDB calls are now visible in APM. The driver was installed under a package name the tracer did not recognise, so every DocumentDB read since July was invisible and the error monitors had never received data.
  • The DocumentDB monitors now work. They page when errors pass a fixed count in five minutes, instead of watching a percentage that DocumentDB's volume kept near zero.
  • The alerts reach people. Every wapi monitor was notifying a Slack channel that did not exist. They now route to #infrastructure-alerts and #core-product-alerts.

Deployment · Continuous delivery

Eleven apps now ship themselves

Merge to main and production updates on its own, behind a post-deploy suite in every environment. When each app adopted Ship It:

April
experiment-router
google-partner
reservation-service
canopy
May
ui-react
route
June
next-prototypes
July
datadog-k8s
August
September
pricing-service New
payment-service New
nexus New
Next
wapi Tests
messaging Tests
4
6
7
8
8
11
 
11 of 27 app repos deployed to prod EKS ship on merge
  • payment-service and pricing-service joined this sprint with new post-deploy suites, and nexus went fully continuous with pilot back as a gate on every merge.
  • wapi and messaging are next, once their post-deploy tests pass in every environment.

Infrastructure · Legacy retirement

Ten legacy systems retired or on Kubernetes

The warehouse-etl host is empty, and the legacy production cluster is down to Snowplow and kafka-ui.

  • MovedAirbyte — production on the prod EKS cluster; dev and pilot retired
  • MovedSnowplow dev — collector on the dev EKS cluster at sp.dev.wanderu.dev
  • MovedWarehouse consumers — all 16 off the warehouse-etl host: 14 on EKS, two dead products retired
  • Stoppedblog-01 — the WordPress server, now that nexus serves the blog
  • Shut downBurrow, the legacy Docker registry, WKE Ops, Rancher, k3s and its database
  • Everything that moved is now watched. The legacy hosts shipped no logs and no metrics. Every process on EKS reports to Datadog, and the warehouse feeds each have a dashboard and an alert.
  • Airbyte kept everything. Same URL, same login, same connections and job history, and nobody noticed the move.
  • Snowplow dev captures its events again. 633 of 657 a day had been rejected on the new hosts before the collector fix.

Infrastructure · Legacy retirement

Legacy decommission: where we are

Thirty-five legacy systems are on the schedule. Most of the rest fall in chains, where one shutdown unlocks the next.

12 of 35 retired, up from 5 at the start of the sprint

Live report: techops.pages.wanderu.dev/reports/legacy-decommission/

  • NextLegacy production cluster — Snowplow prod and kafka-ui move, then it switches off
  • NextLegacy MongoDB — route and google-partner read from DocumentDB, then the writers drop it
  • ChainLegacy Kafka — app event streams and filebeat to MSK, then Kafka, CMAK and Opsdocker
  • ChainCarrier Data Automation and cortex automation — then Jenkins, Rundeck, BackDocker, Active Directory, BIND
  • RestData moves for three RDS instances, prod Redis and Kinesis; warehouse etl after the Mammoth pipelines; six stand-alone hosts