Sprint Demo · September 4 – 18
TechOps
Data · MongoDB to DocumentDB
wapi reads from DocumentDB in production
The largest reader of our trip and reference data is off self-managed MongoDB. Two readers remain, then the old cluster can go.
- Donewapi — exchange rates, carriers, cities, stations and trip search all read from DocumentDB
- Up nextroute — one geospatial query to rewrite, then a connection-string change
- Up nextgoogle-partner — one aggregation stage to remove, then a connection-string change
- ThenWriters drop the MongoDB side and the legacy cluster shuts down
- wapi's DocumentDB calls are now visible in APM. The driver was installed under a package name the tracer did not recognise, so every DocumentDB read since July was invisible and the error monitors had never received data.
- The DocumentDB monitors now work. They page when errors pass a fixed count in five minutes, instead of watching a percentage that DocumentDB's volume kept near zero.
- The alerts reach people. Every wapi monitor was notifying a Slack channel that did not exist. They now route to #infrastructure-alerts and #core-product-alerts.
Deployment · Continuous delivery
Eleven apps now ship themselves
Merge to main and production updates on its own, behind a post-deploy suite in every environment. When each app adopted Ship It:
April
experiment-router
google-partner
reservation-service
canopy
September
pricing-service New
payment-service New
nexus New
Next
wapi Tests
messaging Tests
Infrastructure · Legacy retirement
Ten legacy systems retired or on Kubernetes
The warehouse-etl host is empty, and the legacy production cluster is down to Snowplow and kafka-ui.
- MovedAirbyte — production on the prod EKS cluster; dev and pilot retired
- MovedSnowplow dev — collector on the dev EKS cluster at sp.dev.wanderu.dev
- MovedWarehouse consumers — all 16 off the warehouse-etl host: 14 on EKS, two dead products retired
- Stoppedblog-01 — the WordPress server, now that nexus serves the blog
- Shut downBurrow, the legacy Docker registry, WKE Ops, Rancher, k3s and its database
- Everything that moved is now watched. The legacy hosts shipped no logs and no metrics. Every process on EKS reports to Datadog, and the warehouse feeds each have a dashboard and an alert.
- Airbyte kept everything. Same URL, same login, same connections and job history, and nobody noticed the move.
- Snowplow dev captures its events again. 633 of 657 a day had been rejected on the new hosts before the collector fix.
Infrastructure · Legacy retirement
Legacy decommission: where we are
Thirty-five legacy systems are on the schedule. Most of the rest fall in chains, where one shutdown unlocks the next.
12 of 35
retired, up from 5 at the start of the sprint
Live report: techops.pages.wanderu.dev/reports/legacy-decommission/
- NextLegacy production cluster — Snowplow prod and kafka-ui move, then it switches off
- NextLegacy MongoDB — route and google-partner read from DocumentDB, then the writers drop it
- ChainLegacy Kafka — app event streams and filebeat to MSK, then Kafka, CMAK and Opsdocker
- ChainCarrier Data Automation and cortex automation — then Jenkins, Rundeck, BackDocker, Active Directory, BIND
- RestData moves for three RDS instances, prod Redis and Kinesis; warehouse etl after the Mammoth pipelines; six stand-alone hosts