Sprint Demo · September 21 – October 2

TechOps

Deployment · Branch environments

A full branch environment in about three minutes

Your branch plus every service needed to exercise it, on its own hostnames in dev. It used to rebuild all of them. Now only your service builds.

190 s from 502 whole environment, with all nine required dependencies
61 s from 195 median deploy per service
Stop deploying your branch over main in dev It breaks everyone else's testing and no longer saves time. Use a branch environment instead.
  • BuildsDependencies reuse the image already built for their commit, so the build step is zero on every re-run. ui-react reads its config at runtime instead of rebuilding per environment
  • RolloutsMedian Kubernetes rollout fell from 78 s to 12 s: readiness is checked every 5 s instead of every 60 s
  • Coveragecanopy and nexus come along with the APIs they front, so a wapi branch is testable through checkout and C2C pages
  • CleanupEnvironments idle for 24 hours are removed nightly; re-running keeps yours alive

Reliability · Incident recovery

Every booking from the database incident is repaired

For 42 hours the shared booking database ran at 99% CPU. Bookings went through, but some trips never reached My Trips and their confirmation emails never sent.

356 trips restored to My Trips
237 confirmation emails sent for trips still to depart
100% of the 4,898 bookings in the window now have a complete record
  • Root cause found two days in: a change to the query in the void-auths job, which runs every 20 minutes. The new query never finished, so each run added another stuck query to the database. The query is fixed and each run is now time-limited.
  • Rows were replayed from the failed-query logs, rebuilt from the booking where the log was masked, and each affected session got one complete purchase event.
  • Production databases now page on-call. New RDS monitors page infrastructure on-call when a production instance runs above 85% CPU or past 300 connections.

Security · Carrier certificates

Amtrak and VIA Rail certificates now have one home

The Amtrak and VIA Rail client certificates our booking services present to the carriers now live in carrier-certs-k8s, with the keys in Secrets Manager.

Before
  • Scripts in opstoolbox
  • An Ansible role in the ops repo
  • A Jenkins job on the legacy cluster
  • Confluence runbooks
→
Now
  • Certs in git, keys in Secrets Manager
  • One deploy for dev, pilot and prod
  • Renewal and expiry checks as workflows
  • LiveDeploy refuses a key that does not match its cert, then restarts every workload that mounts one: papi, tripcache, cachescheduler and their branch deploys
  • LiveA monthly check posts to #infrastructure-alerts when a cert is within 90 days of expiry or a cluster has drifted from the repo
  • LiveRenewing the Amtrak dev cert is a workflow that signs the cert, opens the PR and produces the upload Amtrak asks for

Infrastructure · Legacy retirement

The route graph leaves legacy MongoDB and Jenkins

wapi and google-partner already read from DocumentDB. This sprint the job that builds our multi-leg routes stopped building from legacy MongoDB, and stopped running on Jenkins.

  • Movedupdate-routes runs as a GitHub Actions job and writes to the new routes DocumentDB cluster. The Jenkins job it replaces had reported success every week while writing nothing since February 2025
  • Oct 1Its station graph is built from supply DocumentDB. The legacy copy of stations had been going stale since cortex stopped writing to it
  • Sep 30Route indexes are created before the job writes; on the new clusters every write had been a full scan, so the rebuild slowed as it went
  • Sep 29In-account Valkey route caches are ready in dev, pilot and prod, for the jobs to move off the legacy Redis
  • Up nextThe first full production rebuild has been running since October 1. The route service reads from DocumentDB once its Mongo driver upgrade lands

Cost · Snowflake

Bringing Snowflake spend back to contract

We are trying to bring Snowflake from about $5,800 a month to the $4,167 our contract covers. We have a lot of ideas, and we think most of it is compute.

$5,810 to $4,167 September spend, and the monthly contract run rate we are aiming for
78% of September spend was warehouse compute
  • In reviewFirst fix: the hourly warehouse loads list all 20.9 million files in the ETL bucket before loading a handful. They will look only in the day's folder
  • IdeaKeep automated queries, such as Datadog's polling, from holding the analyst warehouse awake all day
  • IdeaRight-size the Amplitude export warehouse, and put spending limits on every warehouse
  • With ACQRun dbt on a schedule instead of continuously
  • IdeaMove old raw log tables out of Snowflake storage into S3