A Pre-Launch DevOps Checklist for New Web Applications

A Pre-Launch DevOps Checklist for New Web Applications DevOps

Launch day has a way of compressing a month of decisions into a single afternoon. The feature work is finished, the design is signed off, and then someone asks the question that matters: if this breaks at 2am, who finds out, and how do we get it back? Answering that properly is not glamorous work, but it separates a calm launch from a long night of guesswork.

What follows is the checklist worth running before pointing a production domain at a new application. Nothing here is exotic. It is the set of things that are easy to skip when you are close to the finish line.

Backups are only real once you have restored one

An automated backup job that has run for six months without complaint tells you almost nothing. What you need to know is how long a restore takes, what exactly is included, and who has permission to trigger it.

Before launch, confirm that database dumps, uploaded files and any object storage buckets are all covered. Check the retention window — thirty days is a common default, but a slow-burn data corruption bug can take longer than that to notice. Keep at least one copy in a separate account or provider, so a compromised credential does not take the backups with it.

Then do a restore. A full run into a scratch environment, timed with a stopwatch. Write the result down. If it takes four hours, that is the number you quote when someone asks, and it belongs in the runbook rather than in someone's memory.

Environment variables and secrets

Configuration drift between staging and production is one of the most common sources of launch-day surprises. The application reads a variable that was set locally, in a CI job, or in a dashboard, and production simply does not have it.

A few habits prevent most of this:

  • Keep every environment variable in one documented place, with a short note on what it does and a safe default where one exists.
  • Never commit secrets. A .env file belongs in .gitignore from the first commit, and any key that has ever been pushed should be rotated before launch.
  • Validate at boot. If a required variable is missing, fail loudly on startup rather than on the first request that needs it.
  • Check for staging leftovers: sandbox API keys, test email addresses, a debug flag still switched on.
  • Store production secrets in a managed secret store or your host's encrypted variable settings, and limit who can read them.

One extra step is worth the ten minutes it takes: compare the list of variables the running application actually sees against your documented list. The gap is usually where the missing value is hiding.

Monitoring: decide what healthy looks like

Monitoring is not a dashboard you glance at when something feels wrong. It is a set of signals that tell you whether the application is doing its job, plus a baseline so you can recognise normal on launch day.

Cover the basics first. An external uptime check hitting a health endpoint every minute, from more than one region. Response time percentiles rather than averages, because an average hides the slowest tenth of your users. Error rates from the application itself, not just the web server. And the resources underneath: CPU, memory, disk, database connections, queue depth.

Log retention matters too. When something goes wrong at launch, you will want at least a fortnight of searchable logs. Make sure timestamps are in UTC and that request IDs are logged, so a single request can be followed through the stack.

Alerting that reaches a person

An alert that lands in a channel nobody reads is decoration. Decide in advance who gets woken up, by what, and at what hour.

Keep the rules few and sharp. A handful of alerts that genuinely mean "someone must act now" will be trusted; forty noisy ones will be muted within a week. Route the urgent cases to a phone or pager, and everything informational to a channel.

Test the whole path before launch. Trigger a fake error and confirm the notification arrives with enough context to act on: which service, which environment, what the threshold was, and a link to the relevant dashboard or logs. Confirm the escalation path too. If the first person does not respond in ten minutes, what happens next?

A rollback plan you have practised

Deployment is usually the easy part. Getting back is where teams lose an hour.

Before go-live, you want immutable build artefacts — the exact thing running in production, stored somewhere you can redeploy in a single command. Blue/green or rolling deployments help here, but the important bit is that the previous version remains runnable.

Database migrations deserve their own paragraph. A migration that drops a column cannot be undone by rolling back the code. Prefer backwards-compatible changes: add the new column, deploy code that writes to both, backfill, then remove the old one in a later release. If a change is genuinely risky, put it behind a feature flag so behaviour can be switched off without a deploy at all.

If you cannot describe the rollback in three steps, you do not have one — you have a hope.

Write the steps down and walk through them once on a staging environment. Time it. Then keep the plan within reach for at least a fortnight after launch, because confidence is not the same as certainty.

Final checks before you switch the DNS

There is a cluster of small, boring items that cause disproportionate pain. Run through them in the last week:

  1. TLS certificates installed and set to renew automatically, with a calendar reminder as a fallback.
  2. DNS records and TTLs lowered in advance, so a change propagates in minutes rather than hours.
  3. Redirects from old URLs, www to non-www (or the reverse), and a sensible 404 page.
  4. Transactional email tested end to end, including SPF, DKIM and DMARC records, so receipts and password resets do not land in spam.
  5. Cron jobs, background workers and scheduled tasks running on the production host, with logs.
  6. Third-party services switched from test to live keys, with rate limits and quotas checked against expected launch traffic.
  7. Robots.txt and a sitemap in place, and analytics or consent handling configured if you use them.

Launch day, and the week after

On the day itself, have two people available and one runbook open. Deploy in the morning if you can. Problems are easier to solve when the people who built the system are awake and the support inbox is quiet.

After the switch, watch error rates, response times and the database for the first hour, then check in at the end of the day and again the next morning. Keep a short note of anything odd, even if it turned out to be nothing. That log becomes the first entry in your operational history, and the basis for the next checklist you write.

Photo: rawpixel / Pixabay