Preparing Your Hosting for Black Friday and Christmas Traffic Peaks

Preparing Your Hosting for Black Friday and Christmas Traffic Peaks Cloud

Black Friday has stretched into a fortnight, and the run-up to Christmas now starts somewhere around the second week of November. For anyone running an online shop, that means your hosting will meet traffic it has never seen before, on the same infrastructure that has comfortably handled a quiet Tuesday for eleven months. The good news is that most peak failures are predictable. They happen in the same handful of places, and they can be found and fixed in advance.

The work is unglamorous: measure the peak, test it, cache aggressively, cap what can be capped, and rehearse the failure. Here is how to approach it without rewriting your platform in October.

Model the peak before you change a single setting

Start with numbers, not adjectives. Pull last year's access and order logs and find the busiest hour, then scale it up by a factor you can defend from your own growth figures. Look at concurrency, not just visits: a hundred people browsing a cached category page is a very different load from twenty people checking out at once.

Write down the shape of the peak in plain terms. Roughly how many requests per second at the top of the hour? How many of those hit the database? What is your current cache hit ratio, and what would it need to be? Include the boring traffic too: sitemap crawls, stock feeds, admin dashboards, and the marketing team's scheduled email that lands at 9am.

Separate your journeys into three groups, because they stress different parts of the stack:

  • Read-heavy: category pages, product pages, search, blog content. These should be cached and served from the edge.
  • Write-heavy: add to basket, checkout, account creation, stock decrements. These hit the database and cannot be cached away.
  • Background: order confirmations, warehouse syncs, payment webhooks, analytics. These belong in a queue, not in the request cycle.

Load test with journeys that resemble your shop

A test that requests the home page ten thousand times tells you almost nothing. Use a tool such as k6, Locust, Gatling or JMeter and script realistic journeys: land on a campaign page, search, open a product, add to basket, apply a discount code, reach the payment step and stop there. Keep sessions and cookies so the application treats each virtual user as one person rather than a new visitor on every request.

Run it from outside your own network, ramp up gradually, hold the peak for at least fifteen minutes, then spike hard and see what recovers. Point it at a staging environment sized like production, or at a production environment during a quiet window with the payment provider sandbox in place. Tell your hosting provider and payment gateway beforehand if you are generating serious volume, or you may find yourself rate-limited or blocked.

The first bottleneck is usually one of four things: database connections, application workers, memory, or a single synchronous call to a third party. Fix the constraint you actually find, then test again. Load testing is a loop, not an event.

Set auto-scaling limits you can live with

Auto-scaling is only useful if it has been tested. Untested scaling policies are guesses, and they tend to fail at the worst moment.

Scale on a metric that reflects user pain: p95 response time, request queue depth, or worker saturation. CPU alone can sit low while requests queue behind a database lock.

Set the floor at your expected busy-day level, not your quietest overnight level. Cold starts during a spike are brutal, and instances that take ninety seconds to boot will not save you. Set the ceiling with your database in mind: if each application server opens a pool of twenty connections and you allow thirty servers, you have asked for six hundred database connections. Most managed databases will refuse well before that. Cap the pool, cap the servers, or move reads to a replica.

Finally, slow the scale-down. Keep extra capacity for at least twenty minutes after the peak passes, and pre-warm with scheduled scaling before a flash sale goes live.

Caching is cheaper than capacity

Every request you serve from cache is a request your database never sees. Work through the layers rather than chasing one magic setting.

  1. Turn on a full-page cache or reverse proxy for anonymous visitors, with sensible exclusions for basket and account pages.
  2. Move sessions and object fragments into Redis or Memcached so every server is not reading local files.
  3. Enable the opcode cache and check it is actually on; plenty of servers have it disabled after a migration.
  4. Send correct headers: long max-age for hashed assets, short TTLs for HTML, and stale-while-revalidate where available.
  5. Warm the cache before the sale. A cold cache meeting a spike is a self-inflicted outage.

CDN details that break under pressure

Most CDN problems are configuration problems. Check that static assets are served with hashed filenames and a long expiry, so a deploy does not force a global re-fetch. Be careful with cookie-based rules: a Vary: Cookie header on HTML can quietly push your hit ratio close to zero. Decide how query strings are handled, or campaign tracking parameters will fragment cached pages into hundreds of near-identical copies.

Test your purge process now, not on the morning a pricing error goes live. Know how long a purge takes, how you would roll back a bad release, and whether your TLS certificates renew automatically. Enable an origin shield if your provider offers one, so a cache miss from several edge locations becomes one request to your server rather than many.

The traffic you did not plan for

Third-party scripts sit on your critical path and you do not control them. A chat widget, review platform or tag manager that slows down or times out can make your pages feel broken while your own servers are healthy. Defer what can be deferred, load analytics after interaction, and set timeouts on anything your checkout calls.

Expect more bots as well. Scrapers, price comparison crawlers and card-testing attempts all climb in November. Add sensible rate limiting on search, login and checkout endpoints, and block obviously malicious patterns at the edge. Move email, PDF generation and stock syncs onto a queue so a slow job cannot hold a customer's confirmation page open.

Rehearse, monitor, and keep one page of notes

Run a short game day: what happens if the database fails over, the CDN has an incident, or the payment provider returns errors? Alert on what customers feel — p95 latency, error rate, queue depth, checkout completion — rather than CPU graphs alone. Give whoever is on call a single dashboard and a one-page plan that says who to contact and what to switch off first. Recommendations widgets, live chat and personalised banners are usually the cheapest things to disable.

Keep an eye on spend, too. Auto-scaling can be expensive if a bot flood keeps your ceiling pinned for hours, so set billing alerts alongside your performance ones. Do the work in October, leave the settings alone in December, and let the infrastructure be the least interesting part of your Christmas.

Photo: webandi / Pixabay