Serverless gets used to mean at least three different things, often in the same conversation. To some people it's a function you upload and forget about. To others it's any managed service where you don't pick a machine. The distinction matters, because the trade-offs you're signing up to depend on which one you mean — and there are workloads where serverless is plainly the right call, and plenty where it quietly costs more than a small server you'd have forgotten about anyway.
What "serverless" actually means
There are still servers. You just don't provision, patch, or scale them. Instead you hand the platform a unit of code or a configuration, and it runs when something triggers it: an HTTP request, a queue message, a schedule, a file upload, a row changing in a database.
Two shapes dominate. Functions as a service run a single handler with no routing of your own. Managed application platforms let you ship a normal container or framework app while the platform handles traffic and scaling. Both share the property that makes the model interesting: scale to zero. No traffic means no running instances and, in most cases, no charge. That's the line between serverless and "just use a managed VM".
Cold starts, without the hand-waving
When nothing is warm, the platform has to build an execution environment: fetch your code, start the runtime, run your initialisation. Requests arriving during that window wait longer than usual. For a lean Node or Python function that's typically tens to a few hundred milliseconds. For a heavy runtime, a bloated dependency bundle, or an application doing a lot of startup work, it can be well over a second.
You'll meet cold starts in two places. After an idle period, when the platform has let your instances go. And during a traffic spike, because every new instance starts cold by definition — the busiest moment is exactly when the platform is creating the most environments.
Warm instances do get reused for a while, which is why caching a database client outside your handler helps. But the retention window isn't something you control, so treat reuse as a bonus rather than a design assumption.
How to blunt cold starts
- Keep deployment bundles small. Strip unused dependencies and avoid importing a heavyweight SDK for one call.
- Do expensive setup lazily, and cache what you can outside the handler — connection pools, SDK clients, configuration.
- Pick a lighter runtime for latency-critical paths. Interpreted languages generally start faster than the JVM.
- Move slow work off the request path entirely: return quickly, push the job onto a queue, let a second function handle it.
- Use the platform's provisioned concurrency or minimum-instance setting if you genuinely need it — just remember you're now paying for idle capacity, which partly undoes the appeal of scale to zero.
Cold starts rarely matter for background jobs, webhooks, and admin tasks. They matter for the first page load of an interactive app, and they compound badly in a chain of functions where every hop can be cold.
The pricing model, and where the surprises live
Providers generally charge per invocation plus compute time measured in gigabyte-seconds, with data transfer, logging, and storage billed on top. The shape of that is straightforward: idle costs nothing, spiky traffic is cheap because you never pay for peak capacity in advance, and low-volume workloads are close to free.
Two things bite. The first is sustained load. A service handling steady, predictable traffic can be cheaper on a reserved instance or a container you keep running, because per-request pricing carries a premium for elasticity you're not using. The second is chattiness. One API call that fans out into five functions is five invocations, plus a queue, plus the logs from all of them. Verbose logging at volume shows up on the bill more often than people expect.
Serverless pricing rewards traffic you can't predict. It punishes traffic you can.
Model it against your real request volume before you commit. Free tiers are generous enough that a prototype tells you nothing about production cost.
Vendor lock-in is a spectrum, not a switch
Pure handler code wrapped around a standard HTTP framework is reasonably portable. The moment you adopt the platform's queue, scheduler, auth, database, and deployment tooling — which is where most of the productivity comes from — migration becomes real work. That isn't automatically a reason to avoid it. Accept lock-in where the service is a commodity you'd struggle to run better yourself, and keep the core business logic in plain modules with no platform imports. Wrap provider SDKs behind thin interfaces, and the boundary stays cheap to move if you ever need to.
Workloads serverless suits well
- Spiky or unpredictable HTTP APIs, especially early on when traffic is anyone's guess.
- Webhooks and third-party integrations, where reliability matters more than latency.
- Scheduled jobs: nightly reports, cleanups, data syncs, reminder emails.
- Event-driven glue: resize an image on upload, update a search index, notify a team channel.
- Internal tools and admin endpoints that get used a few times a day.
- Independent batch items that can be processed in parallel and retried individually.
Workloads you should keep on a server
- Long-running work that exceeds the platform's execution limit — video encoding, large exports, lengthy migrations.
- Steady, high-throughput services, where per-request billing adds up and a warm container is cheaper and more predictable.
- Latency-critical interactive paths with strict tail-latency targets.
- Stateful services: in-memory caches, sticky sessions, game servers, anything holding a persistent socket.
- Connection-heavy apps. Serverless instances multiply quickly and can exhaust a database's connection limit; a pooling proxy helps, but the fan-out is still real.
- Anything needing a tuned operating system, custom binaries, or a persistent local disk.
Most teams end up hybrid. Serverless for the edges, the events, and the glue; a small container or two for the core API and the long-running workers.
A practical way to decide
Stop arguing about serverless in the abstract and run the decision per workload. A workable sequence:
- Sketch the profile: how many requests or events, how steady, how latency-sensitive, how much data per call.
- Estimate cost at both peak and idle for a serverless option and a container option. Peak is where the answer often flips.
- Check the limits you'll actually hit — execution time, payload size, memory, concurrency caps — against your worst realistic case.
- Prototype the one path most likely to break the model: usually the latency-sensitive endpoint or the connection-hungry worker.
- Keep the business logic free of platform imports so the decision stays reversible for a year or two.
Serverless is a good default for things that run rarely or unpredictably, and a poor default for things that run constantly. Treat it as one tool in the drawer, chosen per workload, and it will serve you well.
Photo: valaymtw / Pixabay


