Cloudflare Load Balancing
What it actually does, explained twice — once for anyone in the room, once for the engineer.
In plain English
Load Balancing means you have more than one server capable of answering a request, and Cloudflare decides which one actually gets each request in real time. It constantly checks whether each server is healthy, and if one goes down, Cloudflare stops sending traffic there automatically, no one has to notice, page anyone, or flip a switch.
The two things customers actually care about: their site stays up even if one server or data center fails, and their site stays fast because requests go to whichever healthy server can answer quickest.
Think of it like a host at a busy restaurant. The host knows which tables are open and which section just closed for a cleaning emergency. Guests get seated at whatever open table makes sense right now, not the table they always used to sit at, and definitely not the section that's currently closed.
Capability 1
Spread the traffic out
Instead of one server handling everything, requests get split across several, so no single machine gets overwhelmed.
Traffic steering: off / random / weighted / geo / dynamic (latency) / proximity / least outstanding requests
Capability 2
Notice when something breaks
Cloudflare keeps knocking on each server's door in the background. The moment one stops answering, it's marked unhealthy and traffic stops going there.
Health monitors: HTTP/HTTPS & TCP probes on a defined interval, configurable success criteria
Capability 3
Keep the same user on the same server
For things like a shopping cart, you don't want a user bounced to a different server mid-session and losing their cart. This keeps them pinned to one.
Session affinity: cookie-based or HTTP header-based, TTL 30 min – 7 days (default 23h)
How it's actually built
Three components, in order
This is the real structure behind the dashboard, worth knowing before you configure anything or explain it to an engineer on the customer's side.
1. Health Monitors
Watch each server, mark it healthy or critical
→
2. Endpoint Pools
Groups of servers, usually one pool per data center or region
→
3. Load Balancer (VIP)
The entry point (usually a DNS record) that picks a pool, then picks a server inside it
A request comes in to the Load Balancer's entry point. It first applies a traffic steering policy to pick which pool should handle it (which data center/region), then applies an endpoint steering policy inside that pool to pick the actual server. If a pool or server fails its health check, both steps redo themselves automatically using whatever pool/servers are still healthy.
How To Actually Configure It
The real dashboard workflow, in order
This is the actual sequence Cloudflare's own setup flow follows, five steps, same order every time, whether you're doing it live on a call or on your own.
1
Go to Load Balancing and pick a hostname
In the dashboard, go to the Load Balancing page, select Create load balancer, then choose Public or Private. Pick the website/zone, then enter the hostname the load balancer will answer on (e.g. app.example.com).
dash.cloudflare.com → your account/zone → Load Balancing → Create load balancer
2
Create or attach a Pool
A pool needs a unique name, a description, and an Endpoint Steering choice (how it picks between servers inside itself). Then for each server you add: a unique name, its address or hostname, a weight, and optionally a host header, a destination port, or a Virtual Network (required if the endpoint has a private IP, e.g. behind a Tunnel).
- Gotcha: endpoint addresses must be unique within a pool, even across different ports or virtual networks. Overlapping IPs need separate pools.
3
Attach a Monitor (health check)
Pick a Type: HTTP, HTTPS, or TCP on non-Enterprise; Enterprise adds UDP, ICMP Ping, and SMTP. Set the path and port it checks, then tune interval, timeout, retries, and expected response code(s) under Advanced settings.
- Minimum check interval by plan: 60 seconds (Pro), 15 seconds (Business), 10 seconds (Enterprise), faster detection is a real, tangible plan-tier difference worth mentioning on a call about downtime.
- Retry timing gotcha: retries fire immediately, they don't wait for the next interval. With 5 retries and a 20s timeout, that's 6 total attempts (~120s) before an endpoint is marked unhealthy, plan your interval/retry combo around how fast you actually need failover to kick in.
4
Choose Traffic Steering and a Fallback Pool
This is where you pick the actual steering method from the table below (Off/Random/Geo/Dynamic/Proximity/etc.), and set a Fallback Pool, the pool used if every configured pool is unhealthy at once. Don't skip the fallback pool, it's the difference between graceful degradation and a hard outage during a total regional failure.
5
(Optional) Custom Rules, then Save and Deploy
Add any URL-path/header/ASN-based routing rules if needed, then save. The health monitor's status shows as unknown until the first check completes, that's normal, not a misconfiguration.
All of this is also API-driven, if the customer's team wants Terraform or CI/CD-managed load balancers instead of clicking through the dashboard. Required token permission: Load Balancing: Monitors and Pools Write. Good thing to mention if you're talking to an engineer instead of a manager.
Steering Methods
Which plan actually gets you which routing logic
This is the part that actually determines whether a customer needs to pay for anything beyond the base product. Source: Cloudflare's Load Balancing reference architecture docs.
| Method | What it does | Traffic steering | Endpoint steering | Plan |
| Off / Failover |
Strict priority order — first healthy pool/endpoint wins |
✓ |
— |
All plans |
| Random |
Random selection, can honor configured weights |
✓ |
✓ |
All plans |
| Hash |
Same source IP always lands on the same endpoint |
— |
✓ |
ENTERPRISE or self-serve add-on |
| Geo |
Ties pools to specific countries/regions |
✓ |
— |
ENTERPRISE or self-serve add-on |
| Dynamic (Latency) |
Builds a real round-trip-time profile per pool, routes to lowest RTT |
✓ |
— |
ENTERPRISE or self-serve add-on |
| Proximity |
Routes to the physically closest data center by GPS coordinates |
✓ |
— |
ENTERPRISE or self-serve add-on |
| Least Outstanding Requests |
Routes based on which endpoint currently has the fewest open requests |
✓ |
✓ |
ENTERPRISE or self-serve add-on |
Say it like this: "The basic version just needs someone to be alive and answering. The smarter versions, latency-based, closest-data-center, least-busy-server, are Enterprise, but you don't have to sign a full Enterprise contract just to get them, they're also available as a self-serve add-on if that's the only piece you need."
Other Details Worth Knowing
Session affinity, custom rules, and newer capabilities
- Session affinity only applies to Layer 7 (HTTP/HTTPS) load balancers. Cookie-based affinity's timer counts down from session start regardless of activity; header-based affinity resets the timer on every request. TTL ranges 30 minutes to 7 days, default 23 hours. Note: session affinity overrides weighted steering after the first request, once a user is pinned, weights stop applying to them.
- Custom Load Balancing Rules let you route based on URL path, ASN, headers, or other request attributes, similar to WAF custom rules. Non-Enterprise accounts get one rule per load balancer hostname by default. More rules require Enterprise.
- Monitor Groups (Enterprise) let you combine multiple health checks, like an HTTP check on the API gateway and a TCP check on a dependency, so a pool is only marked healthy if all of its critical components are actually working, not just the front door.
- Pool Sets (newer, Aug 2026) let one load balancer apply different steering policies to different geographies in a single config, e.g., dynamic-latency steering just for Germany while everywhere else uses standard failover order.
Gotcha to flag on a call: Load Balancing isn't on the Free plan. It's a paid product starting at Pro, priced per load balancer plus per monitored origin. Exact current dollar pricing lives in the dashboard's billing page, not something to quote from memory, pull it up live or loop in your account team.
Running This On a Call
The plain-English script, then when to go technical
- Start with the restaurant analogy or something equally simple. Most people in the room, even technical ones, want the concept before the mechanism.
- Ask what's actually failing today. "Is this about a server going down, or about traffic getting slow when it spikes, or both?" The answer tells you whether health checks/failover or steering method is the actual selling point, don't lead with steering policy jargon if their real problem is just "our one server fell over last month."
- Only bring up specific steering methods (Geo, Dynamic, Proximity) once they've described a specific symptom that method solves, global users complaining about latency is Dynamic/Proximity; a data residency requirement is Geo. Don't list all seven methods up front, it's overwhelming and most of the room won't retain it.
- If pricing comes up, be upfront that non-Enterprise access to the smarter steering methods exists as a self-serve add-on, it's a real option, not just an upsell to a full Enterprise contract.
Mistakes SEs Make Here
Anti-patterns that cost credibility on this specific product
1. Presenting every steering method as if they're all equally relevant
Seven methods sound impressive but confuse a discovery call. Diagnose the actual symptom first (downtime vs. latency vs. compliance), then present the one or two methods that solve it.
2. Implying Enterprise-only steering requires a full Enterprise contract
It doesn't. The reference architecture explicitly states these steering methods are also available as a self-serve add-on. Telling a Pro/Business customer they have to go full Enterprise to get Dynamic steering is factually wrong and needlessly kills a smaller, faster deal.
3. Forgetting session affinity overrides weighting after the first request
A customer testing a canary rollout with weighted steering will see skewed results if users are already session-pinned from a prior visit. Mention this explicitly when weighted steering and session affinity are both in play, it's a real gotcha that causes "why isn't my weight working" support tickets.