1. What it is
Load balancing is spreading incoming requests across several servers so no single one gets overwhelmed.
2. Everyday analogy
Picture a busy restaurant on a Friday night. (An imagined restaurant, for illustration.) There are five waiters, each covering a section of tables. At the front door stands a host. As each group walks in, the host decides where to seat them.
A careless host might seat everyone in the first section. That waiter drowns while the other four stand around. A good host spreads groups across sections. A great host also watches the room, notices which waiter is swamped, and sends the next group elsewhere. If a waiter goes home sick, the host simply stops seating people in that section.
On a website, the guests are requests from people's devices, the waiters are servers, and the host is a load balancer. The picture is a fair one, but real load balancers make these choices automatically, for every request. GitHub, for example, said in 2018 that it served "tens of thousands of requests every second" from the edge of its network. The next section covers how load balancers decide.
3. How it actually works
A load balancer is a piece of hardware or software that sits in front of a group of servers. Requests arrive at the load balancer first. For each one, it picks a server and passes the request along. Cloudflare's guide notes that it can be a dedicated device, or software running on a server or in the cloud.
How does it pick? Cloudflare describes two broad families of methods, called algorithms, which here just means rules for choosing. Static methods follow a fixed plan without checking how busy each server is. One simple static method is round robin: requests go to the servers evenly, in order (first request to server 1, next to server 2, and so on, then back to the start). It is the default method in Amazon's Application Load Balancer and in the NGINX web server. Amazon says it is commonly used when the requests and the servers are similar to one another. But if one server is running slowly, its queue can grow long. Dynamic methods look at the current state of each server. One example is least connection, which, as NGINX's documentation describes it, sends each request to the server with the fewest active connections. Dynamic methods keep the spread of work more even, but they are harder to set up.
The second big job is the health check. The load balancer regularly tests each server to see whether it is responding properly. Amazon's documentation for its Elastic Load Balancing service puts it this way: it "monitors the health of its registered targets, and routes traffic only to the healthy targets." When a server fails, traffic moves to the others. This is called failover.
Load balancing also makes scaling easier. Because people only ever talk to the load balancer, the team behind it can add servers when demand rises and remove them when it falls. Amazon's documentation notes that this can be done without disrupting the overall flow of requests.
4. Real-world example
GitHub built its own load balancer, called GLB, for its physical data centers. In a 2018 post on the GitHub Blog, GitHub's Theo Julienne described its design and released part of it, the GLB Director, as open source.
According to the post, GitHub was serving "tens of thousands of requests every second" out of the edge of its network at the time. GLB powered the majority of GitHub's public web and git traffic, and it also sat in front of some critical internal systems, such as its MySQL database clusters.
Two parts of the design connect directly to this lesson. First, GLB was designed to keep disruption to existing connections as small as possible when servers are added or removed. That is hard, because if the rest of a connection's data suddenly lands on a server that doesn't know about that connection, it fails. Second, GLB runs ongoing health checks. A service called glb-healthcheck keeps testing each backend server. When one fails, GLB shifts traffic away from it in a way that gives existing connections the best chance to move over smoothly.
The post goes much deeper into networking details that are beyond this course. The key idea is the one from section 3: spread the work, watch server health, and keep things running while servers come and go.
Source: The GitHub Blog, "GLB: GitHub's open source load balancer" (August 8, 2018): https://github.blog/engineering/infrastructure/glb-director-open-source-load-balancer/
5. Diagram
+--> [ Server 1 ] healthy
|
Users --> [ Load ] +--> [ Server 2 ] healthy
[ balancer] |
+-X [ Server 3 ] failed health check:
| no traffic sent
+--> [ Server 4 ] healthy
Users only see one address. The load balancer picks a server.6. When you'd care
Load balancing matters as soon as one server is not enough, or when you can't afford for one server's failure to take you offline. It goes hand in hand with horizontal scaling, covered in the scalability lesson.
If you work with a technical team, good questions to ask are: "What happens if one server dies?" and "Can we add capacity without downtime?" The load balancer and its health checks are often a big part of the answer.
A common mistake: assuming that simply spreading requests evenly is enough. As Cloudflare's guide points out, a static method like round robin doesn't notice when one server is slow, so some queues can still grow much longer than others. Even spreading works best when the load balancer also pays attention to server health and how busy each server is.
7. Check yourself
Q1. What is the main job of a load balancer?
- A) Translating domain names into IP addresses
- B) Spreading incoming requests across several servers
- C) Storing backups of a database
- D) Making a single server's processor faster
Show answer to question 1
Answer: B. A load balancer distributes requests so no single server is overloaded.
Q2. What does a health check let a load balancer do?
- A) Send traffic only to servers that are responding properly
- B) Repair broken servers automatically
- C) Make web pages smaller
- D) Speed up DNS lookups
Show answer to question 2
Answer: A. Health checks let it detect failing servers and stop sending them traffic.
Q3. Why can plain round robin sometimes cause problems?
- A) It only works with one server
- B) It ignores how busy or slow each server is, so some queues can grow long
- C) It requires users to pick a server themselves
- D) It turns off health checks
Show answer to question 3
Answer: B. Round robin follows a fixed order without checking each server's current state.
8. Sources
- The GitHub Blog, "GLB: GitHub's open source load balancer" (2018): https://github.blog/engineering/infrastructure/glb-director-open-source-load-balancer/
- Amazon Web Services, "What is Elastic Load Balancing?": https://docs.aws.amazon.com/elasticloadbalancing/latest/userguide/what-is-load-balancing.html
- Cloudflare Learning Center, "What is load balancing?": https://www.cloudflare.com/learning/performance/what-is-load-balancing/
- Amazon Web Services, "Edit target group attributes for your Application Load Balancer" (routing algorithms): https://docs.aws.amazon.com/elasticloadbalancing/latest/application/edit-target-group-attributes.html
- NGINX documentation, "HTTP Load Balancing": https://docs.nginx.com/nginx/admin-guide/load-balancer/http-load-balancer/