Pythonium

Python, What else ?

Why scalability is not just a server problem

When an application starts getting more and more users, the first reaction is often quite simple: it needs a more powerful server. More CPU, more RAM, more bandwidth, or several servers behind a load balancer.

Sometimes, that's the right solution. But an application can also have scalability problems while the server is barely using 30% of its CPU.

The problem may be somewhere else.

CPU is not always the problem

Let's imagine an application handling 100 requests per second. Each request performs several operations: it reads data from the database, does some calculations, calls an external API, and then writes the result.

If the CPU is being used at 20%, you might think the application still has plenty of room. But if each request spends 200 ms waiting for a response from the database, adding more CPU won't change much.

The problem may simply be related to I/O. This is one of the first things I try to look at when an application becomes slower: what is it actually waiting for? Which operations are taking the most time? From there, we can identify the real bottleneck and choose the most appropriate solution.

A bad SQL query can cost more than a bad server

Databases are probably one of the first places where scalability problems appear. I've fixed quite a few database performance problems over the years. A query that works perfectly with 10,000 rows can become much more expensive when the table contains 10 million.

A simple SELECT, for example, may end up scanning an entire table when an index would allow it to directly retrieve the few rows being searched for. And even with the right indexes, some queries can become expensive when they have to perform several JOINs, sort large amounts of data, or calculate aggregations. Finally, the database configuration also needs to be properly tuned, as incorrectly configured parameters can limit the effectiveness of indexes.

In this kind of situation, replacing the server with a machine that is twice as powerful may provide some temporary relief, but it doesn't necessarily solve the problem. A better-optimized query or a suitable index can sometimes make a much bigger difference. Especially in terms of cost...

Optimizing a query is a first step. But sometimes, the best query is the one you never run.

Caching can completely change the problem

Caching is another example. Imagine that a page is viewed 10,000 times and performs exactly the same SQL queries every time.

Why query the database 10,000 times? If the result can be cached for a few seconds or a few minutes, a large part of that load can simply disappear.

Redis, an HTTP cache, a CDN, or even a very simple application-level cache can sometimes be much more effective than adding server resources.

But caching also brings its own problems: invalidation, data consistency, expiration, and memory usage. I invite you to read my article about the problems caching can cause.

So it's not always enough to put "a cache somewhere". But at the very least, it's a question worth asking.

Synchronous processing doesn't always scale

Another common problem is trying to do everything during the HTTP request.

Imagine that a user triggers an operation that needs to save an order, generate a PDF, send an email, call an external API, and update several systems.

If all of this is executed before responding to the user, the request can become very slow. And as traffic increases, each request consumes more resources for a longer period of time.

A message queue can be used to move some operations to the background. The user then receives a response quickly while a worker handles the longer-running tasks.

Solutions such as RabbitMQ, Kafka, SQS, or simply a queue based on a database can be used for this, depending on the requirements.

Once performance has been improved, another type of problem often appears: several users perform exactly the same operation at the same time.

Concurrency can also become a problem

An application can also work perfectly well with few users and start running into problems when several requests modify the same data simultaneously.

For example, imagine an application that needs to update the same record from multiple requests. If two transactions try to modify this data at the same time, they can conflict. The database then has to handle this concurrent access, notably through locks.

Locks help prevent several transactions from modifying the same data inconsistently, but they can also become a problem when there are many of them or when they are held for too long. A query may then remain blocked while waiting for another transaction to release a lock.

This type of problem has nothing to do with server power. Adding CPU or RAM won't allow a transaction to acquire a lock held by another transaction any faster.

You then need to look at how transactions are used, how long locks are held, the isolation level, and possibly rethink how the data is being modified.

More servers can even make things more complicated

Adding multiple instances of an application obviously makes it possible to increase its capacity. But it also introduces new constraints.

If the application stores files locally, what happens when the next request arrives on another instance? If a session is stored in memory on one server, how can it be retrieved on another? If several instances write to the same data simultaneously, how can conflicts be avoided?

We therefore often end up having to externalize certain components: shared storage, a database, distributed caching, shared sessions, message queues, etc.

Horizontal scalability is therefore not simply "running the same application twice".

The network can become the new bottleneck

Even when the code and database are fast enough, the network can become a problem. An API returning a few kilobytes does not have the same constraints as an API returning several megabytes with every request.

At scale, a few hundred extra kilobytes per response can represent a huge amount of traffic.

Compression, HTTP caching, CDNs, pagination, and on-demand loading then become important.

Once again, the application server is only one part of the system.

Scalability is a system problem

A scalable application is not simply an application that can use more CPU.

You need to look at the entire chain: application, database, network, storage, cache, external services, and asynchronous processing. And above all, you need to identify the real bottleneck before trying to solve it.

Adding servers is relatively easy (When we're not the ones paying for them ^^). Understanding why an application is slow is often much more interesting and, above all, more economical.

And sometimes, the best optimization isn't adding more power, but simply doing less work.




Laisser un commentaire