When people say “Node is single-threaded,” they usually mean that JavaScript in a typical Node process runs on one main event-loop thread. The rest of the system is still doing work at the same time.
A production request also involves the operating system, network, database, runtime internals, and sometimes worker threads. Node serves many requests by coordinating that work and keeping the main JavaScript thread free while other parts of the system are waiting.
Concurrency is not parallel JavaScript
Imagine three order requests:
Request A: send query ───── wait for PostgreSQL ───── handle result
Request B: send query ─── wait ─── handle result
Request C: call payment ─────── wait ───── handle result
The JavaScript thread performs short pieces of work for A, B, and C. During network waits, the operating system monitors sockets. When a response becomes available, the event loop eventually runs the relevant continuation.
All three requests are in progress at the same time, so they are concurrent. Whenever their JavaScript needs to run, the callbacks still take turns on the same event-loop thread.
await yields only for real asynchronous work
This handler can serve other requests while the database operation is in flight:
app.post('/orders', async (req, res) => {
const order = await orders.create(req.body);
res.status(201).send(order);
});
The promise returned by orders.create represents work that will finish later. While the database is working, await pauses this handler and gives control back to the event loop. When the result arrives, the rest of the handler is scheduled to run.
Now compare CPU-heavy work:
app.post('/reports', async (req, res) => {
const report = generateLargeReportSynchronously(req.body);
res.send(report);
});
Adding async to this function does not help because the calculation never reaches an asynchronous boundary. Every other JavaScript callback has to wait for it to finish.
The operating system handles most socket waiting
Network I/O does not require one libuv thread per socket. Node relies on operating-system mechanisms that can watch many descriptors and report which ones are ready.
This is how one process can keep thousands of mostly idle or occasionally active connections open without creating thousands of JavaScript threads. It is not unlimited. Memory, file descriptors, kernel queues, and the work performed for each connection still set the real limit.
The libuv thread pool has a specific role
Some operations do not have an efficient non-blocking operating-system interface that Node can use consistently. Libuv maintains a small worker pool for selected work such as many file-system operations, some DNS functions, and certain cryptographic tasks.
The pool is only used for specific operations. A normal network socket or database query does not take one libuv worker for the entire time it is waiting.
The pool can still become saturated. Several expensive password hashes or file operations may queue behind one another. A larger pool can help when measurements show that this is the bottleneck, but it also creates more competing work.
Microtasks can starve ordinary progress
Promise continuations and process.nextTick callbacks run with high priority around event-loop phases. A long self-replenishing chain can delay timers and I/O callbacks even though each individual step appears asynchronous.
function starve() {
Promise.resolve().then(starve);
}
Real applications usually create this problem through unbounded promise chains or recursive scheduling rather than a function this obvious. Moving work into another microtask still does not give timers and I/O callbacks a fair chance to run.
CPU work needs another execution strategy
For CPU-intensive JavaScript, common choices are:
- split the work into bounded pieces and yield between them;
- use Node worker threads for parallel computation;
- run a separate process or service;
- move durable background work to a queue consumed by workers;
- use optimized native or external tools when appropriate.
Worker threads run inside the same process and communicate through messages or carefully managed shared memory. Separate processes have their own heaps and stronger isolation. I would choose between them based on how much data has to move, how failures should be contained, and how much parallel work the machine can actually support.
Multiple Node processes use multiple cores
One event-loop thread cannot execute JavaScript on all CPU cores. Production deployments normally run multiple Node processes or containers and distribute connections across them.
load balancer
├── Node process 1 → event loop 1
├── Node process 2 → event loop 2
├── Node process 3 → event loop 3
└── Node process 4 → event loop 4
This gives the application more parallel capacity and isolates failures between processes. It also multiplies everything each process owns. If 20 application instances each open 20 database connections, the database may have to manage 400 connections.
Adding application capacity can therefore move the bottleneck straight to the database.
Throughput is limited by the slowest constrained resource
An event loop can dispatch many I/O operations, but the system cannot complete unlimited work. Requests eventually contend for something:
- database connections;
- CPU time;
- memory;
- external API quotas;
- socket or file descriptors;
- downstream concurrency;
- network bandwidth.
A queue can hide the limit for a while. Once requests arrive faster than the system completes them, the queue keeps growing and latency rises. Increasing concurrency at that point may make things worse because more work now competes for the same saturated dependency.
There is no single useful answer to “how many concurrent requests can Node handle?” A server returning a cached 20-byte response has a very different limit from one hashing passwords, generating PDFs, and running five database queries for every request.
Measure event-loop health
Useful signals include:
- event-loop delay or utilization;
- CPU saturation;
- garbage-collection pauses;
- active and queued work in connection pools;
- libuv pool pressure for relevant operations;
- request latency percentiles;
- request rate and error rate.
If CPU usage and event-loop delay rise together, JavaScript or runtime work may be keeping callbacks from running promptly. If CPU is low while requests are slow, I would look for waiting in a queue, a lock, a connection pool, or another service.
Node works best when each callback does a small, bounded amount of work and leaves the waiting to the operating system or another service.
Our order handler eventually calls the database. The next episode follows that query into PostgreSQL.
