The Rust Async Runtime Footgun That Blocks Your Entire Tokio Thread Pool

· Medium ·

2 min read Original article ↗

Illya Yalovoy

Understanding spawnblocking vs block_on in production services, and why your Rust service silently dies under load*

Your Rust service passes all tests, handles load beautifully in staging, then goes completely silent in production under sustained traffic. Health checks time out. Metrics stop updating. You restart the pod, it works for 20 minutes, then dies again. The culprit isn’t a memory leak or a panic — it’s a three-line function that reads a file.

The Failure Pattern Nobody Warns You About

Your service does not crash. It does not panic. It does not print an error. It just stops responding to requests, slowly at first, then completely. The health check still passes because it runs on its own thread. Your metrics show increasing p99 latency, then timeouts, then cascading failures upstream. And you have no idea why because nothing in your logs suggests anything is wrong.

This is what blocking-in-async looks like in production. Alice Ryhl, a Tokio maintainer, wrote an entire blog post explaining what “blocking” means in async Rust specifically because this question kept appearing in the community over and over. The problem is not obscure. It is common, and it is silent.

The math is straightforward. Tokio’s default multi-threaded runtime creates one worker thread per CPU core. On a 4-core machine, you get 4 worker threads. If one of your request handlers calls…