Skip to content

Report blockOn call sites with -Dcats.effect.detectBlockOn - #4689

Open
jacum wants to merge 1 commit into
typelevel:series/3.xfrom
jacum:detect-block-on
Open

jacum wants to merge 1 commit into
typelevel:series/3.xfrom
jacum:detect-block-on

Conversation

@jacum

@jacum jacum commented Sep 27, 2026 •

Copy link
Copy Markdown

This small enhancement proven to be useful for high-level optimisation of cats effect work stealing pool:

It reports all places where an io-compute thread enters blocking state triggering io-blocker threads to spawn. This can be some syncrhonized {} block, directly or via many layers of dependencies.

Generally deemed safe for production, as multiple blocking attempts of the same stack won't be reported more than once.

Await.result or blocking { } - e.g. http4s blaze:

[WARNING] A Cats Effect worker thread was blocked via `BlockContext.blockOn`
  at scala.concurrent.Await$.result(package.scala:125)
  at cats.effect.std.DispatcherPlatform.unsafeRunTimed(DispatcherPlatform.scala:61)
  at cats.effect.std.DispatcherPlatform.unsafeRunTimed$(DispatcherPlatform.scala:59)
  at cats.effect.std.Dispatcher$$anon$2.unsafeRunTimed(Dispatcher.scala:287)
  at cats.effect.std.DispatcherPlatform.unsafeRunSync(DispatcherPlatform.scala:52)
  at cats.effect.std.DispatcherPlatform.unsafeRunSync$(DispatcherPlatform.scala:51)
  at cats.effect.std.Dispatcher$$anon$2.unsafeRunSync(Dispatcher.scala:287)
  at org.http4s.blaze.server.WebSocketSupport.$anonfun$renderResponse$7(WebSocketSupport.scala:103)
  at scala.concurrent.impl.Promise$Transformation.run(Promise.scala:517)
This is very likely to be due to `scala.concurrent.blocking` or `Await.result`
in `IO.delay` or `IO.apply`. If this is the case then you should use
`IO.blocking` or `IO.interruptible` instead.

and indeed here it is -
https://github.com/http4s/blaze/blob/2ae13a74d55209b6573d5228d1aa94f0361a75d0/blaze-server/src/main/scala/org/http4s/blaze/server/WebSocketSupport.scala#L103-L104

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@djspiewak

Copy link
Copy Markdown
Member

Does the existing blocking detection not catch these cases?

@jacum

jacum commented Sep 28, 2026 •

Copy link
Copy Markdown
Author

Not exactly these ones - not Await.result.
Please disregard my first example regarding UUID was incorrect - this was indeed caught by existing detection.

I was investigating the spikes of io-blocker threads in particular scenarios (like massive client reconnect in http4s/blaze websockets),

The existing blocking detector correctly showed UUID blocks

[WARNING] A Cats Effect worker thread was detected to be in a blocked state (BLOCKED)
  at java.base/sun.security.provider.SecureRandom.engineNextBytes(SecureRandom.java:222)
  at java.base/sun.security.provider.NativePRNG$RandomIO.implNextBytes(NativePRNG.java:530)
  at java.base/sun.security.provider.NativePRNG.engineNextBytes(NativePRNG.java:217)
  at java.base/java.security.SecureRandom.nextBytes(SecureRandom.java:776)
  at java.base/java.util.UUID.randomUUID(UUID.java:151)

However, this detector is probabilistic, it samples a random sibling from the pool's worker array and checks its Thread.State. Anything that goes through BlockContext.blockOn (scala.concurrent.blocking, Await.result, Future-based library code) calls prepareForBlocking() first. It hands the thread's slot to a clone via replaceWorker and renames it to an "io-blocker". By the time it actually blocks, it is no longer in the worker array, so the sampler seem to never pick it.

The code above from http4s/blaze is called only once in websocket lifetime, because there are massive bursts, it's pronounced (not enough io-blocker threads and new got spawned), but the unsafeRunSync thunk itself is a tiny and never got caught by sampler.


TLDR:
The two blocking detectors are complementary: the current sampler catches raw blocking that bypasses the BlockContext (synchronized, file IO, JDBC), probabilistically and only while the lock is long enough to be sampled.

My change catches BlockContext entrances deterministically, each call site once, at the cost of one static boolean check
when off. The blockOn path is invisible today precisely because it, kind of, works: nothing hangs, but every entry rotates a worker out and spins up an io-blocker. With cpu restricted cloud instances you start feeling that impact as latency degradation.

@durban

durban commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

What's not entirely clear to me: why warn about these? The WSTP intentionally works with blockOn. This is, in effect, a supported thing...


It reports all places where an io-compute thread enters blocking state triggering io-blocker threads to spawn.

I'm not sure this is correct: I think it doesn't report, e.g., IO.blocking (and that seems intentional).

@jacum

jacum commented Sep 29, 2026 •

Copy link
Copy Markdown
Author

@durban

It doesn't report, e.g., IO.blocking (and that seems intentional).

Indeed it does not, its concerns are only eventual blocking contexts on the main io-compute workers.

WSTP performs excellently and close to theoretically possible limits, if there are no blocking contexts at all, just running the fibers. blockOn is a mitigation of blocking that still occurs eventually - as it often happens when the codebase is large, and Cats Effect code uses some libraries where such blocking occurs, and when the codebase is hybrid i.e. not everything is pure F[_] but some Java Futures, Akka/Pekko etc. parts are there too.

Spawning the io-blocker threads and keeping a few of them - is not free and does create some overhead, but may be sufficient for the cases when the latency SLO is not too tight, also in cases when there are enough CPU cores to run these. The impact becomes visible when the blockOn spots are hit in bursts, so there are no available 'io-blocker' threads, and new ones keep spawning.

A particular problem for which this tracking was added - a http4s/blaze websocket/edge service implementation, on a cloud pod sharing cores with other processes, getting hit by up to few thousands of the (reconnecting) clients in the same second. Up to a hundred io-blocker threads were seen spawning, resulting in increased reconnection latency.

The problem spot could be found by (agentic) static code analysis, but with blockOn detector it showed it immediately.
Another (rather efficient) mitigation was dropping WSTP together and running compute pool directly on JDK 25 Loom virtual threads. This doesn't spawn the extra threads but also doesn't benefit from WSTP optimisations like fiber inlining.

@durban

durban commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Thanks for the explanation. What's still not clear to me: if the goal is avoiding spawning extra threads (which is indeed a good goal), then why recommend "IO.blocking or IO.interruptible" in the warning text? Those also start extra threads.


Out of curiosity, what was the solution for the issue you've encountered with http4s? Just using Loom, or did you do something else eventually?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants