Investigating load tester limits
On today's Chaos Day, we investigated a puzzling result of our daily load tests, reported in camunda/camunda#62783: without secondary storage, the REST load test is much slower than the same test with Elasticsearch, while for gRPC it is the other way around, as one would expect.

We run several load tests every day (gRPC and REST, each with and without Elasticsearch as secondary storage) against the current main. The "None" variants run without secondary storage to max out the engine and observe its limits. In the results above, None-gRPC completes about 491 process instances per second (PI/s), while None-REST completes only about 85 PI/s, even though its starter reports starting 435 PI/s. With Elasticsearch, the REST test completes about 170 PI/s and is at least in the same range as gRPC.
We had already seen that the starter, the application that creates process instances, was heavily CPU throttled. Today, we wanted to find out why.
TL;DR; The collapse was caused by the load tester itself, not by Camunda.
- The starter's HTTP client has a pool of 100 connections. At a configured rate of 500 PI/s, the starter scheduled more requests than the pool completed and queued the rest in memory, up to about 85,000 requests.
- The full heap made the JVM spend most of its CPU on garbage collection, which slowed the starter down and caused the CPU throttling we saw. Giving the starter more CPU made it worse: with two or more CPUs, the JVM runs the G1 garbage collector instead of the serial one, and the starter then has enough CPU to keep sending at the full rate until the heap is exhausted. The starter ran out of memory, and its scheduler died silently.
- Limiting the number of in-flight requests keeps the starter stable under the configured load, in an open workload model: when the system cannot keep up, the starter skips sending instead of queueing them. With this prototype, we reach about 375 PI/s out of the configured 500 PI/s. This is now limited by Camunda's CPU: all three Camunda pods run at their limit of 3 CPUs. The remaining difference to the 491 PI/s of None-gRPC is expected: the protocols differ, and there are known REST performance issues, for example camunda/camunda#35067.
- We also added HTTP client metrics (camunda/camunda#64477) that made the backlog visible. The investigation produced eight new issues and five draft fixes for the load tester, the Java client, and the load test configuration, listed at the end.




