Skip to content

Benchmarks

Flow-Like includes Criterion benchmarks for single-execution latency, concurrent throughput, and allocator behavior. Use them to compare commits or deployment choices on the same machine.

The throughput benchmark:

  • loads the checked-in test app and board from tests/flow;
  • initializes the current node catalog and local object stores;
  • suppresses normal execution logs;
  • warms the runtime before sampling;
  • runs the same board concurrently at configured in-flight levels;
  • reports Criterion latency and element-throughput distributions.

It measures in-process workflow execution. It does not include ingress, external APIs, remote storage, model inference, databases, queueing, network latency, container startup, or an end-user request path.

The previous internal run used:

SettingRecorded value
CPU16-core Apple M-series machine
Memory32 GB
Operating systemmacOS
BuildOptimized benchmark profile with thin LTO
Allocatormimalloc
WorkloadSmall checked-in workflow; two nodes and five pins
Concurrent executionsApproximate throughputApproximate batch latency
12865,000 exec/s2.0 ms
512100,000 exec/s5.1 ms
1,024112,000 exec/s9.1 ms
2,048121,000 exec/s17 ms
4,096123,000 exec/s33 ms
8,192124,000 exec/s66 ms
Concurrent executionsApproximate throughputApproximate batch latency
12860,000 exec/s2.1 ms
512140,000 exec/s3.6 ms
1,024177,000 exec/s5.7 ms
4,096228,000 exec/s18 ms
8,192238,000 exec/s35 ms
32,768241,000 exec/s135 ms
65,536244,000 exec/s269 ms

These values describe this synthetic in-process workload. A workflow that performs HTTP requests, model inference, storage, database work, or substantial serialization will be dominated by those operations.

The historical high-concurrency allocator run recorded approximately:

AllocatorThroughput at 1,024 concurrent executions
mimalloc222,000 exec/s
System allocator179,000 exec/s

Allocator effects vary by platform, allocation pattern, and concurrency. Run both variants on the deployment target rather than assuming the same percentage improvement.

Run commands from the repository root.

Terminal window
FL_WORKER_THREADS=4 \
FL_CONCURRENCY_LIST="128,512,1024,2048,4096,8192" \
FL_MEASURE_SECS=10 \
RUST_LOG=off \
cargo bench -p flow-like-catalog \
--bench throughput_bench \
--features mimalloc \
-- peak_throughput

The first run may spend significant time compiling the optimized benchmark profile. Later runs reuse Cargo artifacts unless relevant code or features changed.

Environment variableDefaultPurpose
FL_BOARD_IDChecked-in benchmark boardBoard loaded from the test app
FL_START_IDChecked-in start nodeEntry node executed by the benchmark
FL_APP_IDChecked-in test appApp directory under the test store
FL_TESTS_DIR../../tests from the packageLocal benchmark object store
FL_WORKER_THREADSLogical CPU countTokio runtime worker threads
FL_MAX_BLOCKING_THREADSWorker threads × 4Tokio blocking-thread ceiling
FL_CONCURRENCY_LISTAutomatic sweepComma-separated in-flight levels
FL_MAX_CONCURRENCYLogical CPU count × 8Maximum automatic sweep value
FL_MEASURE_SECS10Criterion measurement duration per level
FL_MAX_IN_FLIGHTLogical CPU count × 4In-flight tasks in the raw benchmark
RUST_LOGInherited environmentSet to off for low-noise measurements

Record these details with every result:

  1. Commit SHA and whether the worktree is clean.
  2. Rust and Cargo versions.
  3. Operating system, kernel, CPU model, core topology, and memory.
  4. Power mode, virtualization, and container limits.
  5. Allocator and Cargo feature set.
  6. Board, start node, app ID, and any changes to the test fixture.
  7. Worker, blocking-thread, concurrency, and measurement settings.
  8. Criterion estimate and confidence interval, not only the fastest sample.

Compare two commits on the same machine with the same fixture and environment. Alternating baseline and candidate runs helps reveal thermal or background-load drift.

  • Single execution isolates the overhead of one small in-process run.
  • Throughput shows how efficiently the runtime uses concurrency.
  • Batch latency grows with queue depth even while throughput improves.
  • Peak throughput is not the same as a safe production operating point.
  • External-node workflows require end-to-end benchmarks that include their actual dependencies.

Do not compare these numbers directly with another product unless the workload, durability, logging, retries, isolation, hardware, and measurement boundary are equivalent.

  1. Add or update a target under packages/catalog/benches/.
  2. Keep the workload deterministic and check in its fixture.
  3. Document the measurement boundary and environment variables.
  4. Include before-and-after Criterion output for performance changes.
  5. Explain any change in behavior, durability, or correctness that accompanies the performance result.