Skip to content

benchmarking/locust: capture cluster hardware facts and density frontiers. - #1603

Closed
Nishanth Kotla (Nishanth29) wants to merge 5 commits into
agent-substrate:mainfrom
Nishanth29:benchmarking/locust-telemetry
Closed

Nishanth Kotla (Nishanth29) wants to merge 5 commits into
agent-substrate:mainfrom
Nishanth29:benchmarking/locust-telemetry

Conversation

@Nishanth29

@Nishanth29 Nishanth Kotla (Nishanth29) commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #1590

What this PR does

In alignment with the Actor Density Benchmark Specs, this PR teaches the Locust runner to discover cluster capacity, record actor density frontiers, and harvest server-side Prometheus metrics into stats.jsonl and server_summary.json. Discovery lives in its own module, benchmarking/locust/cluster_facts.py, rather than inside runner.py.

Proposed Changes

Cluster hardware discovery (cluster_facts.py, locust.yaml)

  • Reads allocatable cores, RAM, node count, machine type and worker pod count through the official Kubernetes Python client, so in-cluster and local auth both work without a kubectl subprocess.
  • Adds a ClusterRole with list on nodes and pods. --no-cluster-facts skips discovery, since listing is expensive on a large cluster.
  • Anything unreadable stays null. Nothing is guessed or defaulted, so a real reading is always distinguishable from a missing one.

Density frontiers (trial_summary in stats.jsonl)

  • actors_per_node, actors_per_vcpu, actors_per_gb_ram.
  • Actors per pod as ap_ratio_p50 / p90 / p99 over the steady-state samples rather than one average, so it is visible whether the system actually pegs at 1.
  • Raw readings persist beside the derived ones in raw_configuration, so ratios can be re-derived after the fact.

Server ground truth (server_telemetry.py, server_summary.json)

  • Physical worker assignment from ate_workerpool_workers, as packing percentiles plus the underlying timeseries.
  • Host Linux kernel PSI stalls for CPU, memory and IO, and CFS throttling.
  • Snapshot sizes as P50, P90, P95 and mean, checkpoint counts, restore and checkpoint latencies, and throughput. Counts, mean and throughput cover the steady-state window; the quantiles are a 5m rate at window end.
  • server_telemetry.py uses only the standard library, every call with a timeout. An unreachable Prometheus leaves nulls rather than failing the run, and --prometheus-url retargets it. status.json keeps its existing minimal schema.
  • Depends on benchmarking: expand cAdvisor scrape allowlist for kernel PSI, networ… #1599 and benchmarking: expose raw atelet and ate metrics via Prometheus export… #1601 landing first: the PSI queries need the container="node" relabel and the snapshot queries need the metrics/raw pipeline. Merged ahead of those, the PSI and snapshot fields come back null.

Docs

  • benchmarking/README.md covers the new flags and every field in the output files.

How this was tested

  • 15 unit tests across test_cluster_facts.py and test_server_telemetry.py, covering percentile edges, steady-state detection, counter resets across an atelet restart, and the null-versus-zero rules.
  • Multi-user runs on a benchmark cluster. Frontier fields land in stats.jsonl, server_summary.json is populated, and status.json is unchanged.
  • Emitted values were re-derived by hand from raw Prometheus and matched.

References

Agent Substrate: Actor Density Benchmark Specs

  • Tests pass
  • Appropriate changes to documentation are included in the PR

@maxsmythe Max Smythe (maxsmythe) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this! Left a few initial comments.

Comment thread benchmarking/locust/runner.py Outdated
Comment thread benchmarking/locust/runner.py Outdated
Comment thread benchmarking/locust/runner.py Outdated
Comment thread benchmarking/locust/runner.py Outdated
Comment thread benchmarking/locust/runner.py Outdated
Comment thread benchmarking/locust/runner.py Outdated
@Nishanth29

Nishanth Kotla (Nishanth29) commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks Max, good catches. All six are fixed: reverted the boomer switching,
moved discovery into cluster_facts.py, dropped the env overrides and the
kubectl fallback, switched to the K8s client, and added --cluster-facts /
--no-cluster-facts so the listing can be skipped on big clusters.

Also fixed some telemetry bugs, rewrote the tests and documented the flags and output fields in the README.

@Nishanth29 Nishanth Kotla (Nishanth29) changed the title benchmarking/locust: capture cluster hardware facts and density frontiers in runner benchmarking/locust: capture cluster hardware facts and density frontiers. Sep 14, 2026

@maxsmythe Max Smythe (maxsmythe) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some minor questions about structure. Biggest questions are around some odd things that look like anticipated edge cases (such as worker pods in an unexpected namespace), and a question about how custom load shapes interact with the idea of a steady state.

Comment thread benchmarking/automation/manifests/runner-job.yaml.tmpl Outdated
Comment thread benchmarking/locust/unit_tests/test_cluster_facts.py Outdated
Comment thread benchmarking/locust/unit_tests/test_cluster_facts.py Outdated
Comment thread benchmarking/locust/cluster_facts.py Outdated
Comment thread benchmarking/locust/cluster_facts.py Outdated
Comment thread benchmarking/locust/cluster_facts.py Outdated
Comment thread benchmarking/locust/cluster_facts.py Outdated
Comment thread benchmarking/locust/server_telemetry.py Outdated
@bowei Bowei Du (bowei) added kind/feature An enhancement / feature request or implementation area/benchmarking labels Sep 17, 2026
@Nishanth29

Copy link
Copy Markdown
Contributor Author

superseded by #1723 and #1725, which split this into the kubernetes side and the prometheus side. closing this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/benchmarking kind/feature An enhancement / feature request or implementation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[benchmarking] Capture cluster hardware density frontiers and Prometheus server telemetry in Locust runner

3 participants