bexhoma.clusters module
Kubernetes cluster management for bexhoma experiments.
Provides Kubernetes for managing experiment deployments on Kubernetes,
and AWS for AWS-specific Kubernetes clusters.
Authors: Patrick K. Erdelt Copyright (C) 2020 Patrick K. Erdelt SPDX-License-Identifier: AGPL-3.0-or-later See LICENSE for details.
- class bexhoma.clusters.AWS(clusterconfig='cluster.config', experiments_configfolder='experiments/', yamlfolder='k8s/', context=None, code=None, instance=None, volume=None, docker=None, script=None, queryfile=None, connect=True)
Bases:
KubernetesAWS EKS extension of the Kubernetes cluster manager.
Adds
eksctl-based nodegroup scaling and AWS-specific node label handling.- check_nodegroup(nodegroup_type='', nodegroup_name='', num_nodes_aux_planned=0)
Return whether a nodegroup is at the expected size.
- Parameters:
nodegroup_type –
typelabel value.nodegroup_name – EKS nodegroup name.
num_nodes_aux_planned – Expected node count.
- Returns:
Trueif actual count equalsnum_nodes_aux_planned.
- eksctl(command)
Run an
eksctlcommand and return its stdout.- Parameters:
command – eksctl subcommand string (without the
eksctlprefix).- Returns:
stdout of the eksctl command.
- get_nodegroup_size(nodegroup_type='', nodegroup_name='')
Return the current number of ready nodes in an EKS nodegroup.
- Parameters:
nodegroup_type –
typelabel value.nodegroup_name – EKS nodegroup name.
- Returns:
Number of nodes currently in the nodegroup.
- get_nodes(app='', nodegroup_type='', nodegroup_name='')
Return node objects matching the given selectors.
Overrides
Kubernetes.get_nodes()to use the EKS-specificalpha.eksctl.io/nodegroup-namelabel instead of the genericnamelabel.- Parameters:
app –
applabel value. Defaults toself.appname.nodegroup_type –
typelabel value.nodegroup_name – EKS nodegroup name (
alpha.eksctl.io/nodegroup-name).
- Returns:
List of Kubernetes node objects.
- scale_nodegroup(nodegroup_name, size)
Scale a single EKS nodegroup to the requested number of nodes.
No-ops if the nodegroup is already at the target size.
- Parameters:
nodegroup_name – EKS nodegroup name.
size – Desired number of nodes.
- Returns:
eksctl output, or
Noneif already at the target size.
- scale_nodegroups(nodegroup_names, size=None)
Scale multiple EKS nodegroups.
- Parameters:
nodegroup_names – Dict mapping nodegroup name to default target size.
size – If given, overrides the default size for every nodegroup.
- wait_for_nodegroup(nodegroup_type='', nodegroup_name='', num_nodes_aux_planned=0)
Block until a single EKS nodegroup reaches the target size, polling every 30 s.
- Parameters:
nodegroup_type –
typelabel value.nodegroup_name – EKS nodegroup name.
num_nodes_aux_planned – Desired node count.
- Returns:
Trueonce the nodegroup is at the target size.
- wait_for_nodegroups(nodegroup_names, size=None)
Block until all listed EKS nodegroups reach their target sizes.
- Parameters:
nodegroup_names – Dict mapping nodegroup name to default target size.
size – If given, overrides the default size for every nodegroup.
- bexhoma.clusters.KUBECTL_CP_INTERNAL_RETRIES = 20
Retries passed to
kubectl cp’s own--retriesflag. kubectl’s exec-based tar transport can truncate its stdout stream mid-copy on larger payloads, surfacing aserror: unexpected EOFwith zero retries (the default);--retriesmakes kubectl resume the tar stream at the last byte offset instead of failing outright. See https://github.com/kubernetes/kubectl/issues/1425.
- class bexhoma.clusters.Kubernetes(clusterconfig='cluster.config', experiments_configfolder='experiments/', yamlfolder='k8s/', context=None, code=None, instance=None, volume=None, docker=None, script=None, queryfile=None, connect=True)
Bases:
objectManages bexhoma experiments on a Kubernetes cluster.
Provides Kubernetes API wrappers (pod/job/service/PVC queries and deletions), cluster component lifecycle methods (dashboard, message queue, monitoring, SUT), experiment bookkeeping (code, result folder, experiment list), and pod/job log persistence.
Subclass:
AWS(K8s on AWS with EKS nodegroup scaling).- add_experiment(experiment)
Append an experiment object to this cluster’s experiment list.
- Parameters:
experiment – Experiment object to add.
- add_to_messagequeue(queue, data)
Push
dataonto the tail of a Redis list (message queue).- Parameters:
queue – Redis key (queue name).
data – Value to push onto the queue.
- Returns:
Resulting list length reported by
RPUSH, orNoneif the reply could not be parsed as an integer (e.g. the command failed and returned no numeric reply).- Return type:
int or None
- check_dbms_connection(ip, port)
Test whether a TCP connection to
ip:portcan be established.Used to probe DBMS readiness before starting a benchmark.
- Parameters:
ip – Hostname or IP address to connect to.
port – TCP port number.
- Returns:
Trueif the connection succeeded,Falseotherwise.
- cluster_access()
Initialise Kubernetes API clients using the configured context.
Sets
self.v1core,self.v1apps, andself.v1batches. Prints a warning if the cluster cannot be reached.
- create_object_from_file(filename_source)
Apply a Kubernetes manifest file via
kubectl create.The manifest is copied to the experiment result folder (or, if no experiment
codeis set, directly to the result folder root, e.g. for cluster-wide components like the dashboard) with theBEXHOMA_PACKAGE_VERSIONplaceholder substituted, then applied.- Parameters:
filename_source – Path to the source manifest template file.
- delete_deployment(deployment)
Delete a Kubernetes Deployment by name.
- Parameters:
deployment – Name of the Deployment to delete.
- delete_job(jobname='', app='', component='', experiment='', configuration='', client='')
Delete a Job by name or by matching label selectors.
If
jobnameis empty, the first Job matching the given selectors is deleted.- Parameters:
jobname – Name of the Job to delete (optional).
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.client –
clientlabel value (legacy).
- Returns:
Trueon success, orNoneif the Job was not found.
- delete_job_pods(jobname='', app='', component='', experiment='', configuration='', client='')
Delete all Pods of a Job identified by name or by matching label selectors.
If
jobnameis empty, all Pods matching the given selectors are deleted individually by recursing with their names.- Parameters:
jobname – Pod or Job name (optional).
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.client –
clientlabel value (legacy).
- delete_messagequeue_key(queue)
Delete a Redis key.
Used to clear a message queue before repopulating it, so that any leftover entries from an earlier (e.g. crashed or retried) population attempt cannot desynchronise the 1..N chunk assignment handed out to a fresh batch of pods.
- Parameters:
queue – Redis key to delete.
- delete_pod(name)
Delete a Pod by name. A 404 (already gone) is silently ignored.
- Parameters:
name – Name of the Pod to delete.
- delete_pvc(name)
Delete a PersistentVolumeClaim by name. A 404 (already gone) counts as success.
- Parameters:
name – Name of the PVC to delete.
- Returns:
Trueif deleted (or already gone),Falseon error.
- delete_service(name)
Delete a Service by name.
- Parameters:
name – Name of the Service to delete.
- delete_stateful_set(name)
Delete a StatefulSet by name.
- Parameters:
name – Name of the StatefulSet to delete.
- download_file(filename_remote, filename_local, pod, container='dashboard', max_retries: int = 3) str
Download a file from a Pod container to the local machine using
kubectl cp.On Windows the local destination path is converted to a UNC path first. The command itself is given
--retriesso kubectl resumes a truncated tar stream (seeKUBECTL_CP_INTERNAL_RETRIES) instead of failing outright. On top of that, this method retries the whole command up tomax_retriestimes on transient failures. RaisesRuntimeErrorif all attempts fail.- Parameters:
filename_remote – Source path inside the container.
filename_local – Destination path on the local machine.
pod – Source Pod name.
container – Source container name. Defaults to
dashboard.max_retries – Maximum number of attempts before raising.
- Returns:
Output of the kubectl command.
- Raises:
RuntimeError – If every attempt returns a failure.
- execute_command_in_pod(command, pod='', container='', params='')
Execute a shell command inside a container of a running Pod.
Retries automatically on transient
error dialing backendfailures.- Parameters:
command – Shell command string.
pod – Name of the target Pod.
container – Container name within the Pod (optional for single-container pods).
params – Unused; reserved for future use.
- Returns:
Tuple
("", stdout_str, stderr_str).
- forward_dashboard_ports()
Forward the dashboard Pod’s ports (8050 and 8888) to localhost.
Port 8050 is the DBMSBenchmarker result dashboard; port 8888 is Jupyter.
- forward_sut_port(experiment='', configuration='')
Forward the SUT master service port to localhost.
- Parameters:
experiment – Filter by experiment code (optional).
configuration – Filter by DBMS configuration name (optional).
- get_available_storage_types() list
Return the SUT persistent-volume storage class names valid for the current cluster context.
Noneand''(ephemeral / no PVC) and'ramdisk'(in-memory, handled specially byuse_ramdisk()) are always valid, independent of the cluster. Any other value must be declared in the cluster configuration file undercredentials.k8s.context.<name>.storage_classes, since actual Kubernetes StorageClass names differ from cluster to cluster.- Returns:
Valid values for
--request-storage-type.- Return type:
- get_dashboard_name(app='', component='dashboard')
Build a canonical name string for the dashboard component.
- Parameters:
app – App name. Defaults to
self.appname.component – Component label. Defaults to
dashboard.
- Returns:
Name string in the format
<app>_<component>.
- get_dashboard_pod_name(app='', component='dashboard')
Return the name of the dashboard Pod, or
""if none exists.- Parameters:
app – App name (passed to
get_pods(); unused if empty).component – Component label. Defaults to
dashboard.
- Returns:
Pod name string, or
""if no dashboard Pod was found.
- get_deployments(app='', component='', experiment='', configuration='')
Return names of Deployments matching the given label selectors.
- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value (e.g.sut,monitoring).experiment –
experimentlabel value (experiment code).configuration –
configurationlabel value (DBMS config name).
- Returns:
List of Deployment names.
- get_job_pods(app='', component='', experiment='', configuration='', client='')
Return names of Pods belonging to Jobs matching the given label selectors.
- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.client –
clientlabel value (legacy).
- Returns:
List of Pod names, or
Noneon a 404 error.
- get_job_status(jobname='', app='', component='', experiment='', configuration='', client='')
Return the completion status of a Job.
If
jobnameis empty, the first Job matching the given selectors is used. ReturnsTrueif the completion count has been reached,0if in-progress or on error,"no job"if no matching Job was found.- Parameters:
jobname – Name of the Job (optional).
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.client –
clientlabel value (legacy).
- Returns:
True,0, or"no job".
- get_jobs(app='', component='', experiment='', configuration='', client='')
Return names of Jobs matching the given label selectors.
- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.client –
clientlabel value (legacy, may be unused).
- Returns:
List of Job names, or
Noneon a 404 error.
- get_jobs_labels(app='', component='', experiment='', configuration='', client='')
Return a dict mapping Job name to label dict for Jobs matching the given selectors.
- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.client –
clientlabel value (legacy).
- Returns:
Dict
{job_name: labels_dict}, or[]if no Jobs found.
- get_messagequeue_name(app='', component='messagequeue')
Build a canonical name string for the message-queue component.
- Parameters:
app – App name. Defaults to
self.appname.component – Component label. Defaults to
messagequeue.
- Returns:
Name string in the format
<app>_<component>.
- get_nodes(app='', nodegroup_type='', nodegroup_name='')
Return node objects matching the given label selectors.
- Parameters:
app –
applabel value. Defaults toself.appname.nodegroup_type –
typelabel value (e.g.sut).nodegroup_name –
namelabel value (e.g.sut_high_memory).
- Returns:
List of Kubernetes node objects.
- get_pod_containers(pod)
Return the names of all containers and init-containers in a Pod.
- Parameters:
pod – Name of the Pod.
- Returns:
List of container name strings (regular containers + init containers).
- get_pod_status(pod, app='')
Return the phase of a named Pod.
- Parameters:
pod – Pod name to look up.
app –
applabel value. Defaults toself.appname.
- Returns:
Phase string (e.g.
Running,Succeeded) or""if not found.
- get_pods(app='', component='', experiment='', configuration='', dbms='', status='')
Return names of Pods matching the given label selectors and optional phase.
- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.dbms –
dbmslabel value (DBMS type, e.g.MonetDB).status – Pod phase to filter by (e.g.
Running,Succeeded).
- Returns:
List of Pod names.
- get_pods_labels(app='', component='', experiment='', configuration='')
Return a dict mapping Pod name to label dict for Pods matching the given selectors.
- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.
- Returns:
Dict
{pod_name: labels_dict}.
- get_ports_of_service(app='', component='', experiment='', configuration='')
Return the port numbers exposed by the first Service matching the given selectors.
- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.
- Returns:
List of port number strings from the first matched Service.
- get_pvc(app='', component='', experiment='', configuration='', pvc='')
Return names of PersistentVolumeClaims matching the given selectors.
If
pvcis provided, only the entry with that exact name is returned.- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.pvc – Optional PVC name to filter by.
- Returns:
List of PVC names.
- get_pvc_labels(app='', component='', experiment='', configuration='', pvc='')
Return label dicts of PVCs matching the given selectors.
- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.pvc – Optional PVC name to filter by.
- Returns:
List of label dicts.
- get_pvc_specs(app='', component='', experiment='', configuration='', pvc='')
Return spec objects of PVCs matching the given selectors.
- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.pvc – Optional PVC name to filter by.
- Returns:
List of PVC spec objects.
- get_pvc_status(app='', component='', experiment='', configuration='', pvc='')
Return status objects of PVCs matching the given selectors.
When
pvcis given, returns thestatusof the matching PVC; otherwise returns thespecof all matched PVCs.- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.pvc – Optional PVC name to filter by (returns
statusfor match).
- Returns:
List of PVC status or spec objects.
- get_service_endpoints(service_name='bexhoma-service-monitoring-default')
Return IP addresses of all endpoints for a named Service.
Particularly useful for headless Services (e.g. to enumerate monitoring nodes).
- Parameters:
service_name – Name of the Service to query.
- Returns:
List of endpoint IP strings, or
[]on error.
- get_services(app='', component='', experiment='', configuration='')
Return names of Services matching the given label selectors.
- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.
- Returns:
List of Service names.
- get_stateful_set_pods(stateful_set='')
Return names of Pods belonging to a given StatefulSet.
- Parameters:
stateful_set – Name of the StatefulSet.
- Returns:
List of Pod names.
- get_stateful_sets(app='', component='', experiment='', configuration='')
Return names of StatefulSets matching the given label selectors.
- Parameters:
app –
applabel value. Defaults toself.appname.component –
componentlabel value.experiment –
experimentlabel value.configuration –
configurationlabel value.
- Returns:
List of StatefulSet names.
- is_dashboard_running()
Return whether the dashboard Pod exists and is in the
Runningphase.- Returns:
Trueif the dashboard is running.
- is_messagequeue_running(component='messagequeue')
Return whether the message-queue Pod exists and is in the
Runningphase.- Parameters:
component – Component label to query. Defaults to
messagequeue.- Returns:
Trueif the message queue is running.
- is_monitoring_healthy()
Probe Prometheus by issuing a
query_rangerequest from inside the dashboard pod.Queries
sum(node_memory_MemTotal_bytes)over a 5-minute window ending 4 minutes ago. ReturnsTrueif Prometheus responds with HTTP 200.- Returns:
Trueif Prometheus is reachable and healthy.
- is_pod_ready(pod)
Return whether a Pod’s
Readycondition isTrue.- Parameters:
pod – Name of the Pod to check.
- Returns:
Trueif the Pod is ready,Falseotherwise.
- job_description(jobname)
Return the
kubectl describe joboutput for a given Job.Unlike a Pod’s describe, this captures the Job’s own Events (e.g.
SuccessfulCreate/BackoffLimitExceeded) for every Pod the Job ever spawned – including one that failed and was replaced, even after that Pod object itself has been garbage-collected.- Parameters:
jobname – Name of the Job to describe.
- Returns:
kubectl output string.
- job_description_exists(jobname)
Return whether a cached
describefile exists in the result folder.- Parameters:
jobname – Name of the Job.
- Returns:
Trueif the.describe.job.logfile exists on disk.
- kubectl(command)
Run a
kubectlcommand in the configured context.Decodes output using UTF-8, Latin-1, or CP-1252 (in that order). On an
Unauthorizedresponse the access token is refreshed and the command is retried once.- Parameters:
command – kubectl subcommand string (without
kubectl --context ...prefix).- Returns:
Decoded stdout string, or
Noneon failure.
- list_deployments_full(app: str = '', component: str = '', experiment: str = '', configuration: str = '') list
Return full Deployment objects matching the given label selectors in a single API call.
- list_jobs_full(app: str = '', component: str = '', experiment: str = '', configuration: str = '') list
Return full Job objects matching the given label selectors in a single API call.
Each returned object already carries
spec.completionsandstatus.succeeded/status.conditions, so callers can determine completion without a furtherread_namespaced_job_statuscall per Job.
- list_pods_full(app: str = '', component: str = '', experiment: str = '', configuration: str = '') list
Return full Pod objects matching the given label selectors in a single API call.
- list_pvcs_full(app: str = '', component: str = '', experiment: str = '', configuration: str = '') list
Return full PersistentVolumeClaim objects matching the given label selectors in a single API call.
- list_services_full(app: str = '', component: str = '', experiment: str = '', configuration: str = '') list
Return full Service objects matching the given label selectors in a single API call.
- list_stateful_sets_full(app: str = '', component: str = '', experiment: str = '', configuration: str = '') list
Return full StatefulSet objects matching the given label selectors in a single API call.
- log_experiment(experiment)
Append a step record to the experiment log and persist it to disk.
The record is enriched with cluster-level metadata and appended to
self.experiments. The list is written toexperiments.configin the benchmark result folder.Note
This method should be updated to produce YAML output and to handle detached parallel-loader workflows.
- Parameters:
experiment – Dict describing the current experiment step.
- pod_description(pod, container='')
Return the
kubectl describe podoutput for a given Pod.The
containerparameter is accepted for API compatibility but is not forwarded to kubectl (describe is not container-sensitive).- Parameters:
pod – Name of the Pod to describe.
container – Ignored; kept for API compatibility.
- Returns:
kubectl output string.
- pod_description_exists(pod_name, container='')
Return whether a cached
describefile exists in the result folder.- Parameters:
pod_name – Name of the Pod.
container – Accepted for API compatibility but ignored —
kubectl describeis not container-sensitive.
- Returns:
Trueif the.describe.logfile exists on disk.
- pod_log(pod, container='')
Return the
kubectl logs --tail=-1output for a given Pod or container.- Parameters:
pod – Name of the Pod.
container – Container name within the Pod (optional).
- Returns:
kubectl output string.
- pod_log_exists(pod_name, container='')
Return whether a cached log file exists in the result folder.
- Parameters:
pod_name – Name of the Pod.
container – Container name within the Pod (optional).
- Returns:
Trueif the.logfile exists on disk.
- pvc_exists(name)
Return whether a PVC with the given name exists in the namespace.
- Parameters:
name – Name of the PVC to check.
- Returns:
Trueif the PVC exists,Falseif not found (HTTP 404).
- restart_dashboard(app='', component='dashboard')
Force-restart the dashboard by deleting its Pod (Kubernetes will recreate it).
- Parameters:
app – App name (passed to
get_dashboard_pod_name()).component – Component label. Defaults to
dashboard.
- set_code(code)
Set the unique identifier of the current experiment.
If
codeis notNoneand anexperiments.configfile exists in the result folder, the stored experiment list is loaded from it.- Parameters:
code – Unique experiment identifier string.
- set_connection_management(**kwargs)
Set DBMSBenchmarker connection-management parameters.
Can be overridden per experiment or per DBMS configuration.
- Parameters:
kwargs – Key/value pairs (e.g.
timeout=60,numProcesses=4).
- set_ddl_parameters(**kwargs)
Set DDL template substitution parameters.
Values replace placeholders in DDL scripts executed during loading. Can be overridden per experiment or per DBMS configuration.
- Parameters:
kwargs – Key/value pairs (e.g.
index='btree').
- set_experiment(instance=None, volume=None, docker=None, script=None)
Select a specific instance/volume/docker/script combination for the experiment.
- Parameters:
instance – Instance key within
self.instances(legacy IaaS).volume – Volume key within
self.volumes.docker – Docker image key within
self.dockers.script – Init-script key within the selected volume’s
initscripts.
- set_experiments(instances=None, volumes=None, dockers=None)
Store the top-level experiment catalog from the cluster configuration.
- Parameters:
instances – Dict of IaaS instance specs (legacy, was for IaaS scaling).
volumes – Dict of volume definitions carrying dataset metadata.
dockers – Dict of Docker image descriptors and usage metadata.
- set_experiments_configfolder(experiments_configfolder)
Set the folder that contains experiment configuration sub-folders.
Bexhoma expects sub-folders named by experiment type (e.g.
tpch), each containing aqueries.configfile and per-DBMS DDL schema folders.- Parameters:
experiments_configfolder – Relative path to the experiments folder.
- set_pod_config(key: str, config: dict) None
Store per-pod configuration as a JSON string in Redis.
The JSON object is stored at
keyand can be retrieved by shell scripts running inside pods usingredis-cli get <key>. The value is then expanded intoBEXHOMA_POD_-prefixed environment variables bybexhoma-pod-env.sh.
- set_pod_counter(queue, value=0)
Set a Redis key to an integer value (used as a pod-count synchronisation counter).
- Parameters:
queue – Redis key to set.
value – Integer value to assign. Defaults to
0.
- set_query_management(**kwargs)
Set DBMSBenchmarker query-management parameters.
- Parameters:
kwargs – Key/value pairs forwarded to DBMSBenchmarker (e.g.
numRun=3).
- set_queryfile(queryfile)
Set the query config file for the DBMSBenchmarker benchmarker component.
- Parameters:
queryfile – Path to the DBMSBenchmarker query configuration file.
- set_resources(**kwargs)
Set Kubernetes resource requests/limits for the SUT component.
Can be overridden per experiment or per DBMS configuration.
- Parameters:
kwargs – Key/value pairs (e.g.
requests={'cpu': 4, 'memory': '16Gi'}).
- set_workload(**kwargs)
Set workload metadata for the experiment (e.g.
name,description).- Parameters:
kwargs – Arbitrary key/value pairs stored in
self.workload.
- start_dashboard(app='', component='dashboard')
Start the dashboard Deployment and its Service if not already running.
Manifest template:
deploymenttemplate-bexhoma-dashboard.yml.- Parameters:
app – App name passed to
get_dashboard_name().component – Component label. Defaults to
dashboard.
- start_datadir()
Provision the shared data-source PVC if it does not exist.
The PVC is used by data-generator pods for writing and by loader pods for reading. Manifest template:
pvc-bexhoma-data.yml.
- start_messagequeue(app='', component='messagequeue')
Start the message-queue Deployment if not already running.
Manifest template:
deploymenttemplate-bexhoma-messagequeue.yml.- Parameters:
app – App name passed to
get_messagequeue_name().component – Component label. Defaults to
messagequeue.
- start_monitoring_cluster(app='', component='monitoring')
Start cluster-level monitoring (Prometheus + node exporters) if not healthy.
Waits up to 100 seconds for an existing Prometheus to become healthy before deploying the daemonset. Manifest template:
daemonsettemplate-monitoring.yml.- Parameters:
app – App name (unused; kept for API consistency).
component – Component label. Defaults to
monitoring.
- start_resultdir()
Provision the shared results PVC if it does not exist.
Benchmarker pods write results here; the evaluator pod reads from here. Collected metrics are also stored here. Manifest template:
pvc-bexhoma-results.yml.
- stop_benchmarker(experiment='', configuration='')
Stop all benchmarker Jobs and their Pods in the cluster.
- Parameters:
experiment – Filter by experiment code (optional).
configuration – Filter by DBMS configuration name (optional).
- stop_dashboard(app='', component='dashboard')
Stop the dashboard Deployment and its Service if running.
- Parameters:
app –
applabel value. Defaults toself.appnamevia sub-calls.component –
componentlabel value. Defaults todashboard.
- stop_loading(experiment='', configuration='')
Stop all loading Jobs and their Pods in the cluster.
- Parameters:
experiment – Filter by experiment code (optional).
configuration – Filter by DBMS configuration name (optional).
- stop_messagequeue(app='', component='messagequeue')
Stop the message-queue Deployment and its Service if running.
- Parameters:
app –
applabel value. Defaults toself.appnamevia sub-calls.component –
componentlabel value. Defaults tomessagequeue.
- stop_monitoring(app='', component='monitoring', experiment='', configuration='')
Stop all monitoring Deployments and their Services in the cluster.
- Parameters:
app –
applabel value.component – Component label. Defaults to
monitoring.experiment – Filter by experiment code (optional).
configuration – Filter by DBMS configuration name (optional).
- stop_sut(app='', component='sut', experiment='', configuration='')
Stop all SUT Deployments, StatefulSets, and Services in the cluster.
When
component='sut', worker and pool sub-components are recursively stopped.- Parameters:
app –
applabel value.component – Component label. Defaults to
sut.experiment – Filter by experiment code (optional).
configuration – Filter by DBMS configuration name (optional).
- store_job_description(jobname)
Fetch and persist
kubectl describe joboutput to the result folder.The file is not overwritten if it already exists. Up to 10 retries are attempted in case of transient kubectl failures.
jobnamealready encodesexperiment_run/data_job(seecreate_manifest_job()), so unlikestore_pod_description()no separatenumbersplicing is needed here.- Parameters:
jobname – Name of the Job to describe.
- store_pod_description(pod_name, container='', number=None)
Fetch and persist
kubectl describe podoutput to the result folder.The file is not overwritten if it already exists. Up to 10 retries are attempted in case of transient kubectl failures.
- Parameters:
pod_name – Name of the Pod to describe.
container – Accepted for API compatibility but ignored —
kubectl describeis not container-sensitive.number – Optional experiment-run index, spliced into the filename directly after the experiment code — see
_pod_label().
- store_pod_log(pod_name, container='', number=None)
Fetch and persist
kubectl logsoutput to the result folder.The file is not overwritten if it already exists. Up to 10 retries are attempted in case of transient kubectl failures.
- Parameters:
pod_name – Name of the Pod.
container – Container name within the Pod (optional).
number – Optional experiment-run index, spliced into the filename directly after the experiment code — see
_pod_label().
- upload_file(filename_remote, filename_local, pod, container='dashboard', max_retries: int = 3) str
Upload a local file into a Pod container using
kubectl cp.On Windows the local path is converted to a UNC path first. The command itself is given
--retriesso kubectl resumes a truncated tar stream (seeKUBECTL_CP_INTERNAL_RETRIES) instead of failing outright. On top of that, this method retries the whole command up tomax_retriestimes on transient failures (e.g. context deadline exceeded). RaisesRuntimeErrorif all attempts fail.filename_remotetypically lives on a PVC also mounted by other pods (e.g. parallel benchmarker job pods reading/results/{code}/queries.configfrom their own mount of the same volume).kubectl cpextracts its tar stream directly into the destination path, which is a truncate-then-write from the perspective of any other pod reading that path concurrently. To avoid a concurrent reader observing a partially-written (or empty) file, the upload lands atfilename_remote + '.uploadtmp'first, then anmvexecuted inside the same pod performs an atomic rename into the final path – rename is a single filesystem metadata operation, so any other pod’s mount of the same PVC always sees either the previous complete file or the new complete file, never a truncated one.The rename is verified with a follow-up
test -sinside the pod rather than trusted blindly:execute_command_in_pod()swallows exec failures (e.g. a transient RBAC/permission error onpods/exec) by design, so an uncheckedmvcan silently fail whilerun_pod()goes on to submit the benchmarker Job believing the file is in place – the pods then find nothing at their final path. A failed verification retries the whole upload (not just the rename), since the tmp file may also be gone.- Parameters:
filename_remote – Destination path inside the container.
filename_local – Source path on the local machine.
pod – Target Pod name.
container – Target container name. Defaults to
dashboard.max_retries – Maximum number of attempts before raising.
- Returns:
Output of the kubectl command.
- Raises:
RuntimeError – If every attempt returns a failure.
- wait(sec, silent=False)
Sleep for
secseconds, optionally printing progress messages.- Parameters:
sec – Number of seconds to wait.
silent – If
True, suppress all output.
- bexhoma.clusters.to_unc(path: str) str
Convert a local Windows drive path to a UNC administrative share path.
D:/fooandD:\foobecome\\localhost\D$\foo. On Linux/macOS the normalized path is returned unchanged. A path that is already UNC is returned unchanged.A trailing
/.segment (thekubectl cpmarker meaning “copy this directory’s contents”, not the directory itself, needed to avoid kubectl cp nesting the source directory inside an already-existing destination) is preserved, sincePath()would otherwise silently drop it.