Workload evaluations

Applies to: Kestowv 0.5.5 enterprise pilot package

Updated

A workload evaluation answers a specific operating question: whether an application completes correctly, whether a change improves a defined measure, or whether behavior remains acceptable over time. Choose that question and its acceptance criteria before running the pilot.

Choose a workload type#

Kind Use Required inputs
single A bounded correctness, readiness or failure-handling check. Executable and arguments, input data, working directory where needed, and timeout.
ab Compare a baseline and candidate on equivalent work. Both run definitions, matching input, iteration count, warmup and acceptance thresholds.
soak Observe repeated work over a stated duration. Run definition, duration, interval and expected behavior over the whole period.

List the configured workloads before choosing one:

jruby bin/pilotctl workloads /path/to/pilot.json
jruby bin/pilotctl run /path/to/pilot.json example-workload

Replace example-workload with a declared name. Running a workload executes its command and can affect the target environment. Use a workload whose inputs and effects have been reviewed.

Establish correctness first#

Define what the program must produce and how that result will be checked. An exit code of zero may be necessary, but it is not always sufficient. A service that quickly returns an empty response may be fast and still be wrong.

For a comparison, hold the input and required output constant. Record the machine, operating environment, application versions, workload parameters, relevant limits and competing activity. If those conditions change, explain the difference in the result rather than attributing it automatically to the candidate.

Read the result#

Review the workload's success or pass result, errors, timeout count and applicable latency or comparison measures. Read the complete result record when the terminal presents only a bounded sample of output. An omitted portion of displayed output is not evidence that the underlying operation produced nothing else.

For a soak, report the actual duration, repetitions and failures. A short smoke test cannot establish long-term stability. For fault tests, identify the deliberate fault and the expected recovery behavior separately from an unexpected application failure.

Report a comparison precisely#

A useful report says what completed, whether it was correct, how long it took and under which conditions. Include the number of trials and variation when making a performance comparison. Report failed runs alongside successful ones; excluding them can hide the operational cost of a change.

Throughput alone does not establish responsiveness, energy efficiency or reliability. Those questions require their own observations. Likewise, a Rubian or Kestowv component result applies to the tested configuration; a claim about the integrated Valux operating system requires a corresponding integrated evaluation.

Stop an unsuitable evaluation#

If a workload targets the wrong data, exceeds the approved resource scope or produces an unexpected state change, stop it through the environment's operating procedure and retain the result. Correct the manifest or workload before repeating the run. Do not repeatedly retry a failing production-affecting operation merely to obtain a passing result.