Services, releases and recovery
Applies to: Kestowv 0.5.5 enterprise pilot package
Updated
A Kestowv pilot keeps the intended service configuration, observed health and application release visible to the operator. A successful change requires the intended release to be active and the desired services to satisfy their health checks.
Prepare the change#
Record the application version, service commands, resource requirements, health checks and expected user-visible result. Confirm who may stop or restart the affected services and when the change may occur. Keep the previous known-good application package available.
jruby bin/pilotctl validate /path/to/pilot.json
jruby bin/pilotctl doctor /path/to/pilot.json
jruby bin/pilotctl status /path/to/pilot.json
Resolve validation and readiness errors before deploying. On a first deployment, status may not exist yet. On an existing pilot, preserve the current status and verify that the manifest and state directory identify the intended environment.
Deploy and verify#
jruby bin/pilotctl deploy /path/to/pilot.json
Read the activation result and the command's exit status. A failed health check can return the environment to its previous release; do not report the new release as active solely because staging began.
Start or reprovision the long-running pilot using the service instructions supplied for the environment. pilotctl deploy alone is not a substitute for operating that supervisor. Then read status again and verify the expected application behavior with an actual request or other agreed check.
Keep the release health check distinct from ongoing service health. The former determines whether a change is acceptable to activate. The latter monitors a service during operation. Both should reflect useful application behavior, and both should have bounded timeouts.
Respond to a service failure#
Inspect the service's desired state, current state, health result and recent errors. A stopped service with autostart: false can be intentional. A running process can still fail its health check.
The configured restart policy determines whether another attempt is permitted. never leaves recovery to the operator; on_failure retries failures; always permits restart after an exit. Restart budgets and retry delays prevent unlimited rapid attempts. If the budget is exhausted, investigate the underlying error before restarting the service manually. Raising the budget does not correct a missing executable, invalid credential or application defect.
Roll back an application change#
Confirm that a previous release is available and that restoring it is compatible with the application's current data. Then, during the authorized recovery window:
jruby bin/pilotctl rollback /path/to/pilot.json
Restart or reprovision the running pilot using the environment's service runbook. This is a required part of the standalone rollback procedure: changing the selected release does not by itself establish that running services now use it.
After recovery, confirm the active version, desired services, health checks and application response. An application release rollback does not automatically roll back database migrations or recover user data; those belong to the application's recovery procedure.
Retain an operational record#
Keep the manifest, readiness output, application release identity, before/after status, relevant metrics and observed application result. If the change failed, preserve the failure output and the recovery outcome together. This record should let the next operator tell what changed and whether the service recovered without reconstructing the event from memory.