The service starts. Its process appears in the list. The first job still fails because the application hasn't finished opening its database connection. Restarting it again sends the same job into the same gap.
That's an illustrative failure, but it's a useful question for any operator: what does your system actually know when it says a service is up? A process existing, an application accepting work and a customer's job completing are three different observations. Your startup rules and recovery plan need to say which one matters at each step.
Find out what started means on your host
The systemd service manual makes the distinction explicit. Type=simple considers startup complete after the process is forked. Type=exec waits until the service executable has been invoked successfully. Neither event, by itself, says the application has finished its own initialization. With Type=notify, the application sends READY=1 to report that it's ready; startup of ordered follow-up units waits for that notification.
The notification has to mean something in the application. Changing the unit type won't teach an unsupported program to send it, and sending it too early won't make a database connection usable. Check the service's implementation and the documentation for the installed systemd version before changing production configuration. The linked manual source is the project's current development branch, not a guarantee about every distribution.
An ordering rule isn't a health check
systemd documents ordering and requirement dependencies separately. After= controls order when the units are being started; it doesn't, on its own, start the other unit. Wants= and Requires= express dependency relationships, but don't establish startup order by themselves. Ordering also uses the relevant unit's definition of startup completion.
For a review, draw just the path of the failing job. Where does it enter? Which service receives it? What has to be usable before that service can respond? Put the actual readiness observation beside each dependency. If the only answer is 'we wait a few seconds,' record that as a timing assumption to examine, not a measured condition.
Decide whether to wait, stop sending work or restart
Kubernetes separates these decisions with startup, readiness and liveness probes. A startup probe gives initialization its own check and delays the other probes until it succeeds. Once the configured failure threshold is reached, readiness failure marks the Pod not ready without stopping its container; liveness failure can trigger a restart under the applicable policy. The project's documentation warns that poorly designed liveness checks can cause cascading failures, including restarts under load that put more pressure on the remaining Pods.
That distinction is useful even outside Kubernetes. In the review, ask whether restarting the process can change the failed condition. A temporary dependency outage and a deadlocked application may need different responses. If the response is automatic, write down its limit and the condition that hands the problem to a person. Don't let one generic health label silently choose every action.
Write one service's acceptance record
You don't need to replace the monitoring platform to make this clearer. Start with one important service and one action it must perform. Our suggested review record has five fields. Keep the answers close to the configuration and runbook the next operator will actually use.
- Ready for what: name the operation the service must accept, including the dependency it needs for that operation.
- Observation: identify the check, log event or application response that demonstrates readiness, and who maintains it.
- While waiting: describe where new work goes, how long it can wait and how the caller learns that it hasn't been accepted.
- Recovery boundary: specify the permitted response, retry limit, stopping condition and escalation owner.
- Completion evidence: record how the operator confirms that the original job finished, rather than merely seeing the process return.
Rehearse the gap before automating the response
In an authorized test environment, walk through a slow startup and an unavailable dependency as separate scenarios. Agree on the test, stopping conditions and rollback owner first. Watch when the process appears, when the service declares readiness and when the selected operation succeeds. Use test data and a harmless operation; an acceptance check shouldn't place a real order or repeat a customer's payment.
Then ask the receiving operator to explain the result from the record alone. If they need the person who built the service to interpret every step, the handoff still has work in it. If only a tabletop exercise is possible, use it to expose missing decisions and leave the untested behavior clearly marked.
Start with the evidence you already have
For one host and up to three related services, Valen Systems' Reliability Baseline Review examines supplied logs and configuration summaries, health signals, recovery dependencies and response limits. You get written findings and a proposed recovery-test plan, with a review call to work through the next decision.
The review doesn't change production or run a live incident exercise. If your question spans a larger data-center environment or needs hands-on implementation, talk with us about the scope first.
Sources
- systemd.service manual source: service startup types (development branch)Publication date not stated. Retrieved September 11, 2026.
- systemd.unit manual source: ordering and requirement dependencies (development branch)Publication date not stated. Retrieved September 11, 2026.
- Kubernetes: Liveness, Readiness, and Startup ProbesPublication date not stated. Retrieved September 11, 2026.
Put this to work.
Review one host's reliability baseline
Have a different situation? Talk with Valen Systems before choosing a scope.