2026-07-29
Stage a systemd 261 OOM ruleset for a restartable service
Attach one named pressure and swap rule to a restartable service, observe it in dry-run mode, then gate live userspace OOM enforcement.
A named OOM rule is safe to enable only when the monitored cgroup is an intentional kill boundary and the service has documented restart-recovery evidence. This article attaches one rule to one restartable service, evaluates per-cgroup PSI and system-wide swap together, observes the selected cgroup in dry-run mode, and requires an explicit gate before live kills.
systemd-oomd is an earlier userspace action, not a guarantee that it always beats the kernel OOM killer. Abrupt pressure, missing swap, a stalled daemon, or poor thresholds can still let the kernel act first.
Evidence boundary
Source reviewed on 2026-07-29 against the systemd v261 tag at commit de9dbc37ad4aa637e200ac02a0545095997055df. The retained upstream TEST-55-OOMD.sh is source evidence for systemd's own test coverage. It was not run for this article.
The authoring host reports systemd 259 (259.5-0ubuntu3). The authoring host therefore cannot execute or validate the v261 .oomrule and OOMRules= enforcement path. Every v261 command, configuration block, log line, and acceptance check here is an operator template unless a future v261 canary run is recorded and this disclosure is updated with exact host, kernel, systemd build, workload, and observations.
The pinned v261 ruleset manual, OOMRules= reference, implementation, and retained test source describe the behavior used here. Source-reviewed behavior is not locally executed evidence.
Make one service the kill boundary
The example uses Action=kill-all for one restartable-worker.service. In v261, kill-all sends SIGKILL to every process in the monitored unit cgroup hierarchy, including descendants. That action fits only when the complete service cgroup is disposable and restartable.
Do not put unrelated processes, irreplaceable in-memory state, an untested singleton, or a process whose abrupt loss can corrupt shared state in that cgroup. Restart=on-failure restarts a process terminated by SIGKILL unless another unit setting changes that classification or blocks restart, as documented in the v261 systemd.service manual. Existing start limits, readiness, idempotency, durable state, and dependency recovery still need service-specific validation.
kill-all is the primary example because it marks the monitored unit cgroup and its descendants for killing, matching the reviewed service boundary. kill-by-pgscan and kill-by-swap are not alternate rollout paths here. For either action, the monitored cgroup is a candidate when it has no subgroups. Otherwise, v261 walks its descendants and ranks eligible leaf cgroups or descendants with memory.oom.group=1: by recent page-scan rate for kill-by-pgscan, and by swap usage for kill-by-swap. The pgscan action also requires recent reclaim activity. The swap action requires configured swap and the selected cgroup to use more than 5 percent of total system swap. This broader selection boundary needs a separately reviewed cgroup tree.
Verify the host before installing a rule
Run this read-only preflight on the intended canary:
systemctl --version
test "$(stat -fc %T /sys/fs/cgroup)" = cgroup2fs
test -r /sys/fs/cgroup/cgroup.controllers
test -r /proc/pressure/memory
systemctl show --property=DefaultMemoryAccounting --value
swapon --show
systemctl status systemd-oomd.serviceThe version output must show systemd 261 for .oomrule and OOMRules=. The cgroup tests require the unified cgroup v2 hierarchy. The PSI test requires kernel support and readable memory PSI. This preflight establishes host prerequisites, not rule loading or a safe kill boundary.
Require memory accounting for every monitored unit as part of the service's separately reviewed existing policy. The v261 systemd-oomd manual also strongly recommends swap so the daemon has time to react.
This example uses SwapUsageMax=, so require real swap from swapon --show. Without configured swap, that condition cannot be meaningfully satisfied. Pressure-only action remains possible without swap, but pressure can rise abruptly and needs separate tuning. systemctl status reports the daemon's current loaded and runtime state; it does not establish that the daemon can start successfully or that the new rule is loaded.
Put the canary daemon in dry-run mode
--dry-run changes the whole systemd-oomd daemon. It logs destructive actions instead of killing cgroups, and it is not scoped to one rule or unit. Use it only on a dedicated canary, a disposable lab host, or an approved window where suppressing every existing systemd-oomd kill is acceptable.
Inspect the vendor unit before overriding it:
systemctl cat systemd-oomd.serviceCopy the complete installed vendor ExecStart= command, including every argument, then append --dry-run. Clearing ExecStart= first replaces the vendor command rather than adding a second command. The following is an upstream v261 example only if systemctl cat shows the installed vendor command is exactly /usr/lib/systemd/systemd-oomd with no arguments. A distribution command with a different path or arguments must be copied in full before --dry-run is appended.
Save the resulting override as /etc/systemd/system/systemd-oomd.service.d/10-dry-run.conf:
[Service]
ExecStart=
ExecStart=/usr/lib/systemd/systemd-oomd --dry-run/usr/lib/systemd/systemd-oomd is the executable path shown in the v261 manual. The v261 dry-run documentation states that --dry-run logs a triggered kill instead of killing the cgroup.
Activate the override and confirm the reported daemon mode:
sudo systemctl daemon-reload
sudo systemctl restart systemd-oomd.service
oomctl | grep -Fx 'Dry Run: yes'The final command confirms that the daemon reports dry-run mode, not a rule subscription or workload outcome. Stop if Dry Run: yes is absent, and do not attach the canary workload until it is present.
Define one named rule and attach it
The following percentages and time value are illustrative syntax, not safe defaults or measured recommendations. Derive thresholds from the service's normal and failure windows, then repeat dry-run observation.
Save /etc/systemd/oomd/rules.d/restartable-pressure.oomrule:
[Rule]
MemoryPressureAbove=60%
SwapUsageMax=80%
Action=kill-all
LastingSec=30sThe filename supplies the name restartable-pressure; the unit property omits the .oomrule extension. OOMRules= is plural and accepts a space-separated list even though this article attaches one name. NEWS says singular OOMRule=, but the v261 manual, parser, D-Bus property, implementation, and tests use plural OOMRules=.
At least one condition and Action= are mandatory. A 100 percent value for either threshold causes the entire ruleset to be ignored with a warning because that condition can never be exceeded. Zero percent is accepted but is usually not useful. Do not copy the upstream zero-percent stress fixture into a production example. The v261 ruleset validation is the source for those limits.
Save this drop-in as /etc/systemd/system/restartable-worker.service.d/40-oomrule.conf:
[Service]
OOMRules=restartable-pressureThis drop-in intentionally contains only OOMRules=. Memory accounting and restart behavior are separately reviewed existing service policy. Before applying it, review the effective MemoryAccounting=, Restart=, RestartPreventExitStatus=, SuccessExitStatus=, start limits, state ownership, and readiness. Do not add or alter that policy merely to attach a rule.
Apply the rule and inspect the installed unit state:
sudo systemctl reload systemd-oomd.service
sudo systemctl daemon-reload
sudo systemctl restart restartable-worker.service
systemctl show restartable-worker.service \
--property=OOMRules \
--property=MemoryAccounting \
--property=Restart \
--property=RestartPreventExitStatus \
--property=SuccessExitStatus \
--property=ControlGroupReloading systemd-oomd reparses .oomrule files and resets in-progress LastingSec= timers. Restarting the example service gives a clear configuration boundary for the planned canary. systemctl show reports the manager's unit properties and cgroup path; use the systemd-oomd journal to check for undefined-ruleset or ignored-configuration warnings before continuing.
Read the AND gate and duration correctly
MemoryPressureAbove= applies to PSI memory full avg10 for the monitored service cgroup, while SwapUsageMax= applies to used swap across the whole system. LastingSec= requires every configured condition to remain true continuously for that ruleset and monitored cgroup. When both thresholds are configured, both must be true at the same time. If either falls to or below its threshold, the timer resets. LastingSec=30s is continuous time, not thirty samples or accumulated intermittent pressure. The v261 evaluation path runs on a periodic loop, so the action is not a real-time deadline.
The primary example chooses kill-all. The two selective actions have the candidate and eligibility boundary described above; kill-by-swap additionally rejects candidates at or below 5 percent of total system swap. The example values do not define an accepted production policy.
Prove the dry run on the intended cgroup
Use a bounded, disposable canary replay that represents the service's known overload mode and cannot affect unrelated workloads. Do not use a generic memory bomb or a copy-paste exhaustion command.
Observe the daemon, subscription, cgroup PSI, swap context, and journal during the bounded window:
oomctl | sed -n '1,8p'
systemctl show restartable-worker.service \
--property=OOMRules \
--property=MemoryAccounting \
--property=ControlGroup
cgroup_path="$(systemctl show restartable-worker.service --property=ControlGroup --value)"
cat "/sys/fs/cgroup${cgroup_path}/memory.pressure"
awk '/^Swap(Total|Free):/ { print }' /proc/meminfo
sudo journalctl -u systemd-oomd.service -b --since '10 minutes ago'oomctl reports daemon state and system context. systemctl show reports the subscribed name, effective accounting setting, and target path. The pressure and swap reads expose inputs at the time of observation, and the journal records parser messages and proposed action. They do not create overload or establish restart recovery.
Treat a dry-run trigger as useful only when the following evidence is recorded:
oomctlstill saysDry Run: yes.- The unit reports exactly
OOMRules=restartable-pressure, memory accounting enabled, and the expected cgroup. - The journal has no parse, ignored-rule, or undefined-ruleset warning.
- A trigger window logs
Rule 'restartable-pressure' conditions metandoomd dry-run: Would have tried to kill .... - The path in the dry-run log is the service cgroup reviewed as disposable.
- The service remains alive because dry-run does not send the kill.
- Non-overload windows do not repeatedly trigger the rule.
- Record the systemd build, kernel, swap size and policy, rule content, cgroup path, workload window, and journal interval.
In v261, oomctl documents and dumps daemon state, but the dump implementation does not list the named ruleset registry or its monitored rules cgroups. Use it to confirm Dry Run: and system context. Use systemctl show OOMRules plus the systemd-oomd journal to establish the named-rule path.
Gate live enforcement on restart recovery
Do not restore live daemon behavior until all of these are true:
- The exact service cgroup is the intended kill boundary.
- The service has no unrelated descendants.
- Abrupt loss and restart have been tested through the service's own failure-injection procedure.
- Durable state, idempotency, readiness, dependencies, and backlog recovery pass application-specific checks.
- Start limits and alerting prevent an invisible restart loop.
- Dry-run triggers align with intended overload windows and remain absent during accepted normal windows.
- A rollback owner and maintenance window are named.
- An operator explicitly approves restoring live daemon behavior.
This recoverable change disables only the dry-run drop-in created here:
sudo mv \
/etc/systemd/system/systemd-oomd.service.d/10-dry-run.conf \
/etc/systemd/system/systemd-oomd.service.d/10-dry-run.conf.disabled
sudo systemctl daemon-reload
sudo systemctl restart systemd-oomd.service
oomctl | grep -Fx 'Dry Run: no'Restarting the daemon changes enforcement for every policy it manages. Stop if Dry Run: no is absent. Run the first live exercise only on the accepted canary and capture the systemd-oomd journal, unit restart evidence, application readiness, and cgroup path. No such exercise was performed for this article.
Roll back by detaching first
Reduce exposure before changing the shared rule registry:
- Clear
OOMRules=in the service drop-in or disable only40-oomrule.conf. - Run
systemctl daemon-reload. - Restart the canary service in its approved window.
- Verify that
OOMRules=is empty and that the effective memory-accounting and restart properties still match the separately reviewed service policy. - Rename
restartable-pressure.oomruleto a suffix that is not.oomrule. - Reload systemd-oomd.
- Move
10-dry-run.conf.disabledback to10-dry-run.confand restart systemd-oomd if the host must return to dry-run mode.
For the example, disable the unit subscription first, then remove the rule from daemon discovery:
sudo mv \
/etc/systemd/system/restartable-worker.service.d/40-oomrule.conf \
/etc/systemd/system/restartable-worker.service.d/40-oomrule.conf.disabled
sudo systemctl daemon-reload
sudo systemctl restart restartable-worker.service
systemctl show restartable-worker.service \
--property=OOMRules \
--property=MemoryAccounting \
--property=Restart \
--property=RestartPreventExitStatus \
--property=SuccessExitStatus
sudo mv \
/etc/systemd/oomd/rules.d/restartable-pressure.oomrule \
/etc/systemd/oomd/rules.d/restartable-pressure.oomrule.disabled
sudo systemctl reload systemd-oomd.serviceDo not use systemctl revert, because it could remove unrelated administrator drop-ins. Do not remove the ruleset first while a unit still references it: v261 keeps the subscription name and warns that the undefined ruleset is ignored.
Acceptance rule
- Keep the daemon in dry-run mode until the named rule, exact cgroup target, quiet-window behavior, overload-window trigger, and restart recovery have evidence from systemd 261.
- Enable live mode only when the complete selected kill boundary is safe to lose and recover.
- If the boundary or recovery evidence is unclear, detach
OOMRules=and stop.