Session OOM isolation
An out-of-memory kill of a process running inside Contour — a compiler, a language server — must end that process and nothing else. This page explains why that took code, and how it works.
Why Contour died with its children
Desktops start applications as systemd units (app-org.contourterminal.Contour@….service).
Every process a session spawns lands in that unit's cgroup, beside Contour. Two things then
act on the unit as a whole:
OOMPolicy=stop(systemd'sDefaultOOMPolicy): when the kernel OOM-kills any process in the unit, systemd stops the whole unit — Contour and every tab.- systemd-oomd kills whole cgroups under memory pressure, so a session's cgroup that contains Contour makes Contour a candidate.
What happens now
vtpty::Process::start() forks the child and parks it on a gate — a socket pair — before
exec. The parent hands the child, as a ParkedChild, to the injected ProcessPlacement:
SystemdScopePlacement(Linux,CONTOUR_WITH_SYSTEMD=ON) asks the user's systemd, over sd-bus, for a transientcontour-session-<contour pid>-<child pid>-<n>.scopeinapp.slicewithOOMPolicy=continue, waits for the job, and releases the child. Everything the child forks from then on is born in that scope. systemd before 253 knows neitherPIDFDsnorOOMPolicy=on scopes and refuses a request carrying either, so the first such refusal switches the connection toPIDs=and noOOMPolicy=. Its scopes do not react to an out-of-memory kill at all, which is what we want.NoPlacement(other platforms, Flatpak, builds without systemd) releases at once.
Before parking, the child also sets the oom_score_adj its placement supplies -- Contour's own plus 100 -- so in a global
out-of-memory the kernel picks a session's process before Contour.
Invariants worth not breaking
- Every path releases the child. Release is
ParkedChild's destructor: an exception, a refusal, a timeout or shutdown cannot leave a shell parked. - Release is a byte sent with
MSG_NOSIGNAL. Contour does not ignoreSIGPIPE, and a child killed while parked must not take Contour with it. A byte, not end-of-file, because a sibling forked meanwhile inherits the gate's parent end. - The child waits on the gate before anything else, using async-signal-safe calls only.
placeThenReleasenever blocks and never throws:Process::start()runs on the GUI thread and reports failures as values. One worker thread owns the sd-bus connection, and starts with the first session, with every signal blocked: the daemon waits forSIGTERMwithsigwait().- A child that exited while queued is not asked about. Its pid may name another process by then; the pidfd says whether it is gone.
- Each request listens for its own unit's
JobRemovedonly, through a match filtered on the unit and removed afterwards; nothing queues up on the connection between spawns. - A wedged bus costs one deadline (1 s), not one per tab: the breaker skips requests for 30 s, then lets one trial request decide.
Testing
vtpty_test "[placement]" covers the gate, the breaker and the worker through a fake ScopeBus.
vtpty_session_oom_isolation (CTest label systemd) proves the requirement under a real user
manager: a driver running as a unit with OOMPolicy=stop must end in oom-kill without placement
and in success with it. CI runs it under a lingering user session and fails if it skipped.