Skip to content

LinuxContainer.create() leaks the VZ VM/XPC process when vm.start() throws — no do/catch around it #804

Description

@davidwkeith

Summary

LinuxContainer.create() (Sources/Containerization/LinuxContainer.swift, pinned 0.35.0) calls the raw vm.start() outside the do/catch block that otherwise cleans up on failure. If vm.start() itself throws — e.g. a vmnet allocation failure surfaced through the network device attachment — the underlying VZVirtualMachine (and its com.apple.Virtualization.VirtualMachine.xpc helper process) is never stopped. Since vm is a local variable inside create(), nothing outside the function ever holds a reference to it, so there is no way for a caller to stop or clean it up after the fact either.

public func create() async throws {
    try await self.state.withLock { state in
        ...
        let vm = try await self.vmm.create(config: creationConfig)
        let relayManager = UnixSocketRelayManager(vm: vm, log: self.logger)

        try await vm.start()   // <-- NOT inside the do/catch below; a throw here leaks `vm`
        do {
            try await vm.withAgent { agent in ... }
            state = .created(.init(vm: vm, relayManager: relayManager, ...))
        } catch {
            try? await relayManager.stopAll()
            try? await vm.stop()
            state.setErrored(error: error)
            throw error
        }
    }
}

Because state is never transitioned away from .initialized when vm.start() throws here, a caller's subsequent container.stop() also can't help — stop() requires .created/.started state (via state.createdState("stop") / state.startedState("stop")), both of which throw against .initialized, so stop() is a no-op in this exact failure mode.

Impact

In our case (Anglesite, a macOS app embedding this package), a transient vmnet NAT-subnet contention (a known single-consumer limitation — see apple/container's vmnet plugin) causes exactly this: vm.start() throws vmnet_return_t(rawValue: 1002), and the resulting leaked VZ process permanently holds the vmnet lease. Every subsequent boot attempt — including retries and full app relaunches — fails identically, because the orphaned com.apple.Virtualization.VirtualMachine.xpc process survives even a full quit/relaunch of the embedding app (it isn't reaped as a child of our process). The only recovery is manually finding and killing that PID.

Suggested fix

Wrap vm.start() itself in a do/catch that calls try? await vm.stop() (and relayManager.stopAll()) on failure, mirroring the existing catch that already covers the post-start setup steps — e.g.:

let vm = try await self.vmm.create(config: creationConfig)
let relayManager = UnixSocketRelayManager(vm: vm, log: self.logger)

do {
    try await vm.start()
    try await vm.withAgent { agent in ... }
    state = .created(.init(vm: vm, relayManager: relayManager, ...))
} catch {
    try? await relayManager.stopAll()
    try? await vm.stop()
    state.setErrored(error: error)
    throw error
}

Environment

  • Package version: 0.35.0 (pinned via .upToNextMinor(from: "0.35.0"))
  • macOS: Apple Silicon, Virtualization.framework via VZVirtualMachineManager

Happy to send a PR with the fix above if useful — wanted to check the intended failure-handling shape first, since I'm not certain whether vm.stop() is safe to call on a VM that never finished start().

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions