feat: Add support for Landlock based Sandbox for Linux (#238)

* feat: Initial implementation of landlock based sandbox driver

* fix: Handle seccom probe failure

* fix: Remove unnecessary seccomp probe

* fix: Use file based policy load

* fix: Keep bpf filter in memory

* fix: Use TSYNC for seccom filter

* fix: Use TSYNC for seccom filter

* fix: Update landlock translator

* fix: Landlock sandbox implementation

* fix: Landlock + seccomp based sandboxing on Linux

* fix: Misc fixes

* fix: Cleanup sandbox files

* fix: Handle mandatory deny API change post merge

* fix: Landlock write access translation

* chore: Fix linter issues

* ci: Use /tmp for npm cache for landlock
This commit is contained in:
Abhisek Datta
2026-05-07 12:42:28 +05:30
committed by GitHub
parent 5122a1594c
commit 4c42ceca0e
27 changed files with 4356 additions and 122 deletions
+141
View File
@@ -0,0 +1,141 @@
# Landlock Sandbox: Developer Notes
How the Linux Landlock driver works and why. User docs: [sandbox.md](./sandbox.md).
## Why Landlock + seccomp
Landlock is positive allow-list. Our profiles are negative on top of broad allow:
`allow_read: /` plus implicit deny on `~/.ssh`, `~/.aws`, `.env`, `.git/hooks`. Landlock
cannot subtract from a subtree, so we layer seccomp-notify on top:
- Landlock: kernel-native allow-list, fast, applies to most syscalls.
- seccomp-notify: intercepts `openat`/`openat2`/`execve`/`execveat`, resolves the path arg
by reading the trapping process's memory, matches against the deny list, responds
`EACCES` or `CONTINUE`.
## Architecture
```
pmg main ──fork+exec──► pmg __landlock_sandbox_exec [helper, unfiltered]
│ runs supervisor loop
clone(CLONE_NEWUSER, uid=0→host)
pmg __landlock_shim [single-threaded,
├ install seccomp (no NNP) uid 0 in ns,
├ apply Landlock CAP_SYS_ADMIN]
├ send notify_fd via SCM_RIGHTS
└ execve target
target ─► child ─► grandchild [filter inherited,
dumpable=1]
```
The helper has no filter on itself, so it can read `/proc/<pid>/mem` for any descendant
to resolve `openat` paths.
### Code layout
| File | Role |
|------|------|
| `cmd/landlock/landlock_sandbox_exec_linux.go` | Helper subcommand wrapper |
| `cmd/landlock/landlock_shim_linux.go` | Shim subcommand wrapper |
| `sandbox/platform/landlock_linux.go` | `Sandbox` impl, command rewrite |
| `sandbox/platform/landlock_translator_linux.go` | PMG policy → `landlockExecPolicy` |
| `sandbox/platform/landlock_helper_linux.go` | Helper: forks shim, runs supervisor |
| `sandbox/platform/landlock_shim_linux.go` | Shim: installs seccomp+Landlock, execve |
| `sandbox/platform/landlock_seccomp_linux.go` | BPF, supervisor loop, deny matchers, memfd cache |
| `sandbox/platform/landlock_abi_linux.go` | Kernel ABI probe |
## Key decisions
### Shim runs in `CLONE_NEWUSER` so seccomp can be installed without NNP
Unprivileged seccomp install requires `PR_SET_NO_NEW_PRIVS`. NNP plus `execve` triggers
`LSM_UNSAFE_NO_NEW_PRIVS` and the kernel sets `dumpable=0`. With `dumpable=0`,
`/proc/<pid>/mem` opens require `CAP_SYS_PTRACE`, which the helper does not have.
Result: supervisor cannot resolve openat paths for descendants.
The shim boots inside a fresh user namespace mapped `0 → host_uid`. As uid 0 in the ns
it has `CAP_SYS_ADMIN`, which lets seccomp install skip NNP. No NNP, no dumpable reset,
memfd reads work for the whole tree. The mapping preserves host uid for filesystem
ownership; tools that gate on `getuid()` see no change.
### Landlock is applied in the shim, after seccomp install
Earlier the helper installed seccomp first, then ran `landlock.RestrictPaths`.
`BestEffort()` probes via `openat`. Each probe trapped through the supervisor in the
same process. Go's GC stop-the-world needs every thread at a safepoint; a thread
suspended inside `seccomp_do_user_notification` cannot reach one. Helper hung after a
handful of notifications.
Now seccomp + Landlock both live in the shim, which is single-threaded by virtue of
being just-exec'd Go. No GC pressure during setup. Helper is unfiltered.
### No `TSYNC` on the filter
`SECCOMP_FILTER_FLAG_TSYNC` applies the filter to every thread in the group. Go runtime
threads (GC, sysmon, netpoll) routinely `openat`, all would trap and deadlock the same
way the unsandboxed-helper variant did.
Without TSYNC the filter is only on the installing thread. Descendants inherit it via
`clone()` and `execve()` anyway, so we get the same coverage without polluting the Go
runtime threads.
### `Stop()` wakes the supervisor via an eventfd
Closing `notifyFd` does not wake an `ioctl(SECCOMP_IOCTL_NOTIF_RECV)` blocker. We
`ppoll` over `notifyFd` + an eventfd; `Stop()` writes to the eventfd. See
`waitForNotif` in `landlock_seccomp_linux.go`.
### `landlockReadAccess` includes `EXECUTE`
Bubblewrap's `--ro-bind` permits execve implicitly. Landlock requires explicit
`AccessFSExecute`. Without it `allow_read: /` blocks every binary load. We bake EXECUTE
into read access; deny-exec is still enforced by the seccomp supervisor.
### Per-PID `/proc/<pid>/mem` cache, invalidated on execve
`execve` reshapes the address space; the cached fd returns EOF afterwards. We
invalidate on each execve notification (`seccompPhase.invalidateMemFd`) and lazily
reopen via `memFdFor`. Grandchildren get their own entries.
### Deny matcher treats a path as its own subtree
`GetMandatoryDenyPatterns` emits `/home/user/.ssh` (no trailing slash). The matcher
covers the path itself and anything beneath `entry+"/"`, so `~/.ssh/id_rsa` is caught.
Trailing-slash entries still prefix-match.
## Go-specific nuances
The Landlock+seccomp pattern was designed around the C/Rust threading model. Go pays a
constant tax that maps to most of the decisions above:
- **Multi-threaded from `main()`.** Go always has GC, sysmon, netpoll threads. There is
no single-threaded mode. TSYNC turns those threads into traffic for our supervisor.
- **GC stop-the-world vs. seccomp wait.** A goroutine suspended by the kernel inside
a seccomp trap cannot reach a GC safepoint. STW blocks. The supervisor goroutine,
which would unblock the trap, never runs. Rust has no GC, no STW.
- **No code injection between fork and execve.** Go's `exec.Cmd` does
`clone()` + a hardcoded sequence + `execve()`. There is no `PreExecFn` field. The
shim subcommand exists to provide a hookpoint that doesn't exist in `os/exec`. In
Rust this is inline post-`fork()`.
- **`unshare(CLONE_NEWUSER)` rejects multi-threaded callers.** A Go program cannot
enter a new user namespace from `main()`. We route the namespace through
`clone(CLONE_NEWUSER)` on the child path of `cmd.Start` instead.
- **`runtime.LockOSThread` is mandatory** wherever per-thread state matters
(NNP, seccomp install, the supervisor's `ppoll`/`ioctl` loop). Otherwise Go's
scheduler will move the goroutine and the per-thread state goes with the wrong
thread.
## Limitations
- **Unprivileged user namespaces required.** On distros that disable them, `clone()`
returns EPERM. We don't yet probe and fall back to bubblewrap (TODO).
- **Network filtering not enforced.** Landlock V4 does TCP ports, not hostnames. Use
proxy-mode.
- **PID/IPC namespace isolation is best-effort.** Retried without on EPERM.
- **Audit events are dropped.** Wired but consumed by `io.Discard`.
- **TOCTOU between path read and deny response.** Microseconds. Adequate for benign
install scripts; not a hardened defense.
+57 -3
View File
@@ -48,7 +48,8 @@ like `${HOME}/**` do not opt out of `${HOME}/.aws`. The unnamed absolute form st
## Requirements
- Bubblewrap on Linux
- Linux kernel 5.13+ with Landlock enabled (default, no external dependencies)
- Bubblewrap on Linux (fallback for kernels < 5.13, or when `PMG_SANDBOX_DRIVER=bubblewrap` is set)
- Seatbelt on MacOS
<details>
@@ -203,13 +204,66 @@ Next time you run `pmg pnpm install`, the custom policy template will be used in
| Platform | Supported | Implementation |
| -------- | --------- | ----------------------------------- |
| MacOS | Yes | Seatbelt sandbox-exec |
| Linux | Yes | Bubblewrap with namespace isolation |
| Linux | Yes | Landlock (default, kernel 5.13+) or Bubblewrap (fallback) |
| Windows | No | Not yet supported |
### Platform-Specific Limitations
<details>
<summary>Linux (Bubblewrap)</summary>
<summary>Linux (Landlock, default)</summary>
**Default sandbox on kernel 5.13+**: Landlock provides kernel-native filesystem access control
without requiring external binaries or unprivileged user namespaces.
For the architecture, design tradeoffs, and known limitations see
[sandbox-landlock.md](./sandbox-landlock.md).
**Deny enforcement**: Deny rules (DenyRead, DenyWrite, DenyExec) are enforced via seccomp
user notifications. This introduces a small TOCTOU window (microseconds) between reading
the path and responding.
**Deny enforcement across the process tree**: seccomp-notify resolves the path argument of
an intercepted `openat(2)` by reading `/proc/<pid>/mem` of the trapping process. PMG ships
this in a two-stage architecture so enforcement applies to direct targets AND every
descendant (grandchildren, great-grandchildren, etc.):
1. The helper process (`pmg __landlock_sandbox_exec`) clones a tiny shim
(`pmg __landlock_shim`) with `CLONE_NEWUSER` and a uid/gid map of `0 -> host uid`.
The shim runs as uid 0 inside a fresh user namespace so it has `CAP_SYS_ADMIN` in that
namespace.
2. The shim installs the seccomp-notify filter **without** `PR_SET_NO_NEW_PRIVS` (permitted
by `CAP_SYS_ADMIN` in the ns). It then applies Landlock and `execve`s the real target.
3. Because `NO_NEW_PRIVS` was never set, subsequent `execve` calls in the tree do **not**
reset `dumpable` to 0, so the helper can keep opening `/proc/<pid>/mem` for any
descendant. Deny rules like `~/.ssh` are enforced for the full process tree.
The user namespace is purely a capability vehicle. Host uid/gid are preserved through the
mapping, so targets see the same filesystem ownership they normally would. Tools that
refuse to run as root (npm's root-in-container warning) are unaffected because the
outside-view uid never changes.
**Requirements**: unprivileged user namespaces must be enabled (`unprivileged_userns_clone=1`
on Debian/Ubuntu; default on most modern distros). If disabled, the helper fails with an
EPERM on `clone()` and the sandbox falls back to Bubblewrap.
**Network filtering**: Not enforced. Landlock supports TCP port filtering only (V4+, no hostname).
Use `--proxy-mode` for network control.
**PID/IPC namespace isolation**: Applied best-effort via `CLONE_NEWPID|CLONE_NEWIPC|CLONE_NEWNS`.
If unavailable, a warning is printed and the command continues. Set `PMG_SANDBOX_DRIVER=bubblewrap`
to force Bubblewrap if namespace isolation is required.
**`/proc` access**: The sandbox supervisor requires `/proc` read access. When PID namespace
isolation succeeds, `/proc` is scoped to the child's namespace. When it fails, `/proc`
exposes all system processes.
**Fallback**: If Landlock is unavailable (kernel < 5.13), Bubblewrap is used automatically.
Set `PMG_SANDBOX_DRIVER=bubblewrap` to force Bubblewrap.
</details>
<details>
<summary>Linux (Bubblewrap, fallback)</summary>
**Filesystem permissions are coarse-grained**: [Bubblewrap](https://github.com/containers/bubblewrap) uses bind mounts for filesystem isolation.