Files
agent-skills/skills/proxmox-admin/SKILL.md
T
MalinandClaude Sonnet 5 1d24660303 skills(wordpress-plugin-conventions): merge security/lifecycle checks from wordpress/agent-skills
Per the 2026-08-15 skills.sh and autoskills.sh scans, both flagged
wordpress/agent-skills' wp-plugin-development module (Automattic-origin,
now WordPress-org-hosted) as high-value source material for iWP's plugin
skill: nonce+capability dual-check discipline, late escaping, prepared
SQL, cron idempotency, and uninstall-vs-deactivation guardrails.

Adapted (not copied) against real iWP plugin code in wp-plugins/:
- nonce+capability must-both framing, cited against
  class-iwp-cache-db-cleanup.php's actual AJAX handler
- late-escaping and wp_unslash()/explicit-key superglobal reading
- %i identifier-placeholder version gate (WP 6.2+, most iWP plugins
  floor at 6.0 or lower)
- new "Admin settings" section documenting the real Settings-API vs.
  AJAX-dashboard split across the suite, since the source's generic
  Settings-API-first prescription doesn't match roughly half of iWP's
  plugins
- new cron idempotency section citing the existing wp_next_scheduled()
  guard already used consistently in iwp-cache/iwp-woosales/iwp-booking
- new uninstall-vs-deactivation section flagging that only 3 of ~15
  plugins ship uninstall.php despite most creating options/tables
- new release-packaging checklist tied to iWP's actual IWP_Updater
  version-wiring convention (header/constant/updater param must agree)

Provenance noted inline with source URL. Left out the source's generic
architecture/Settings-API prescription and its detect_plugins.mjs
script (skill's house style is prose-only, no bundled scripts).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 21:26:46 +02:00

214 lines
11 KiB
Markdown

---
name: proxmox-admin
description: Use when administering a Proxmox VE host or cluster via CLI (qm, pct, pvesm, pvecm, pveam, pvesh) — inspecting, creating, or mutating VMs, containers, storage, networking, or cluster state. Safety-first: read-only discovery and explicit target validation before any mutating command, a destructive-command tier list, and real gotchas from hands-on fleet use (reset-vs-reboot tradeoff, Tailscale hostname collisions, storage free-space checks).
license: MIT
source: adapted from https://github.com/bastos/skills/tree/main/proxmox-admin (MIT) — command taxonomy and CLI coverage — plus this fleet's own direct hands-on Proxmox operations across two hosts, which supplied the safety framing, the real gotchas below, and a correction to the source's quick-reference style (see "Why this skill differs from the source" at the bottom)
---
# Proxmox VE Administration
Proxmox VE is a server virtualization platform built on Debian, managing
KVM virtual machines (`qm`) and LXC containers (`pct`) plus storage
(`pvesm`), clustering (`pvecm`), templates (`pveam`), and the API shell
(`pvesh`). This skill covers CLI-based administration.
**Not for:** Web UI-only workflows (use the CLI equivalents below and in
`references/command-reference.md`).
**Fleet-specific values (actual VMIDs, storage pool names, bridge/VLAN
IDs, node names) live in this fleet's own `docs/server-*.md` files, not
here.** Every command below uses a placeholder (`<vmid>`, `<node>`,
`<storage>`, `<bridge>`) — resolve the real value from live discovery
commands or fleet docs before running anything, never by guessing or
reusing an example number from this file or from upstream documentation.
## SAFETY FIRST — the order every task follows
1. **Read-only discovery first, always.** Before touching a single VM or
container, know what actually exists and what state it's in.
2. **Validate the target explicitly.** Never assume a VMID, node name, or
storage name is correct — confirm it via `qm config`/`pct config`/
`qm list` immediately before acting on it, every time, even if you
"already know" the ID from earlier in the conversation.
3. **Know which tier a command is in before running it.** Destructive-tier
commands (see below) should never run unattended — expect to be
blocked by this environment's own permission classifier, and treat
that as correct behavior to work *with*, not around.
## Read-only discovery commands (run these before anything else)
```bash
qm list # all VMs on this node: ID, name, status
qm config <vmid> # full config of one VM — confirm identity before acting
qm status <vmid> # current run state
pct list # all containers on this node
pct config <ctid> # full config of one container
pct status <ctid>
pvesm status # storage pools: type, active, TOTAL/USED/AVAIL
pvecm status # cluster quorum/membership (no-op, harmless, on single-node too)
pvecm nodes # cluster node list
pvesh get /cluster/resources # everything, via the API shell — useful for cross-node views
```
Run the relevant subset of these before every task in this skill, not
just once at the start of a session — state changes, and a VMID that was
stopped five minutes ago may be running now.
## Never trust a VMID or hostname without confirming it first
`qm list` gives you IDs and names; `qm config <vmid>` gives you the full
picture (disks, storage backing, network, memory). **Always run `qm
config <vmid>` (or `pct config <ctid>`) immediately before any mutating
command against that ID**, even if the ID was already established earlier
in the same task — configs can change between when you last checked and
when you act, and a copy-pasted or misremembered ID silently targeting
the wrong guest is a much worse failure mode than a few extra seconds of
discovery.
## The Tailscale hostname-collision gotcha
**Confirmed real risk, hit twice in this session** across the `vulture`
and `ginger` hosts: a bare Tailscale MagicDNS hostname (e.g. `ssh
<hostname>`) can resolve to the **wrong physical host** if there's a
stale or duplicate device name registered in the tailnet. This is
dangerous specifically because it fails silently — the SSH connection
succeeds, a shell prompt comes back, and everything looks normal; you
just end up running Proxmox commands against a different host's guests
than the one you meant.
**Before running any command against what you believe is a specific
host** (and especially before anything in the destructive tier below),
verify the connection actually went where you think:
```bash
ssh -v <host> 2>&1 | grep "Connecting to"
```
If a hostname's identity is ever in doubt, skip MagicDNS entirely and
connect by the host's direct Tailscale IP instead — an IP can't collide
the way a device name can.
## Snapshot vs. backup — know which one you need
These are not interchangeable, and confusing them is a common source of
false confidence before a destructive operation:
- **Snapshot** (`qm snapshot`, `qm rollback`, `qm listsnapshot`,
`qm delsnapshot`) — fast, live, stored *on the same storage backend* as
the disk it snapshots. Good for "undo the last five minutes of change"
during an upgrade or config edit. **Does not protect you if the
underlying storage itself fails, fills up, or gets deleted** — the
snapshot dies with it.
- **Backup** (`vzdump`, restored via `qmrestore`/`pct restore`) — a full
archive, ideally written to storage genuinely separate from the guest's
own disk. Slower, but survives storage-level failure and is what you
actually want before anything in the destructive tier below.
Rule of thumb: before any destructive-tier command, you want a **real
backup on separate storage**, not just a same-storage snapshot. A
snapshot is a convenience for reversible experiments, not a safety net
for irreversible ones.
## Reset vs. reboot — the real tradeoff learned this session
- `qm reboot <vmid>` sends a graceful ACPI shutdown-then-restart request
*into the guest*. If the guest is hung, has no ACPI daemon running, or
is otherwise unresponsive, **this can hang or time out waiting
indefinitely** — it depends entirely on the guest cooperating.
- `qm reset <vmid>` is a hard reset — functionally identical to yanking
power. It was used exactly once this session, deliberately, as the real
fallback when `reboot` wasn't a safe option — but **only after first
confirming the filesystems inside the guest were already cleanly
unmounted**. A hard reset against a guest with dirty/mounted
filesystems risks the same corruption a real power loss would.
**Do not reach for `reset` just because `reboot` seems slow.** Check what
the guest is actually doing first — console (`qm monitor <vmid>`), guest
agent (`qm agent <vmid> ping`), or a direct login — and only escalate to
a hard reset once you've confirmed a clean-unmount state (or have
independently decided the guest's disk state doesn't matter, e.g. it's
about to be destroyed anyway).
## Storage checks before any disk-affecting operation
`pvesm status` reports both which storage pools exist **and** their
free space (`TOTAL`/`USED`/`AVAIL` columns) — check both, not just
"does the storage exist." This matters concretely for restores: this
fleet's real VM disks run 200GB+, and a large `qmrestore` can fail
partway through — or silently fill the target storage and start
affecting other guests sharing it — if free space wasn't checked first.
Confirm `pvesm status` shows enough headroom on the *target* storage
before starting any restore, clone, or disk-resize operation, not just
before the operation you're focused on.
## Destructive-command tier list — expect to be blocked, and that's correct
**This environment's own permission classifier already blocks
`qmrestore` and other VM-creating/mutating commands, requiring explicit
human confirmation before they run.** This is real, enforced behavior in
this session's environment, not a hypothetical policy. Treat it as a
guardrail to design around, not a friction to route past — if a
destructive-tier command gets blocked, **stop, explain to the human
exactly what the command would do and why it's in this tier, and let
them decide.** Do not look for an equivalent path that avoids the
prompt (a raw `pvesh` API call doing the same mutation, hand-editing
`/etc/pve/qemu-server/<vmid>.conf` directly, chaining lower-privilege
commands to reach the same end state) — that defeats the guardrail's
purpose rather than satisfying it.
**Never run these without explicit human confirmation first:**
- Any `*destroy*`: `qm destroy`, `pct destroy` (with or without
`--purge`) — deletes the guest and its disks, generally unrecoverable
without a separate backup
- Any restore that creates or overwrites a guest: `qmrestore`,
`pct restore`
- Any storage mutation: `pvesm add`, `pvesm remove`
- Any cluster mutation: `pvecm create`, `pvecm add`, `pvecm delnode`,
`pvecm expected` (force-quorum — a single-node recovery action with
real risk of split-brain if used carelessly on a healthy cluster)
- Any migration: `qm migrate`, `pct migrate` — moves a guest between
nodes, can fail mid-flight, and affects a second node's state
- `qm template <vmid>` — converting a VM to a template is irreversible
- `qm reset <vmid>` — lower severity than destroy, but still surface it
given the hang-vs-corruption tradeoff above; confirm clean-unmount
status even where the environment doesn't hard-block it
- Any bulk/loop pattern (batch-create N containers, batch-destroy,
scripted mass reconfiguration) — higher risk than the single-target
equivalent because the blast radius multiplies silently and a mid-loop
failure can leave a mix of applied and unapplied state
**Generally fine to run directly** (read-only or narrowly-scoped,
reversible): `qm list`/`config`/`status`, `pct list`/`config`/`status`,
`pvesm status`, `pvecm status`/`nodes`, `pvesh get ...`, `qm agent <vmid>
ping`/`get-osinfo`, `qm start`/`shutdown`/`stop` on a guest whose
identity you've just confirmed, and snapshot creation (still confirm the
target first — a snapshot against the wrong VMID is a wasted, confusing
action even if not a destructive one).
## Full command reference
See `references/command-reference.md` for the complete `qm`/`pct`/
`pvesm`/`pvecm`/`pveam`/backup/troubleshooting command tables, adapted
from the upstream source with all example VMIDs, storage names, and
bridge names replaced by placeholders.
## Why this skill differs from the source
The upstream `bastos/skills@proxmox-admin` skill this was adapted from is
a solid 427-line command taxonomy, but it presents everything —
including `qm destroy`, `pct destroy`, storage mutation, and cluster
operations — in flat quick-reference style with no confirmation, target-
validation, or backup guardrails, and its examples hardcode specific
VMIDs/storage/bridge names as if they were universal defaults to copy.
This version keeps the command coverage but restructures around what
this session's actual hands-on Proxmox work across two hosts (`vulture`,
`ginger`) surfaced as the real operating discipline: discovery before
mutation, explicit target confirmation, a real destructive-tier list that
matches this environment's own enforced permission classifier, and three
gotchas (Tailscale hostname collisions, reset-vs-reboot, storage
free-space checks) that cost real time or carried real risk this session.