Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
3098342331 | ||
|
|
11ddc0c981 | ||
|
|
14d4f2e21a | ||
|
|
ec83e5bace | ||
|
|
7066cc8154 | ||
|
|
fcf7663366 | ||
|
|
e07431f127 | ||
|
|
52bbbcdd86 | ||
|
|
6566ec23d3 | ||
|
|
ed1fbb7f4c | ||
|
|
dcd51f16fe | ||
|
|
d833f4adab | ||
|
|
7693a046ee | ||
|
|
710173f965 | ||
|
|
0d95a81d35 | ||
|
|
e2c09c9a7e | ||
|
|
dc95e915c5 | ||
|
|
686c0cce64 | ||
|
|
e9cb8435d6 | ||
|
|
a8846964b5 | ||
|
|
6919dc2c4f | ||
|
|
e6cce23ed6 | ||
|
|
ec4f664910 | ||
|
|
d28be46e50 | ||
|
|
335ab56dea | ||
|
|
8a5ab60ef0 | ||
|
|
bceaa0cf6d | ||
|
|
19f7b5a0c0 | ||
|
|
5f079ab7d8 | ||
|
|
51e1990340 | ||
|
|
15465d30b5 | ||
|
|
daa8ec3732 | ||
|
|
e7ee55f936 | ||
|
|
bd36cdc953 | ||
|
|
452055463b | ||
|
|
a2cb2f95e8 | ||
|
|
c1e1ffdd47 | ||
|
|
14eea5d5ee | ||
|
|
cdb0d0c613 | ||
|
|
9663344b79 | ||
|
|
d900443b84 | ||
|
|
5b875b8e2e | ||
|
|
79a14ee79f | ||
|
|
c0e3c1e140 | ||
|
|
79285a4c00 | ||
|
|
7dc87d0a32 | ||
|
|
516f917619 | ||
|
|
ae02d2a560 | ||
|
|
2fb7c91183 | ||
|
|
704a460ef8 | ||
|
|
4dad1afe84 | ||
|
|
70ea40dced | ||
|
|
959b943ad4 | ||
|
|
6d71b9ec6f | ||
|
|
e0bb39f267 | ||
|
|
93d9a439b4 | ||
|
|
e8139f86e3 | ||
|
|
6ebdb375c4 | ||
|
|
62446eb525 | ||
|
|
62e77e66ff | ||
|
|
f60a45a98d | ||
|
|
c1ac0664c9 | ||
|
|
a325c64dd4 | ||
|
|
8753bf90d5 | ||
|
|
25d22f864a | ||
|
|
32f37af81a | ||
|
|
ddf212173a | ||
|
|
22e59bf067 | ||
|
|
efd490a5fc | ||
|
|
eb447dad2e | ||
|
|
baf10366f6 | ||
|
|
b5e8cdaf1f | ||
|
|
bdcba4c6ef | ||
|
|
22f75e3de6 | ||
|
|
f0dd780305 | ||
|
|
cd4009effa | ||
|
|
6ce47d4445 | ||
|
|
46de578a7e | ||
|
|
f68111acc3 | ||
|
|
f06348cbd3 | ||
|
|
e8eae0ff28 | ||
|
|
655b7d9c54 | ||
|
|
2422bbe59c | ||
|
|
18e45320a0 | ||
|
|
63a91ecadb | ||
|
|
e8b69797a8 | ||
|
|
279310971a | ||
|
|
21b15183ed | ||
|
|
754795ada4 | ||
|
|
dce88c2f2e | ||
|
|
e843518438 | ||
|
|
7b2bed2d0b | ||
|
|
8b8068d8ca | ||
|
|
1ce38a2295 | ||
|
|
4f2314ae43 | ||
|
|
5786dd56cc | ||
|
|
a2dca8194d | ||
|
|
b67c404e80 | ||
|
|
3c1743ea11 | ||
|
|
086c43077d | ||
|
|
4e1fdca273 | ||
|
|
f62ac792ac | ||
|
|
37ffdd4253 | ||
|
|
abaad10665 | ||
|
|
b3468f4333 | ||
|
|
1b2ee4128e | ||
|
|
eda0a1e7c6 | ||
|
|
c28067e22d | ||
|
|
ed706fa098 | ||
|
|
509935697b | ||
|
|
ef6050c5e1 | ||
|
|
e031c1e440 | ||
|
|
83ec61c509 | ||
|
|
ec7ce3614b | ||
|
|
47cd8d72de | ||
|
|
56875e883f | ||
|
|
de331644ef | ||
|
|
b685570740 | ||
|
|
13f0d4c49a | ||
|
|
31fa95478d | ||
|
|
a9dc067796 | ||
|
|
31d804e5f2 | ||
|
|
a660db02c9 | ||
|
|
b24b915e0a | ||
|
|
57a9b489a1 | ||
|
|
84cc3c14fa | ||
|
|
c2ed24b3e5 | ||
|
|
5ad8f6d684 | ||
|
|
4c82699c26 | ||
|
|
d8d53115d0 | ||
|
|
205e537103 | ||
|
|
adafaba5cd | ||
|
|
e50325b040 | ||
|
|
23ce800ad6 | ||
|
|
4fe5a5d70a | ||
|
|
507ec0af22 | ||
|
|
4345efd45a | ||
|
|
7e89c6a3cd | ||
|
|
b1ca6818f9 | ||
|
|
5af591a028 | ||
|
|
b50aa6d4b3 | ||
|
|
fea439dc14 | ||
|
|
d63b47673c | ||
|
|
3897f40b97 | ||
|
|
52fec218a5 | ||
|
|
1e0fb7f3dd | ||
|
|
e192972632 | ||
|
|
7b4649fad0 | ||
|
|
63e797cd75 | ||
|
|
f7ef823b91 | ||
|
|
4c6dad829a | ||
|
|
56f4047ab3 | ||
|
|
3ed7b732b0 | ||
|
|
600e8efcd6 | ||
|
|
80a317e472 | ||
|
|
6a6c144f28 | ||
|
|
8be1bdd4c7 | ||
|
|
8bab5c739e | ||
|
|
c80f8d7fe1 | ||
|
|
12e78c5201 | ||
|
|
ea89000141 | ||
|
|
24e6ee9ea8 | ||
|
|
0bbd46c390 | ||
|
|
a311eab867 | ||
|
|
bce42f7294 | ||
|
|
8783f63296 | ||
|
|
7e2c94c541 | ||
|
|
8fb2f013a9 | ||
|
|
e8015c384c | ||
|
|
b6a7a3f584 | ||
|
|
79fa46a4a0 | ||
|
|
cfd76590bd | ||
|
|
5580f8e28f | ||
|
|
7906016383 | ||
|
|
9449c8a05a | ||
|
|
92f892f2b9 | ||
|
|
bc351bb7c5 | ||
|
|
558576cdcd | ||
|
|
1be0340955 | ||
|
|
782d17c24a | ||
|
|
18427c26ab | ||
|
|
602d1639c6 | ||
|
|
a9e62b48c3 | ||
|
|
4e919547fb | ||
|
|
a62d6602c3 | ||
|
|
26c25e39b3 | ||
|
|
79ef2933ac | ||
|
|
fecea29b1f | ||
|
|
6f6772ab1b | ||
|
|
e6df0d6055 | ||
|
|
42b3030c69 | ||
|
|
bd671cc907 | ||
|
|
4591fb5396 | ||
|
|
e0c7501619 | ||
|
|
a5b624ca4f | ||
|
|
3e8b60e570 | ||
|
|
7d6aa3ac65 | ||
|
|
e1e5e8b13a | ||
|
|
8195e1a4a0 | ||
|
|
a97af68741 | ||
|
|
f6ca007f49 | ||
|
|
bf828cd0fa | ||
|
|
64c4e2113f | ||
|
|
6588ad1e26 | ||
|
|
784df686b4 | ||
|
|
7de87db9e6 | ||
|
|
1254e714e6 | ||
|
|
6391a702b7 | ||
|
|
2b10489888 | ||
|
|
001075b420 | ||
|
|
03f3dcd723 | ||
|
|
f3c710377c | ||
|
|
9ee0e352c6 | ||
|
|
d6fbf67db3 | ||
|
|
1b15fc7170 | ||
|
|
dfcdbc0eee | ||
|
|
7e8b5f36cf | ||
|
|
d37c921f29 | ||
|
|
9a8e9c4959 | ||
|
|
636695bd0f | ||
|
|
071926541d | ||
|
|
0875a71f8d | ||
|
|
c512cb8487 | ||
|
|
981b6069d8 | ||
|
|
4863bc20d9 | ||
|
|
a3b255235d | ||
|
|
4cfbd9c4ba | ||
|
|
c3ca528c76 | ||
|
|
984f63fa9e | ||
|
|
58ca4a6d3e | ||
|
|
942ecb87a8 | ||
|
|
edcefa328e | ||
|
|
bca1e756d2 | ||
|
|
6971824898 |
@@ -1,20 +0,0 @@
|
||||
{
|
||||
"name": "caveman-repo",
|
||||
"interface": {
|
||||
"displayName": "Caveman Repo"
|
||||
},
|
||||
"plugins": [
|
||||
{
|
||||
"name": "caveman",
|
||||
"source": {
|
||||
"source": "local",
|
||||
"path": "./plugins/caveman"
|
||||
},
|
||||
"policy": {
|
||||
"installation": "AVAILABLE",
|
||||
"authentication": "ON_INSTALL"
|
||||
},
|
||||
"category": "Productivity"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"$schema": "https://anthropic.com/claude-code/marketplace.schema.json",
|
||||
"name": "caveman",
|
||||
"description": "Ultra-compressed communication mode for Claude Code. Cuts ~75% of tokens while keeping full technical accuracy.",
|
||||
"description": "Ultra-compressed communication mode for Claude Code. Cuts 65% of output tokens (measured) while keeping full technical accuracy.",
|
||||
"owner": {
|
||||
"name": "Julius Brussee",
|
||||
"url": "https://github.com/JuliusBrussee"
|
||||
@@ -9,7 +9,7 @@
|
||||
"plugins": [
|
||||
{
|
||||
"name": "caveman",
|
||||
"description": "Talk like caveman. Cut ~75% tokens. Keep all technical accuracy.",
|
||||
"description": "Talk like caveman. Cut 65% output tokens (measured). Keep all technical accuracy.",
|
||||
"source": "./",
|
||||
"category": "productivity"
|
||||
}
|
||||
|
||||
@@ -1,8 +1,34 @@
|
||||
{
|
||||
"name": "caveman",
|
||||
"description": "Ultra-compressed communication mode. Cuts ~75% of tokens while keeping full technical accuracy by speaking like a caveman.",
|
||||
"description": "Ultra-compressed communication mode. Cuts 65% of output tokens (measured) while keeping full technical accuracy by speaking like a caveman.",
|
||||
"author": {
|
||||
"name": "Julius Brussee",
|
||||
"url": "https://github.com/JuliusBrussee"
|
||||
},
|
||||
"hooks": {
|
||||
"SessionStart": [
|
||||
{
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "node \"${CLAUDE_PLUGIN_ROOT}/src/hooks/caveman-activate.js\"",
|
||||
"timeout": 5,
|
||||
"statusMessage": "Loading caveman mode..."
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"UserPromptSubmit": [
|
||||
{
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "node \"${CLAUDE_PLUGIN_ROOT}/src/hooks/caveman-mode-tracker.js\"",
|
||||
"timeout": 5,
|
||||
"statusMessage": "Tracking caveman mode..."
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,5 @@
|
||||
[features]
|
||||
# Both keys: codex-cli renamed codex_hooks → hooks; old versions (<=0.120.0)
|
||||
# silently ignore unknown keys, so shipping both activates on either (#617).
|
||||
hooks = true
|
||||
codex_hooks = true
|
||||
@@ -0,0 +1,17 @@
|
||||
{
|
||||
"hooks": {
|
||||
"SessionStart": [
|
||||
{
|
||||
"matcher": "startup|resume",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "echo 'CAVEMAN MODE ACTIVE. Rules: Drop articles/filler/pleasantries/hedging. Fragments OK. Short synonyms. Pattern: [thing] [action] [reason]. [next step]. Not: Sure! I would be happy to help you with that. Yes: Bug in auth middleware. Fix: Code/commits/security: write normal. User says stop caveman or normal mode to deactivate.'",
|
||||
"timeout": 5,
|
||||
"statusMessage": "Loading caveman mode"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,9 @@
|
||||
root = true
|
||||
|
||||
[*]
|
||||
indent_style = space
|
||||
indent_size = 2
|
||||
end_of_line = lf
|
||||
charset = utf-8
|
||||
trim_trailing_whitespace = true
|
||||
insert_final_newline = true
|
||||
@@ -0,0 +1,15 @@
|
||||
# These are supported funding model platforms
|
||||
|
||||
github: JuliusBrussee
|
||||
patreon: # Replace with a single Patreon username
|
||||
open_collective: # Replace with a single Open Collective username
|
||||
ko_fi: # Replace with a single Ko-fi username
|
||||
tidelift: # Replace with a single Tidelift platform-name/package-name e.g., npm/babel
|
||||
community_bridge: # Replace with a single Community Bridge project-name e.g., cloud-foundry
|
||||
liberapay: # Replace with a single Liberapay username
|
||||
issuehunt: # Replace with a single IssueHunt username
|
||||
lfx_crowdfunding: # Replace with a single LFX Crowdfunding project-name e.g., cloud-foundry
|
||||
polar: # Replace with a single Polar username
|
||||
buy_me_a_coffee: # Replace with a single Buy Me a Coffee username
|
||||
thanks_dev: # Replace with a single thanks.dev username
|
||||
custom: # Replace with up to 4 custom sponsorship URLs e.g., ['link1', 'link2']
|
||||
@@ -0,0 +1,46 @@
|
||||
name: CI
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
pull_request:
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
jobs:
|
||||
node:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 15
|
||||
strategy:
|
||||
matrix:
|
||||
node-version: [18, 20, 22]
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/setup-node@v4
|
||||
with:
|
||||
node-version: ${{ matrix.node-version }}
|
||||
- name: Installer test suite
|
||||
run: npm test
|
||||
- name: Standalone hook/tool tests
|
||||
run: |
|
||||
for f in tests/test_*.js; do
|
||||
echo "== $f"
|
||||
node "$f"
|
||||
done
|
||||
|
||||
python:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/setup-python@v5
|
||||
with:
|
||||
python-version: '3.11'
|
||||
# Several python tests shell out to node for hook checks — don't rely
|
||||
# on the runner image happening to ship it.
|
||||
- uses: actions/setup-node@v4
|
||||
with:
|
||||
node-version: 20
|
||||
- name: Python test suite
|
||||
run: python -m unittest discover -s tests -v
|
||||
@@ -1,10 +1,14 @@
|
||||
name: Sync SKILL.md
|
||||
name: Sync SKILL.md and rules
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
paths:
|
||||
- skills/caveman/SKILL.md
|
||||
- skills/cavecrew/SKILL.md
|
||||
- agents/cavecrew-*.md
|
||||
- skills/caveman-compress/SKILL.md
|
||||
- skills/caveman-compress/scripts/**
|
||||
|
||||
concurrency:
|
||||
group: sync-skill
|
||||
@@ -26,18 +30,41 @@ jobs:
|
||||
git config user.email "github-actions[bot]@users.noreply.github.com"
|
||||
git pull --rebase origin main
|
||||
|
||||
- name: Copy to duplicate locations
|
||||
- name: Sync SKILL.md copies
|
||||
run: |
|
||||
cp skills/caveman/SKILL.md caveman/SKILL.md
|
||||
cp skills/caveman/SKILL.md plugins/caveman/skills/caveman/SKILL.md
|
||||
|
||||
- name: Rebuild caveman.skill ZIP
|
||||
- name: Sync caveman-compress skill to plugin
|
||||
run: |
|
||||
cd skills && zip -r ../caveman.skill caveman/
|
||||
# Plugin distribution mirrors source verbatim — no rename, no sed.
|
||||
mkdir -p plugins/caveman/skills/caveman-compress
|
||||
cp skills/caveman-compress/SKILL.md plugins/caveman/skills/caveman-compress/SKILL.md
|
||||
rm -rf plugins/caveman/skills/caveman-compress/scripts
|
||||
cp -r skills/caveman-compress/scripts plugins/caveman/skills/caveman-compress/scripts
|
||||
rm -rf plugins/caveman/skills/caveman-compress/scripts/__pycache__
|
||||
|
||||
- name: Sync cavecrew skill + agents to plugin
|
||||
run: |
|
||||
mkdir -p plugins/caveman/skills/cavecrew plugins/caveman/agents
|
||||
cp skills/cavecrew/SKILL.md plugins/caveman/skills/cavecrew/SKILL.md
|
||||
cp agents/cavecrew-investigator.md plugins/caveman/agents/cavecrew-investigator.md
|
||||
cp agents/cavecrew-builder.md plugins/caveman/agents/cavecrew-builder.md
|
||||
cp agents/cavecrew-reviewer.md plugins/caveman/agents/cavecrew-reviewer.md
|
||||
|
||||
- name: Rebuild caveman.skill ZIP
|
||||
run: mkdir -p dist && cd skills && zip -r ../dist/caveman.skill caveman/
|
||||
|
||||
- name: Commit and push if changed
|
||||
run: |
|
||||
git diff --quiet && exit 0
|
||||
git add caveman/SKILL.md plugins/caveman/skills/caveman/SKILL.md caveman.skill
|
||||
git commit -m "chore: sync SKILL.md copies and caveman.skill [skip ci]"
|
||||
git add \
|
||||
skills/caveman-compress/ \
|
||||
plugins/caveman/skills/caveman-compress/ \
|
||||
plugins/caveman/skills/caveman/SKILL.md \
|
||||
plugins/caveman/skills/cavecrew/SKILL.md \
|
||||
plugins/caveman/agents/cavecrew-investigator.md \
|
||||
plugins/caveman/agents/cavecrew-builder.md \
|
||||
plugins/caveman/agents/cavecrew-reviewer.md \
|
||||
dist/caveman.skill
|
||||
git diff --staged --quiet && exit 0
|
||||
git commit -m "chore: sync SKILL.md copies [skip ci]"
|
||||
git push
|
||||
|
||||
@@ -3,5 +3,15 @@ __pycache__/
|
||||
*.pyc
|
||||
.venv/
|
||||
.env.local
|
||||
caveman-compress.md
|
||||
**/.DS_Store
|
||||
.claude/worktrees/
|
||||
evals/snapshots/*.html
|
||||
evals/snapshots/*.png
|
||||
context/refs/research-brief-caveman-code-efficiency.md
|
||||
|
||||
# Build artifacts
|
||||
dist/*
|
||||
!dist/caveman.skill
|
||||
|
||||
# Local star-history chart scratch (not part of the product)
|
||||
tmp-starcharts/
|
||||
|
||||
@@ -0,0 +1,4 @@
|
||||
@./skills/caveman/SKILL.md
|
||||
@./skills/caveman-commit/SKILL.md
|
||||
@./skills/caveman-review/SKILL.md
|
||||
@./skills/caveman-compress/SKILL.md
|
||||
@@ -0,0 +1,306 @@
|
||||
# CLAUDE.md — caveman
|
||||
|
||||
## README is a product artifact
|
||||
|
||||
README = product front door. Non-technical people read it to decide if caveman worth install. Treat like UI copy.
|
||||
|
||||
**Rules for any README change:**
|
||||
|
||||
- Readable by non-AI-agent users. If you write "SessionStart hook injects system context," invisible to most — translate it.
|
||||
- Keep Before/After examples first. That the pitch.
|
||||
- Install table always complete + accurate. One broken install command costs real user.
|
||||
- What You Get table must sync with actual code. Feature ships or removed → update table.
|
||||
- Preserve voice. Caveman speak in README on purpose. "Brain still big." "Cost go down forever." "One rock. That it." — intentional brand. Don't normalize.
|
||||
- Benchmark numbers from real runs in `benchmarks/` and `evals/`. Never invent or round. Re-run if doubt.
|
||||
- Adding new agent to install table → add detail block in `<details>` section below.
|
||||
- Readability check before any README commit: would non-programmer understand + install within 60 seconds?
|
||||
|
||||
---
|
||||
|
||||
## Project overview
|
||||
|
||||
Caveman makes AI coding agents respond in compressed caveman-style prose — cuts 65% output tokens (measured), full technical accuracy. Ships as Claude Code plugin, Codex plugin, Gemini CLI extension, agent rule files for Cursor, Windsurf, Cline, Copilot, 40+ others via `npx skills`.
|
||||
|
||||
---
|
||||
|
||||
## What lives where
|
||||
|
||||
Post-cleanup layout. Sources of truth at the top, distribution mirrors below, build outputs in `dist/`, human docs alongside each skill.
|
||||
|
||||
```
|
||||
caveman/
|
||||
├── README.md # Front door (product pitch)
|
||||
├── INSTALL.md # Per-agent install commands
|
||||
├── CONTRIBUTING.md # Dev guide
|
||||
├── CLAUDE.md # This file (maintainer instructions)
|
||||
├── AGENTS.md / GEMINI.md # Autodiscovery files (must stay at root)
|
||||
│
|
||||
├── install.sh / install.ps1 # 30-line shims → cli/install.js
|
||||
│
|
||||
├── cli/ # Unified installer
|
||||
│ ├── install.js # Single source for all 30+ agents (PROVIDERS array)
|
||||
│ └── lib/settings.js # JSONC-tolerant settings.json reader/writer
|
||||
│
|
||||
├── skills/ # ALL skills, single source of truth
|
||||
│ ├── caveman/{SKILL.md, README.md}
|
||||
│ ├── caveman-commit/{SKILL.md, README.md}
|
||||
│ ├── caveman-review/{SKILL.md, README.md}
|
||||
│ ├── caveman-help/{SKILL.md, README.md}
|
||||
│ ├── caveman-stats/{SKILL.md, README.md}
|
||||
│ ├── caveman-compress/{SKILL.md, README.md, scripts/}
|
||||
│ └── cavecrew/{SKILL.md, README.md}
|
||||
│
|
||||
├── agents/ # cavecrew subagents (single source — kept at root for plugin auto-discovery)
|
||||
├── commands/ # Codex/Gemini TOML command stubs (root for plugin auto-discovery)
|
||||
│
|
||||
├── src/ # Internal source — not auto-discovered by plugin
|
||||
│ ├── hooks/ # Claude Code hooks (installer reads here)
|
||||
│ ├── rules/ # Auto-activation rule body (single source)
|
||||
│ ├── tools/ # caveman-init.js (per-repo rule writer)
|
||||
│ └── mcp-servers/ # caveman-shrink npm-published MCP middleware
|
||||
│
|
||||
├── .claude-plugin/ # Claude Code plugin manifest (REQUIRED at root)
|
||||
├── plugins/caveman/ # Claude Code plugin distribution (CI-mirrored)
|
||||
│ ├── skills/ # ← from skills/
|
||||
│ └── agents/ # ← from agents/
|
||||
│
|
||||
├── dist/ # Build artifacts (gitignored)
|
||||
│ └── caveman.skill # ZIP of skills/caveman/, rebuilt by CI
|
||||
│
|
||||
├── tests/ # All tests (Node + Python)
|
||||
├── benchmarks/ # Real token measurements through Claude API
|
||||
├── evals/ # Three-arm eval harness
|
||||
├── docs/ # User-facing docs site
|
||||
└── .github/workflows/ # CI sync
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## File structure and what owns what
|
||||
|
||||
### Single source of truth files — edit only these
|
||||
|
||||
| File | What it controls |
|
||||
|------|-----------------|
|
||||
| `skills/caveman/SKILL.md` | Caveman behavior: intensity levels, rules, wenyan mode, auto-clarity, persistence. Only file to edit for behavior changes. |
|
||||
| `src/rules/caveman-activate.md` | Always-on auto-activation rule body. Consumed by `src/tools/caveman-init.js` when a user runs `npx caveman --with-init` (per-repo IDE rule files). Edit here, not in any per-agent rule copy. |
|
||||
| `src/rules/caveman-openclaw-bootstrap.md` | Marker-fenced bootstrap snippet appended to `~/.openclaw/workspace/SOUL.md` by `cli/lib/openclaw.js`. Drives always-on caveman through the OpenClaw gateway. Must include the SENTINEL `Respond terse like smart caveman` and stay well under OpenClaw's 12K-per-bootstrap-file cap. |
|
||||
| `cli/lib/openclaw.js` | OpenClaw install/uninstall helper. Frontmatter merge (`version`, `always: true`), SOUL.md marker append/strip, idempotent. Shared by `cli/install.js` and `src/tools/caveman-init.js`. |
|
||||
| `skills/caveman-commit/SKILL.md` | Caveman commit message behavior. Fully independent skill. |
|
||||
| `skills/caveman-review/SKILL.md` | Caveman code review behavior. Fully independent skill. |
|
||||
| `skills/caveman-help/SKILL.md` | Quick-reference card. One-shot display, not a persistent mode. |
|
||||
| `skills/caveman-compress/SKILL.md` | Compress sub-skill behavior. |
|
||||
| `skills/cavecrew/SKILL.md` | Cavecrew decision guide — when to delegate to caveman subagents vs vanilla. Edit only here. |
|
||||
| `agents/cavecrew-investigator.md` | Read-only locator subagent (haiku). Output contract: `path:line — symbol — note`. |
|
||||
| `agents/cavecrew-builder.md` | Surgical 1-2 file editor subagent. Refuses 3+ file scope. |
|
||||
| `agents/cavecrew-reviewer.md` | Diff/file reviewer subagent (haiku). One-line findings with severity emoji. |
|
||||
| `src/plugins/opencode/plugin.js` | opencode native plugin. ESM Bun module — `session.created` writes flag, `tui.prompt.append` parses slash/natural-language activation and appends per-prompt reinforcement. Reuses `caveman-config.js` via `createRequire`. |
|
||||
| `src/plugins/opencode/commands/*.md` | Six opencode slash-command prompt templates (`/caveman`, `/caveman-{commit,review,compress,stats,help}`). |
|
||||
|
||||
### Auto-generated / auto-synced — do not edit directly
|
||||
|
||||
We removed the agent-specific dotdir mirrors at the repo root (`.cursor/`, `.windsurf/`, `.clinerules/`, `.github/copilot-instructions.md`, root `caveman/SKILL.md`). They were never read by the installer — only used to self-apply caveman to this repo when a maintainer opened it in Cursor/Windsurf/Cline. Devs who want caveman in their editor while editing this repo should run `npx caveman --with-init` once (writes per-repo rule files from `src/rules/caveman-activate.md` via `src/tools/caveman-init.js`). For per-user installs through the upstream skills CLI, `npx caveman --only <agent>` runs `npx skills add ... -a <profile>`.
|
||||
|
||||
A handful of dotdir leftovers (`.junie/`, `.kiro/`, `.roo/`, `.agents/`) still hold a stale `cavecrew/SKILL.md` mirror from before the cleanup. They aren't read by anything in the current install path; remove on sight, no migration needed.
|
||||
|
||||
What's left is the Claude Code plugin distribution (required by the plugin loader) and the release ZIP.
|
||||
|
||||
| File | Synced from |
|
||||
|------|-------------|
|
||||
| `plugins/caveman/skills/caveman/SKILL.md` | `skills/caveman/SKILL.md` |
|
||||
| `plugins/caveman/skills/caveman-compress/SKILL.md` (+ `scripts/`) | `skills/caveman-compress/SKILL.md` (+ `scripts/`) |
|
||||
| `plugins/caveman/skills/cavecrew/SKILL.md` | `skills/cavecrew/SKILL.md` |
|
||||
| `plugins/caveman/agents/cavecrew-*.md` | `agents/cavecrew-*.md` |
|
||||
| `dist/caveman.skill` | ZIP of `skills/caveman/` directory (gitignored; rebuilt by CI on release) |
|
||||
|
||||
Skills not in this table (`caveman-commit`, `caveman-review`, `caveman-help`, `caveman-stats`) are not mirrored into the Claude Code plugin distribution by CI. They reach Claude Code through the standalone hook + skill install path, and reach other agents via `npx skills add`. A `plugins/caveman/skills/caveman-stats/` directory is currently checked in as a hand-committed copy; the sync workflow does not touch it, so don't rely on edits there to propagate.
|
||||
|
||||
---
|
||||
|
||||
## CI sync workflow
|
||||
|
||||
`.github/workflows/sync-skill.yml` triggers on main push when `skills/**/SKILL.md` or `agents/cavecrew-*.md` changes.
|
||||
|
||||
What it does:
|
||||
1. Copies `skills/caveman/SKILL.md` and `skills/cavecrew/SKILL.md` into their `plugins/caveman/skills/<name>/` mirrors so the Claude Code plugin loader sees the latest behavior.
|
||||
2. Copies `skills/caveman-compress/SKILL.md` and its `scripts/` into `plugins/caveman/skills/caveman-compress/`.
|
||||
3. Copies `agents/cavecrew-*.md` into `plugins/caveman/agents/`.
|
||||
4. Rebuilds `dist/caveman.skill` (ZIP of `skills/caveman/`) for the release artifact.
|
||||
5. Commits and pushes with `[skip ci]` to avoid loops.
|
||||
|
||||
CI bot commits as `github-actions[bot]`. After PR merge, wait for workflow before declaring release complete.
|
||||
|
||||
The old steps that mirrored SKILL.md and rules into root dotdirs (`.cursor/`, `.windsurf/`, `.clinerules/`, `.github/copilot-instructions.md`) are gone — those mirrors no longer exist. The old `caveman-compress/` → `skills/compress/` rename-on-sync is also gone now that compress lives at `skills/caveman-compress/`.
|
||||
|
||||
---
|
||||
|
||||
## Hook system (Claude Code)
|
||||
|
||||
Three hooks in `src/hooks/` plus a `caveman-config.js` shared module and a `package.json` CommonJS marker. Communicate via flag file at `$CLAUDE_CONFIG_DIR/.caveman-active` (falls back to `~/.claude/.caveman-active`).
|
||||
|
||||
```
|
||||
SessionStart hook ──writes "full"──▶ $CLAUDE_CONFIG_DIR/.caveman-active ◀──writes mode── UserPromptSubmit hook
|
||||
│
|
||||
reads
|
||||
▼
|
||||
caveman-statusline.sh
|
||||
[CAVEMAN] / [CAVEMAN:ULTRA] / ...
|
||||
```
|
||||
|
||||
`src/hooks/package.json` pins the directory to `{"type": "commonjs"}` so the `.js` hooks resolve as CJS even when an ancestor `package.json` (e.g. `~/.claude/package.json` from another plugin) declares `"type": "module"`. Without this, `require()` blows up with `ReferenceError: require is not defined in ES module scope`.
|
||||
|
||||
All hooks honor `CLAUDE_CONFIG_DIR` for non-default Claude Code config locations.
|
||||
|
||||
### `src/hooks/caveman-config.js` — shared module
|
||||
|
||||
Exports:
|
||||
- `getDefaultMode()` — resolves default mode in order: `CAVEMAN_DEFAULT_MODE` env var → repo-local config (`<cwd>/.caveman/config.json` or `<cwd>/.caveman.json`, walking up to the filesystem root) → user config (`$XDG_CONFIG_HOME/caveman/config.json` / `~/.config/caveman/config.json` / `%APPDATA%\caveman\config.json`) → `'full'`. The env var short-circuits before any cwd walk. Repo-local config lets a team check in a per-project default without polluting every contributor's env or user config.
|
||||
- `findRepoConfigPath(start)` — walks up from `start` (default `process.cwd()`) looking for the first `.caveman/config.json` or `.caveman.json`. Bounded to 64 ancestors. Refuses symlinked files (symmetric with `safeWriteFlag` / `readFlag`).
|
||||
- `safeWriteFlag(flagPath, content)` — symlink-safe flag write. Refuses if flag target or its immediate parent is a symlink. Opens with `O_NOFOLLOW` where supported. Atomic temp + rename. Creates with `0600`. Protects against local attackers replacing the predictable flag path with a symlink to clobber files writable by the user. Used by both write hooks. Silent-fails on all filesystem errors.
|
||||
|
||||
### `src/hooks/caveman-activate.js` — SessionStart hook
|
||||
|
||||
Runs once per Claude Code session start. Three things:
|
||||
1. Writes the active mode to `$CLAUDE_CONFIG_DIR/.caveman-active` via `safeWriteFlag` (creates if missing). Branches on the hook payload's `source` field (#691): `startup` resets to the configured default; `resume`/`clear`/`compact` re-fires preserve a valid existing flag so mid-session `/caveman <level>` switches survive.
|
||||
2. Emits caveman ruleset as hidden stdout — Claude Code injects SessionStart hook stdout as system context, invisible to user
|
||||
3. Checks `settings.json` for statusline config; if missing, appends nudge to offer setup — once per install, gated by a `.caveman-nudge-shown` marker file (#661)
|
||||
|
||||
Silent-fails on all filesystem errors — never blocks session start.
|
||||
|
||||
### `src/hooks/caveman-mode-tracker.js` — UserPromptSubmit hook
|
||||
|
||||
Reads JSON from stdin. Three responsibilities:
|
||||
|
||||
**1. Slash-command activation.** If prompt starts with `/caveman`, writes mode to flag file via `safeWriteFlag`:
|
||||
- `/caveman` → configured default (see `caveman-config.js`, defaults to `full`)
|
||||
- `/caveman lite` → `lite`
|
||||
- `/caveman ultra` → `ultra`
|
||||
- `/caveman wenyan` or `/caveman wenyan-full` → `wenyan` (alias) / `wenyan-full`
|
||||
- `/caveman wenyan-lite` → `wenyan-lite`
|
||||
- `/caveman wenyan-ultra` → `wenyan-ultra`
|
||||
- `/caveman-commit` → `commit`
|
||||
- `/caveman-review` → `review`
|
||||
- `/caveman-compress` → `compress`
|
||||
|
||||
**2. Natural-language activation/deactivation.** Matches phrases like "activate caveman", "turn on caveman mode", "talk like caveman" and writes the configured default mode. Matches "stop caveman", "disable caveman", "normal mode", "deactivate caveman" etc. and deletes the flag file. README promises these triggers, the hook enforces them.
|
||||
|
||||
**3. Per-turn reinforcement.** When flag is set to a non-independent mode (i.e. not `commit`/`review`/`compress`), emits a small `hookSpecificOutput` JSON reminder so the model keeps caveman style after other plugins inject competing instructions mid-conversation. The full ruleset still comes from SessionStart — this is just an attention anchor.
|
||||
|
||||
### `src/hooks/caveman-statusline.sh` — Statusline badge
|
||||
|
||||
Reads flag file at `$CLAUDE_CONFIG_DIR/.caveman-active`. Outputs colored badge string for Claude Code statusline:
|
||||
- `full` or empty → `[CAVEMAN]` (orange)
|
||||
- anything else → `[CAVEMAN:<MODE_UPPERCASED>]` (orange)
|
||||
|
||||
Then appends the lifetime-savings suffix (`⛏ 12.4k`) read from `$CLAUDE_CONFIG_DIR/.caveman-statusline-suffix` — written by `caveman-stats.js` on every `/caveman-stats` run. **Default on**; users opt out with `CAVEMAN_STATUSLINE_SAVINGS=0`. The suffix file is absent until `/caveman-stats` runs at least once, so fresh installs render no fake number.
|
||||
|
||||
Configured in `settings.json` under `statusLine.command`. PowerShell counterpart at `src/hooks/caveman-statusline.ps1` for Windows. Both scripts symlink-refuse and whitelist-validate the flag/suffix file contents — never echo arbitrary bytes.
|
||||
|
||||
### Hook installation
|
||||
|
||||
**Plugin install** — hooks wired automatically by plugin system.
|
||||
|
||||
**Standalone install** — `cli/install.js` (the unified Node installer) copies hook files into `$CLAUDE_CONFIG_DIR/hooks/` and merges SessionStart + UserPromptSubmit + statusline into `settings.json`. Uses the JSONC-tolerant helpers in `cli/lib/settings.js` so a commented `settings.json` no longer crashes the merge. Defensive `validateHookFields` runs before every write to prevent a single malformed hook from poisoning the entire file (Claude Code Zod silently discards the whole `settings.json` on schema mismatch).
|
||||
|
||||
The `install.sh` / `install.ps1` shims at the repo root delegate to `cli/install.js` via `node` (local clone) or `npx -y github:JuliusBrussee/caveman` (curl|bash). No legacy fallback path remains — earlier `install.sh.legacy` / `install.ps1.legacy` files were removed.
|
||||
|
||||
**Uninstall** — `npx -y github:JuliusBrussee/caveman -- --uninstall` (or `node cli/install.js --uninstall` from a clone). Strips caveman hook entries from `settings.json` via substring marker `caveman`, deletes hook files, and removes the Claude plugin / Gemini extension. Also removes state files from `$CLAUDE_CONFIG_DIR` (`.caveman-active`, `.caveman-active.prev`, `.caveman-mode-log.jsonl`, `.caveman-statusline-suffix`, `.caveman-nudge-shown`); keeps `.caveman-history.jsonl` (lifetime savings data) with a printed note (#635). Skill installs done via `npx skills add` must be removed via the IDE's skill manager (we don't track them).
|
||||
|
||||
---
|
||||
|
||||
## Skill system
|
||||
|
||||
Skills = Markdown files with YAML frontmatter consumed by Claude Code's skill/plugin system and by `npx skills` for other agents.
|
||||
|
||||
Each skill has a human-facing `README.md` alongside the LLM-facing `SKILL.md`. The README explains what the skill does for users browsing GitHub; the SKILL.md is the prompt body the agent loads. Don't merge them — different audiences, different formats.
|
||||
|
||||
### Intensity levels
|
||||
|
||||
Defined in `skills/caveman/SKILL.md`. Six levels: `lite`, `full` (default), `ultra`, `wenyan-lite`, `wenyan-full`, `wenyan-ultra`. Persists until changed or session ends.
|
||||
|
||||
### Auto-clarity rule
|
||||
|
||||
Caveman drops to normal prose for: security warnings, irreversible action confirmations, multi-step sequences where fragment ambiguity risks misread, user confused or repeating question. Resumes after. Defined in skill — preserve in any SKILL.md edit.
|
||||
|
||||
### caveman-compress
|
||||
|
||||
Sub-skill in `skills/caveman-compress/SKILL.md`. Takes file path, compresses prose to caveman style, writes to original path, saves backup at `<filename>.original.md`. Validates headings, code blocks, URLs, file paths, commands preserved. Retries up to 2 times on failure with targeted patches only. Requires Python 3.10+.
|
||||
|
||||
The slash command is `/caveman-compress` everywhere — same name in plugin and standalone install. CI no longer renames the directory on sync (the old `caveman-compress/` → `skills/compress/` sed rename is gone now that the source lives at `skills/caveman-compress/`).
|
||||
|
||||
### caveman-commit / caveman-review
|
||||
|
||||
Independent skills in `skills/caveman-commit/SKILL.md` and `skills/caveman-review/SKILL.md`. Both have own `description` and `name` frontmatter so they load independently. caveman-commit: Conventional Commits, ≤50 char subject. caveman-review: one-line comments in `L<line>: <severity> <problem>. <fix>.` format.
|
||||
|
||||
---
|
||||
|
||||
## Agent distribution
|
||||
|
||||
How caveman reaches each agent type:
|
||||
|
||||
| Agent | Mechanism | Auto-activates? |
|
||||
|-------|-----------|----------------|
|
||||
| Claude Code | Plugin (hooks + skills) or standalone hooks | Yes — SessionStart hook injects rules |
|
||||
| Codex | Plugin in `plugins/caveman/` plus repo `.codex/hooks.json` and `.codex/config.toml` | Yes on macOS/Linux — SessionStart hook |
|
||||
| Gemini CLI | Extension with `GEMINI.md` context file | Yes — context file loads every session |
|
||||
| opencode | Native plugin (`src/plugins/opencode/`) copied into `~/.config/opencode/plugins/caveman/` + `AGENTS.md` ruleset + skills/agents/commands directories. Plugin uses `session.created` and `tui.prompt.append` lifecycle hooks. No statusline (opencode TUI exposes no plugin-writable badge). | Yes — `session.created` writes flag, `AGENTS.md` carries always-on ruleset |
|
||||
| OpenClaw | Workspace skill at `~/.openclaw/workspace/skills/caveman/SKILL.md` (frontmatter merged with `version` + `always: true`) plus a marker-fenced bootstrap block in `~/.openclaw/workspace/SOUL.md`. Both writes go through `cli/lib/openclaw.js`; workspace path is overridable via `OPENCLAW_WORKSPACE`. | Yes — SOUL.md is auto-injected each turn under "Project Context" (subject to OpenClaw's 12K-per-file / 60K-total bootstrap caps) |
|
||||
| Cursor | `npx skills add ... -a cursor` (default via `--only cursor`) writes the upstream skill profile; per-repo `.cursor/rules/caveman.mdc` via `--with-init` (calls `src/tools/caveman-init.js`) | Yes — always-on rule |
|
||||
| Windsurf | `npx skills add ... -a windsurf` (default via `--only windsurf`); per-repo `.windsurf/rules/caveman.md` via `--with-init` | Yes — always-on rule |
|
||||
| Cline | `npx skills add ... -a cline` (default via `--only cline`); per-repo `.clinerules/caveman.md` via `--with-init` | Yes — Cline auto-discovers `.clinerules/` |
|
||||
| Copilot | `npx skills add ... -a github-copilot` (soft probe — pass `--only copilot`); per-repo `.github/copilot-instructions.md` + `AGENTS.md` via `--with-init` | Yes — repo-wide instructions |
|
||||
| Others (Junie, Trae, Warp, Tabnine, Mistral, Qwen, Devin, Droid, ForgeCode, Bob, Crush, iFlow, OpenHands, Qoder, Rovo Dev, Replit, Antigravity, …) | `npx skills add JuliusBrussee/caveman -a <profile>` | No — user must say `/caveman` each session |
|
||||
|
||||
opencode reaches Tier 1 minus the statusline (opencode's TUI has no plugin-writable badge). Mode flag lives at `~/.config/opencode/.caveman-active` for any external tooling that wants to surface it.
|
||||
|
||||
For agents without hook systems, the always-on snippet lives in `INSTALL.md`'s "Want it always on?" section — keep current with `src/rules/caveman-activate.md`.
|
||||
|
||||
**Adding a new agent.** Edit the `PROVIDERS` array in `cli/install.js` — single source of truth, no more bash/PS1 dual-source drift. Each entry has `id`, `label`, `mech`, `detect` (clause spec like `command:foo||dir:$HOME/x`), optional `profile` (vercel-labs/skills slug), optional `soft: true` (config-dir-only detection).
|
||||
|
||||
1. The profile slug must exist in upstream [vercel-labs/skills](https://github.com/vercel-labs/skills). Verify against the README before merging — wrong slugs cause `npx skills add` to fail at runtime, not at install-script load.
|
||||
2. Run `node cli/install.js --list` to confirm the new row renders correctly.
|
||||
3. Soft probes (config-dir-only) are fine but tag them with `soft: true`. They render with `(soft)` in `--list` so users know detection is best-effort.
|
||||
|
||||
---
|
||||
|
||||
## Evals
|
||||
|
||||
`evals/` has three-arm harness:
|
||||
- `__baseline__` — no system prompt
|
||||
- `__terse__` — `Answer concisely.`
|
||||
- `<skill>` — `Answer concisely.\n\n{SKILL.md}`
|
||||
|
||||
Honest delta = **skill vs terse**, not skill vs baseline. Baseline comparison conflates skill with generic terseness — that cheating. Harness designed to prevent this.
|
||||
|
||||
`llm_run.py` calls `claude -p --system-prompt ...` per (prompt, arm), saves to `evals/snapshots/results.json`. `measure.py` reads snapshot offline with tiktoken (OpenAI BPE — approximates Claude tokenizer, ratios meaningful, absolute numbers approximate).
|
||||
|
||||
Add skill: drop `skills/<name>/SKILL.md`. Harness auto-discovers. Add prompt: append line to `evals/prompts/en.txt`.
|
||||
|
||||
Snapshots committed to git. CI reads without API calls. Only regenerate when SKILL.md or prompts change.
|
||||
|
||||
---
|
||||
|
||||
## Benchmarks
|
||||
|
||||
`benchmarks/` runs real prompts through Claude API (not Claude Code CLI), records raw token counts. Results committed as JSON in `benchmarks/results/`. Benchmark table in README generated from results — update when regenerating.
|
||||
|
||||
To reproduce: `uv run python benchmarks/run.py` (needs `ANTHROPIC_API_KEY` in `.env.local`).
|
||||
|
||||
---
|
||||
|
||||
## Key rules for agents working here
|
||||
|
||||
- Edit `skills/<name>/SKILL.md` for behavior changes. Never edit synced copies under `plugins/caveman/skills/`.
|
||||
- Edit `src/rules/caveman-activate.md` for auto-activation rule changes. Never edit any per-agent rule copy a user has on their machine.
|
||||
- Edit `src/rules/caveman-openclaw-bootstrap.md` for the OpenClaw SOUL.md bootstrap snippet. Keep the `<!-- caveman-begin -->` / `<!-- caveman-end -->` markers and the `Respond terse like smart caveman` sentinel — `cli/lib/openclaw.js` keys idempotency off both. If you change the embedded fallback in `cli/lib/openclaw.js`, keep it byte-equivalent to the file.
|
||||
- Per-skill human docs live in `skills/<name>/README.md`. The LLM-facing body is in `SKILL.md`. Don't merge them — different audiences.
|
||||
- Build artifacts go in `dist/`. Never check files into `dist/` manually — CI rebuilds them on push, and `dist/` is gitignored.
|
||||
- README most important file for user-facing impact. Optimize for non-technical readers. Preserve caveman voice.
|
||||
- `INSTALL.md` is the per-agent install reference. Keep the install table in `README.md` short and link out to `INSTALL.md` for the full matrix.
|
||||
- Benchmark and eval numbers must be real. Never fabricate or estimate.
|
||||
- CI workflow commits back to main after merge. Account for when checking branch state.
|
||||
- Hook files must silent-fail on all filesystem errors. Never let hook crash block session start.
|
||||
- Any new flag file write must go through `safeWriteFlag()` in `caveman-config.js`. Direct `fs.writeFileSync` on predictable user-owned paths reopens the symlink-clobber attack surface.
|
||||
- Hooks must respect `CLAUDE_CONFIG_DIR` env var, not hardcode `~/.claude`. Same for `cli/install.js` / statusline scripts.
|
||||
- `cli/install.js` is the only installer source. `install.sh` / `install.ps1` at repo root are 30-line shims that delegate to it. Never re-add per-OS install logic to the shims — that's how we got the Windows quoting bug (#249).
|
||||
- Any settings.json read in installer or hooks must go through `cli/lib/settings.js` `readSettings()` so JSONC comments don't crash the merge. Any settings.json write must run through `validateHookFields()` first.
|
||||
@@ -0,0 +1,55 @@
|
||||
# Contributor Covenant Code of Conduct
|
||||
|
||||
## Our Pledge
|
||||
|
||||
We as members, contributors, and leaders pledge to make participation in our
|
||||
community a harassment-free experience for everyone, regardless of age, body
|
||||
size, visible or invisible disability, ethnicity, sex characteristics, gender
|
||||
identity and expression, level of experience, education, socio-economic status,
|
||||
nationality, personal appearance, race, caste, color, religion, or sexual
|
||||
identity and orientation.
|
||||
|
||||
We pledge to act and interact in ways that contribute to an open, welcoming,
|
||||
diverse, inclusive, and healthy community.
|
||||
|
||||
## Our Standards
|
||||
|
||||
Examples of behavior that contributes to a positive environment for our
|
||||
community include:
|
||||
|
||||
* Demonstrating empathy and kindness toward other people
|
||||
* Being respectful of differing opinions, viewpoints, and experiences
|
||||
* Giving and gracefully accepting constructive feedback
|
||||
* Accepting responsibility and apologizing to those affected by our mistakes,
|
||||
and learning from the experience
|
||||
* Focusing on what is best for the overall community, not just for us as
|
||||
individuals
|
||||
|
||||
Examples of unacceptable behavior include:
|
||||
|
||||
* The use of sexualized language or imagery, and unwelcome sexual attention or
|
||||
advances
|
||||
* Trolling, insulting or derogatory comments, and personal or political attacks
|
||||
* Public or private harassment
|
||||
* Publishing others' private information, such as a physical or email
|
||||
address, without their explicit permission
|
||||
* Other conduct which could reasonably be considered inappropriate in a
|
||||
professional setting
|
||||
|
||||
## Enforcement Responsibilities
|
||||
|
||||
Community leaders are responsible for clarifying and enforcing our standards of
|
||||
acceptable behavior and will take appropriate and fair corrective action in
|
||||
response to any behavior that they deem inappropriate, threatening, offensive,
|
||||
or harmful.
|
||||
|
||||
## Scope
|
||||
|
||||
This Code of Conduct applies within all community spaces, and also applies when
|
||||
an individual is officially representing the community in public spaces.
|
||||
|
||||
## Reporting & Contact
|
||||
|
||||
Instances of abusive, harassing, or otherwise unacceptable behavior may be
|
||||
reported to the repository owners. All complaints will be reviewed and investigated
|
||||
promptly and fairly.
|
||||
@@ -1,20 +1,192 @@
|
||||
# Contributing
|
||||
# Contributing to caveman
|
||||
|
||||
Improvements to the SKILL.md prompt are welcome — open a PR with before/after examples showing the change.
|
||||
Thanks for considering a contribution. Caveman is a multi-agent skill that
|
||||
makes 30+ AI coding agents talk in compressed caveman-style prose. Most
|
||||
contributions fall into one of three buckets:
|
||||
|
||||
## How
|
||||
1. **Editing skill prose** — change how caveman speaks, what intensity levels do, what slash commands trigger.
|
||||
2. **Adding a new agent** — wire a fresh editor/CLI/IDE into the unified installer.
|
||||
3. **Fixing the hooks or installer** — Claude Code hooks, the Node installer, the per-repo init script.
|
||||
|
||||
1. Fork repo
|
||||
2. Edit `skills/caveman/SKILL.md` — this is the only copy you need to touch
|
||||
3. Open PR with:
|
||||
- **Before:** what caveman say now
|
||||
- **After:** what caveman say with change
|
||||
- One sentence why change better
|
||||
Caveman like simple. Small focused PR > big rewrite.
|
||||
|
||||
> **Note:** `caveman/SKILL.md`, `plugins/caveman/skills/caveman/SKILL.md`, and `caveman.skill` are auto-synced by CI after merge. Do not edit them directly.
|
||||
---
|
||||
|
||||
Small focused change > big rewrite. Caveman like simple.
|
||||
## Quick orientation
|
||||
|
||||
The repo distributes one skill (caveman) plus a handful of sub-skills
|
||||
(caveman-commit, caveman-review, caveman-compress, cavecrew-*) to many
|
||||
agents through different distribution mechanisms (Claude Code plugin, Codex
|
||||
plugin, Gemini extension, Cursor/Windsurf/Cline rule files, `npx skills` for
|
||||
the long tail). A single Node installer at `cli/install.js` detects which
|
||||
agents are on the user's machine and installs the right thing for each.
|
||||
|
||||
Sources of truth live at the **top level** of the repo. Agent-specific
|
||||
copies live under `plugins/caveman/` and similar mirror dirs — those are
|
||||
**rebuilt by CI** and edits there are reverted.
|
||||
|
||||
---
|
||||
|
||||
## What to edit (sources of truth)
|
||||
|
||||
| I want to change... | Edit this file |
|
||||
|---|---|
|
||||
| Caveman behavior (intensity levels, voice, rules) | `skills/caveman/SKILL.md` |
|
||||
| Caveman commit-message format | `skills/caveman-commit/SKILL.md` |
|
||||
| Caveman code-review format | `skills/caveman-review/SKILL.md` |
|
||||
| Caveman compress logic | `skills/caveman-compress/SKILL.md` and `skills/caveman-compress/scripts/` |
|
||||
| Caveman quick-reference card | `skills/caveman-help/SKILL.md` |
|
||||
| Cavecrew decision guide (when to delegate to subagents) | `skills/cavecrew/SKILL.md` |
|
||||
| cavecrew subagent definitions | `agents/cavecrew-investigator.md`, `agents/cavecrew-builder.md`, `agents/cavecrew-reviewer.md` |
|
||||
| Auto-activation rule body (Cursor/Windsurf/Cline/Copilot) | `src/rules/caveman-activate.md` |
|
||||
| Add support for a new agent | `cli/install.js` (PROVIDERS array) |
|
||||
| Per-repo init script (drops rule files into a user's repo) | `src/tools/caveman-init.js` |
|
||||
| Claude Code hooks | `src/hooks/caveman-activate.js`, `src/hooks/caveman-mode-tracker.js`, `src/hooks/caveman-config.js`, `src/hooks/caveman-statusline.sh`, `src/hooks/caveman-statusline.ps1` |
|
||||
| Settings.json read/write helpers | `cli/lib/settings.js` |
|
||||
| MCP shrink server | `src/mcp-servers/caveman-shrink/` |
|
||||
|
||||
That's it. Every other markdown file with `SKILL.md` in the path is a copy.
|
||||
|
||||
---
|
||||
|
||||
## What NOT to edit (CI-generated mirrors)
|
||||
|
||||
Edits to these files are wiped by the next CI run. The
|
||||
`.github/workflows/sync-skill.yml` job rebuilds them from the sources above
|
||||
on every push to `main`.
|
||||
|
||||
| Path | Rebuilt from |
|
||||
|------|--------------|
|
||||
| `plugins/caveman/skills/caveman/SKILL.md` | `skills/caveman/SKILL.md` |
|
||||
| `plugins/caveman/skills/caveman-compress/{SKILL.md, scripts/}` | `skills/caveman-compress/{SKILL.md, scripts/}` |
|
||||
| `plugins/caveman/skills/cavecrew/SKILL.md` | `skills/cavecrew/SKILL.md` |
|
||||
| `plugins/caveman/agents/cavecrew-*.md` | `agents/cavecrew-*.md` |
|
||||
| `dist/caveman.skill` | ZIP of `skills/caveman/` (gitignored; rebuilt by CI on each push to `main`) |
|
||||
|
||||
`caveman-commit`, `caveman-review`, `caveman-help`, and `caveman-stats` are **not** mirrored under `plugins/caveman/skills/` by CI. Claude Code reaches them through the standalone hook + skill install path and `npx skills` carries them to other agents. If you see `plugins/caveman/skills/caveman-stats/` checked in, treat it as a legacy hand-committed copy — the workflow in `.github/workflows/sync-skill.yml` does not touch it.
|
||||
|
||||
When in doubt: if the file lives under `plugins/`, `dist/`, or any agent
|
||||
dotdir mirror, it's a build artifact. Edit the top-level source instead.
|
||||
|
||||
---
|
||||
|
||||
## Adding a new agent
|
||||
|
||||
The unified Node installer at `cli/install.js` is the **single source of
|
||||
truth** for the supported-agent list. The README and `INSTALL.md` install
|
||||
tables mirror it by hand — bash and PowerShell shims at the repo root just
|
||||
delegate to it.
|
||||
|
||||
1. Confirm the agent has a distribution path. Either:
|
||||
- it has a profile slug in upstream [vercel-labs/skills](https://github.com/vercel-labs/skills) (most common), or
|
||||
- it has a native plugin / extension / rule-file mechanism we can target.
|
||||
2. Append a row to the `PROVIDERS` array in `cli/install.js`. Each row needs:
|
||||
- `id` — short kebab-case identifier (e.g. `windsurf`)
|
||||
- `label` — human display name (e.g. `Windsurf`)
|
||||
- `mech` — distribution mechanism (`plugin`, `extension`, `rules-file`, `skills-cli`, …)
|
||||
- `detect` — clause spec like `command:foo||dir:$HOME/x` describing how to detect the agent
|
||||
- `profile` — the vercel-labs/skills slug, if applicable
|
||||
- `soft: true` — set when detection is config-dir-only (best-effort)
|
||||
3. Run `node cli/install.js --list` and confirm the new row renders correctly. Soft probes should show as `(soft)`.
|
||||
4. Add a row to the install tables in `README.md` and `INSTALL.md`.
|
||||
5. No CI changes needed — the workflow re-reads `cli/install.js` automatically.
|
||||
|
||||
Bad slug? `npx skills add` fails at install **runtime**, not at install-script
|
||||
load. Always verify the slug against the vercel-labs/skills README before
|
||||
merging.
|
||||
|
||||
---
|
||||
|
||||
## Adding a new skill
|
||||
|
||||
1. Create `skills/<name>/SKILL.md` with frontmatter:
|
||||
```yaml
|
||||
---
|
||||
name: <name>
|
||||
description: <one sentence, present tense>
|
||||
---
|
||||
```
|
||||
2. Create `skills/<name>/README.md` — human-facing summary, install hint, example.
|
||||
3. Add `skills/<name>/scripts/` if the skill ships helpers (Python or Node).
|
||||
4. If the skill should be in the Claude Code plugin, add a sync step to `.github/workflows/sync-skill.yml` so CI mirrors it into `plugins/caveman/skills/<name>/`.
|
||||
5. If it's user-invocable as a slash command, add a row to the slash-command table in `README.md` and `INSTALL.md`.
|
||||
6. Add an eval prompt to `evals/prompts/en.txt` if you want the eval harness to score it.
|
||||
|
||||
---
|
||||
|
||||
## Running tests
|
||||
|
||||
```bash
|
||||
# Installer unit + e2e tests (Node)
|
||||
npm test
|
||||
|
||||
# Compress-skill safety tests (Python)
|
||||
python3 -m unittest tests.test_compress_safety
|
||||
|
||||
# Per-repo init tests
|
||||
node tests/test_caveman_init.js
|
||||
|
||||
# Flag-file symlink-safety tests
|
||||
node tests/test_symlink_flag.js
|
||||
```
|
||||
|
||||
CI runs all of the above on every PR. If any test depends on a network or
|
||||
external SDK, it must skip cleanly when the dependency is missing — never
|
||||
gate the whole suite on optional creds.
|
||||
|
||||
---
|
||||
|
||||
## Running benchmarks and evals
|
||||
|
||||
Benchmarks hit the real Claude API and record raw token counts:
|
||||
|
||||
```bash
|
||||
uv run python benchmarks/run.py # needs ANTHROPIC_API_KEY in .env.local
|
||||
```
|
||||
|
||||
Evals are a three-arm offline harness (`__baseline__`, `__terse__`, each skill):
|
||||
|
||||
```bash
|
||||
python evals/llm_run.py # regenerates evals/snapshots/results.json
|
||||
python evals/measure.py # reads snapshot, prints token deltas
|
||||
```
|
||||
|
||||
Snapshots are committed to git. Only regenerate when a `SKILL.md` or
|
||||
`evals/prompts/en.txt` changes. Numbers in `README.md` and any docs come from
|
||||
real runs — never invent or round.
|
||||
|
||||
---
|
||||
|
||||
## Pull-request guidelines
|
||||
|
||||
- **Conventional Commits** for the commit subject. See `skills/caveman-commit/SKILL.md` for the format we use here.
|
||||
- **One concern per PR.** A README copy-edit and an installer fix go in separate PRs.
|
||||
- **Update `package.json` `files`** if you add a new top-level directory the installer needs to ship to npm. Files outside that array don't get published.
|
||||
- **Show before/after** for prose changes to any `SKILL.md`. One sentence on why the new wording is better.
|
||||
- **Mention the CI sync.** If you edited a source-of-truth file, note it: "CI will resync `plugins/caveman/skills/...` on merge."
|
||||
|
||||
PR descriptions don't need to be long. Caveman style fine. Just say what change, why.
|
||||
|
||||
---
|
||||
|
||||
## Code style
|
||||
|
||||
A handful of invariants that have bitten us before. Keep them.
|
||||
|
||||
- **Hooks must silent-fail on filesystem errors.** A `try/catch` that swallows the error is correct here. A hook that throws blocks Claude Code session start — that's user-facing breakage. See existing patterns in `src/hooks/caveman-activate.js`.
|
||||
- **Settings.json reads and writes go through `cli/lib/settings.js`.** It tolerates JSONC comments. Direct `JSON.parse` on a user's `settings.json` will crash on a single `// comment`.
|
||||
- **Validate hook entries before writing.** Use `validateHookFields()` in `cli/lib/settings.js`. Claude Code's Zod schema silently discards the **entire** `settings.json` on a single bad hook entry — one malformed write poisons the user's whole config.
|
||||
- **Symlink-safe flag writes via `safeWriteFlag()`** in `src/hooks/caveman-config.js`. The flag file lives at a predictable path under `$CLAUDE_CONFIG_DIR/`; without `O_NOFOLLOW` and a parent-symlink check, a local attacker can clobber any file the user can write.
|
||||
- **Honor `CLAUDE_CONFIG_DIR`.** Hooks, the installer, and the statusline scripts must respect it — never hardcode `~/.claude`.
|
||||
- **`install.sh` and `install.ps1` at the repo root are 30-line shims** that delegate to `cli/install.js`. Don't re-add per-OS install logic to them. Quoting bugs that way lie.
|
||||
|
||||
---
|
||||
|
||||
## Ideas
|
||||
|
||||
See [issues labeled `good first issue`](../../issues?q=label%3A%22good+first+issue%22) for starter tasks.
|
||||
See [issues labeled `good first issue`](../../issues?q=label%3A%22good+first+issue%22)
|
||||
for starter tasks. Or grep `TODO` / `FIXME` in `src/hooks/`, `cli/`, `src/tools/` —
|
||||
each one is a real lead.
|
||||
|
||||
Caveman like contribution. You bring rock, caveman put rock in pile. Pile
|
||||
get bigger. Brain still big.
|
||||
|
||||
@@ -0,0 +1,4 @@
|
||||
@./skills/caveman/SKILL.md
|
||||
@./skills/caveman-commit/SKILL.md
|
||||
@./skills/caveman-review/SKILL.md
|
||||
@./skills/caveman-compress/SKILL.md
|
||||
@@ -0,0 +1,261 @@
|
||||
# Install caveman
|
||||
|
||||
One install. Works for every AI coding agent on your machine.
|
||||
|
||||
If just want it to work, run the one-liner. If want to know what gets touched, scroll down.
|
||||
|
||||
## One-liner
|
||||
|
||||
**macOS / Linux / WSL / Git Bash**
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash
|
||||
```
|
||||
|
||||
**Windows (PowerShell 5.1+)**
|
||||
|
||||
```powershell
|
||||
irm https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.ps1 | iex
|
||||
```
|
||||
|
||||
> Piping a script straight into a shell runs it sight-unseen. If you'd rather read it first, download then run: `curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh -o install.sh` (review it) `&& bash install.sh`. The installer downloads hook files from a pinned release tag and verifies them against a committed SHA-256 manifest before writing.
|
||||
|
||||
What it does:
|
||||
|
||||
- Auto-detects every supported agent installed on your machine (Claude Code, Cursor, Codex, etc.).
|
||||
- For each one, runs that agent's native install path (plugin / extension / rule file / `npx skills add`).
|
||||
- Wires Claude Code hooks and statusline badge on top. (`caveman-shrink` MCP middleware is opt-in via `--with-mcp-shrink` — see flag table below.)
|
||||
- Skips anything you don't have. Safe to re-run. ~30 seconds end-to-end.
|
||||
|
||||
Want to preview before installing? Use `--dry-run`:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash -s -- --dry-run
|
||||
```
|
||||
|
||||
## Per-agent install
|
||||
|
||||
If you want to install for one agent (or want to know exactly what command runs under the hood), use the table below. Every row also works as `--only <id>` to the unified installer.
|
||||
|
||||
| Agent | Install command | Auto-activates? |
|
||||
|---|---|:-:|
|
||||
| **Claude Code** | `claude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@caveman` | Yes |
|
||||
| **Gemini CLI** | `gemini extensions install https://github.com/JuliusBrussee/caveman --consent` | Yes |
|
||||
| **opencode** | `node cli/install.js --only opencode` *(or `npx -y github:JuliusBrussee/caveman -- --only opencode`)* | Yes (plugin + AGENTS.md) |
|
||||
| **OpenClaw** | `npx -y github:JuliusBrussee/caveman -- --only openclaw` | Yes (workspace skill + SOUL.md) |
|
||||
| **Hermes Agent** | `npx -y github:JuliusBrussee/caveman -- --only hermes` *(or `node cli/install.js --only hermes` from a clone)* | Yes (native skills, enabled on load) |
|
||||
| **Codex CLI** | `npx skills add JuliusBrussee/caveman -a codex` | Per-session: `/caveman` |
|
||||
| **Cursor** | `npx skills add JuliusBrussee/caveman -a cursor` | Per-session by default; `--with-init` for an always-on rule file |
|
||||
| **Windsurf** | `npx skills add JuliusBrussee/caveman -a windsurf` | Per-session by default; `--with-init` for an always-on rule file |
|
||||
| **Cline** | `npx skills add JuliusBrussee/caveman -a cline` | Per-session by default; `--with-init` for an always-on rule file |
|
||||
| **GitHub Copilot** *(soft probe)* | `npx -y github:JuliusBrussee/caveman -- --only copilot --with-init` | Repo-wide instructions via `--with-init` |
|
||||
| **Continue** | `npx skills add JuliusBrussee/caveman -a continue` | No — say `/caveman` |
|
||||
| **Kilo Code** | `npx skills add JuliusBrussee/caveman -a kilo` | No |
|
||||
| **Roo Code** | `npx skills add JuliusBrussee/caveman -a roo` | No |
|
||||
| **Augment Code** | `npx skills add JuliusBrussee/caveman -a augment` | No |
|
||||
| **Aider Desk** | `npx skills add JuliusBrussee/caveman -a aider-desk` | No |
|
||||
| **Sourcegraph Amp** | `npx skills add JuliusBrussee/caveman -a amp` | No |
|
||||
| **IBM Bob** | `npx skills add JuliusBrussee/caveman -a bob` | No |
|
||||
| **Crush** | `npx skills add JuliusBrussee/caveman -a crush` | No |
|
||||
| **Devin (terminal)** | `npx skills add JuliusBrussee/caveman -a devin` | No |
|
||||
| **Droid (Factory)** | `npx skills add JuliusBrussee/caveman -a droid` | No |
|
||||
| **ForgeCode** | `npx skills add JuliusBrussee/caveman -a forgecode` | No |
|
||||
| **Block Goose** | `npx skills add JuliusBrussee/caveman -a goose` | No |
|
||||
| **iFlow CLI** | `npx skills add JuliusBrussee/caveman -a iflow-cli` | No |
|
||||
| **Kiro CLI** | `npx skills add JuliusBrussee/caveman -a kiro-cli` | No |
|
||||
| **Mistral Vibe** | `npx skills add JuliusBrussee/caveman -a mistral-vibe` | No |
|
||||
| **OpenHands** | `npx skills add JuliusBrussee/caveman -a openhands` | No |
|
||||
| **Qwen Code** | `npx skills add JuliusBrussee/caveman -a qwen-code` | No |
|
||||
| **Atlassian Rovo Dev** | `npx skills add JuliusBrussee/caveman -a rovodev` | No |
|
||||
| **Tabnine CLI** | `npx skills add JuliusBrussee/caveman -a tabnine-cli` | No |
|
||||
| **Trae** | `npx skills add JuliusBrussee/caveman -a trae` | No |
|
||||
| **Warp** | `npx skills add JuliusBrussee/caveman -a warp` | No |
|
||||
| **Replit Agent** | `npx skills add JuliusBrussee/caveman -a replit` | No |
|
||||
| **JetBrains Junie** *(soft probe)* | `npx skills add JuliusBrussee/caveman -a junie` | No |
|
||||
| **Qoder** *(soft probe)* | `npx skills add JuliusBrussee/caveman -a qoder` | No |
|
||||
| **Google Antigravity** *(soft probe)* | `npx skills add JuliusBrussee/caveman -a antigravity` | No |
|
||||
|
||||
"Soft probe" = installer won't auto-detect these without `--only <id>` because there's no reliable always-on signal (Copilot subscription state is auth-gated; the others have no CLI / config-dir-only). Pass the flag when you want them.
|
||||
|
||||
For "auto-activates? No" agents, type `/caveman` once per session (or use natural-language triggers like "talk like caveman", "caveman mode").
|
||||
|
||||
**Finding a profile slug for `npx skills add ... -a <profile>`?** Either read the table above, or print the live matrix from the installer:
|
||||
|
||||
```bash
|
||||
# Either of these works (install.sh / install.ps1 are thin shims that
|
||||
# forward all flags to cli/install.js):
|
||||
bash install.sh --list # macOS / Linux / WSL, from a local clone
|
||||
pwsh install.ps1 --list # Windows / PowerShell, from a local clone
|
||||
node cli/install.js --list # any platform, from a local clone
|
||||
npx -y github:JuliusBrussee/caveman -- --list # no clone needed
|
||||
```
|
||||
|
||||
Each row prints the agent id, profile slug (where applicable), and whether it was auto-detected on your machine. Full agent matrix (with detection rules) is also defined in `cli/install.js` under the `PROVIDERS` array.
|
||||
|
||||
## Manual install (no `curl | bash`)
|
||||
|
||||
If you'd rather see exactly what runs:
|
||||
|
||||
```bash
|
||||
# Clone the repo
|
||||
git clone https://github.com/JuliusBrussee/caveman.git
|
||||
cd caveman
|
||||
|
||||
# Preview every command the installer would run
|
||||
node cli/install.js --dry-run --all
|
||||
|
||||
# Inspect the agent matrix
|
||||
node cli/install.js --list
|
||||
|
||||
# Install for everything detected
|
||||
node cli/install.js --all
|
||||
```
|
||||
|
||||
Useful flags:
|
||||
|
||||
| Flag | What |
|
||||
|---|---|
|
||||
| `--all` | Plugin + hooks + statusline + per-repo rule files in `$PWD`. (MCP shrink is opt-in — see `--with-mcp-shrink` below.) |
|
||||
| `--minimal` | Plugin / extension only. No hooks, no MCP shrink, no per-repo rules. |
|
||||
| `--only <id>` | One agent only. Repeatable: `--only claude --only cursor`. |
|
||||
| `--dry-run` | Print every command. Write nothing. |
|
||||
| `--with-init` | Drop always-on rule files into the current repo (`.cursor/`, `.windsurf/`, `.clinerules/`, `.github/copilot-instructions.md`, `.opencode/AGENTS.md`, `AGENTS.md`) and, if OpenClaw is on the box, append the bootstrap block to `~/.openclaw/workspace/SOUL.md`. |
|
||||
| `--with-mcp-shrink="<upstream cmd>"` | Register `caveman-shrink` MCP proxy wrapping the given upstream MCP server. **Off by default.** A value is required — caveman-shrink is a proxy and exits immediately without one. Example: `--with-mcp-shrink="npx @modelcontextprotocol/server-filesystem /tmp"`. The value is split on whitespace; for paths-with-spaces, install via `node cli/install.js` from a clone or edit `~/.claude.json` after a stub install. |
|
||||
| `--no-mcp-shrink` | Skip MCP-shrink registration. (Default.) |
|
||||
| `--with-hooks` / `--no-hooks` | Force-on or force-off the Claude Code hook installer. (Default: on.) |
|
||||
| `--skip-skills` | Don't run the npx-skills auto-detect fallback when nothing else matched. |
|
||||
| `--config-dir <path>` | Claude Code config dir for hook files + `settings.json`. **Does NOT scope** `claude plugin install`, `gemini extensions install`, opencode (`XDG_CONFIG_HOME`), or openclaw (`OPENCLAW_WORKSPACE`) — those use their own paths. Default: `$CLAUDE_CONFIG_DIR` or `~/.claude`. `~` is expanded. |
|
||||
| `--non-interactive` | Never prompt; use defaults. (Auto when stdin is not a TTY.) |
|
||||
| `--no-color` | Disable ANSI colors. |
|
||||
| `--list` | Print full agent matrix and exit. |
|
||||
| `--force` | Re-run even if already installed. |
|
||||
| `--uninstall` | Remove everything. See below. |
|
||||
|
||||
## Always-on rules
|
||||
|
||||
For agents without a hook system (Cursor, Windsurf, Cline, Copilot, and friends), the always-on path is a static rule file. Two ways:
|
||||
|
||||
```bash
|
||||
# Drop rule files into the current repo
|
||||
node cli/install.js --with-init
|
||||
|
||||
# Or pull the rule body straight in (manual)
|
||||
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/src/rules/caveman-activate.md \
|
||||
> .cursor/rules/caveman.mdc # or .windsurf/rules/caveman.md, .clinerules/caveman.md, .github/copilot-instructions.md
|
||||
```
|
||||
|
||||
`--with-init` writes the rule into every supported per-agent location it can detect (`.cursor/rules/`, `.windsurf/rules/`, `.clinerules/`, `.github/copilot-instructions.md`, `.opencode/AGENTS.md`, `AGENTS.md`). It also installs the OpenClaw workspace bootstrap (skill folder + SOUL.md marker block) when `~/.openclaw/workspace/` exists. Single source: [`src/rules/caveman-activate.md`](src/rules/caveman-activate.md).
|
||||
|
||||
## Verify
|
||||
|
||||
After install, three quick checks:
|
||||
|
||||
**1. See what got installed.**
|
||||
|
||||
```bash
|
||||
node cli/install.js --list
|
||||
```
|
||||
|
||||
You should see ~30 rows. Detected agents are marked. Anything you wanted but isn't marked → not detected (likely the binary isn't on `PATH`).
|
||||
|
||||
**2. Talk to Claude Code.**
|
||||
|
||||
Open Claude Code, type `/caveman`. Response should be terse fragments — "Got it. Caveman mode on." or similar. Try a real question: "What is closures in JS?" — answer should drop articles and read like grunts.
|
||||
|
||||
**3. Check the flag file.**
|
||||
|
||||
```bash
|
||||
cat "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/.caveman-active"
|
||||
# expected output: full
|
||||
```
|
||||
|
||||
If it's missing or empty, the SessionStart hook didn't fire. See troubleshooting below.
|
||||
|
||||
Statusline should show `[CAVEMAN]` (orange) at the bottom of Claude Code. After your first `/caveman-stats` run it appends a savings counter like `[CAVEMAN] ⛏ 12.4k`.
|
||||
|
||||
## Uninstall
|
||||
|
||||
```bash
|
||||
npx -y github:JuliusBrussee/caveman -- --uninstall
|
||||
```
|
||||
|
||||
What it removes:
|
||||
|
||||
- Caveman hook entries from `$CLAUDE_CONFIG_DIR/settings.json` (default `~/.claude/`; matched by the substring `caveman`).
|
||||
- Hook files in `$CLAUDE_CONFIG_DIR/hooks/` (`caveman-activate.js`, `caveman-mode-tracker.js`, `caveman-stats.js`, `caveman-config.js`, `caveman-statusline.{sh,ps1}`, plus the dir's `package.json` marker).
|
||||
- The Claude Code plugin and the Gemini CLI extension (if installed).
|
||||
- The opencode native plugin (`~/.config/opencode/plugins/caveman/`, the `plugin` and `mcp.caveman-shrink` entries from `opencode.json`, our skill/agent/command files, the caveman block from `AGENTS.md`, and the opencode flag file).
|
||||
- The OpenClaw workspace skill folder and the marker-fenced block from `~/.openclaw/workspace/SOUL.md` (when present).
|
||||
- The `.caveman-active` flag file.
|
||||
|
||||
What it does **not** remove:
|
||||
|
||||
- Skills installed via `npx skills add` — the `skills` CLI manages those. Run `npx skills remove caveman` (or use your IDE's skill manager).
|
||||
- Per-repo rule files written by `--with-init` (`.cursor/rules/`, `.windsurf/rules/`, `.clinerules/`, `.github/copilot-instructions.md`, `.opencode/AGENTS.md`, `AGENTS.md`). Delete by hand if you want.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**"Install script broke. What now?"**
|
||||
|
||||
Open your agent in this repo and say:
|
||||
|
||||
> "Read CLAUDE.md and INSTALL.md. Install caveman for me."
|
||||
|
||||
Agent read repo. Agent run install. Caveman make agent talk less — agent first job is install caveman to talk less. Snake eat tail.
|
||||
|
||||
Still broken? [Open an issue](https://github.com/JuliusBrussee/caveman/issues).
|
||||
|
||||
**"I ran the installer but Claude Code isn't talking caveman."**
|
||||
|
||||
1. Run `node cli/install.js --list` — confirm `claude` is on the detected list. If not, `claude` isn't on `PATH`. Fix that first.
|
||||
2. Open `$CLAUDE_CONFIG_DIR/settings.json` (default `~/.claude/settings.json`) and look for `"hooks"` containing `caveman-activate.js` and `caveman-mode-tracker.js`. If missing, re-run with `--force`.
|
||||
3. Check `$CLAUDE_CONFIG_DIR/.caveman-active` exists with content `full`. If not, the SessionStart hook silent-failed — check `$CLAUDE_CONFIG_DIR/hooks/` for the JS files and try `node $CLAUDE_CONFIG_DIR/hooks/caveman-activate.js < /dev/null` to see if it errors.
|
||||
4. Restart Claude Code. The SessionStart hook only fires on session start, not mid-session.
|
||||
|
||||
**"Hooks failing on Windows."**
|
||||
|
||||
- Use `install.ps1`, not `install.sh`. Git Bash works for the shell version, but the hook side wires PowerShell counterparts (`caveman-statusline.ps1`).
|
||||
- PowerShell 5.1 minimum. Check with `$PSVersionTable.PSVersion`.
|
||||
- If `irm | iex` blocks on execution policy: `Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass` for the install session, then re-run.
|
||||
- Long-running issues: see `docs/install-windows.md` in the repo for manual fallback.
|
||||
|
||||
**"My `settings.json` got mangled."**
|
||||
|
||||
The installer uses a JSONC-tolerant parser (`cli/lib/settings.js`) so comments and trailing commas don't crash the merge. It also runs `validateHookFields()` before every write so a malformed hook can't poison the file. If something still went wrong:
|
||||
|
||||
1. Check for a backup at `$CLAUDE_CONFIG_DIR/settings.json.bak` (installer writes one before any merge).
|
||||
2. If no backup, restore from your shell history or version control.
|
||||
3. File an issue with the broken `settings.json` content (redacted) — that file passing validation but breaking Claude Code is a bug we want to fix.
|
||||
|
||||
**"I'm in a managed env where I can't install hooks."**
|
||||
|
||||
Use the rule-file-only path. Hooks are Claude Code-specific; everything else works via static rule files:
|
||||
|
||||
```bash
|
||||
# Just install for one agent, no Claude hooks
|
||||
node cli/install.js --only cursor
|
||||
|
||||
# Or write rule files into the current repo only (no global state)
|
||||
node cli/install.js --with-init --only cursor --only windsurf
|
||||
```
|
||||
|
||||
This drops `.cursor/rules/caveman.mdc` (and friends) into your repo. No hooks, no global config, nothing outside the repo.
|
||||
|
||||
**"`npx skills add` errored on a profile slug."**
|
||||
|
||||
The profile slug must exist in [vercel-labs/skills](https://github.com/vercel-labs/skills). If a row in the table above 404s, the upstream profile was renamed or removed — open an issue, we'll update.
|
||||
|
||||
## Privacy
|
||||
|
||||
The installer doesn't phone home. It writes to:
|
||||
|
||||
- `$CLAUDE_CONFIG_DIR` (default `~/.claude/`) — hooks, flag file, `settings.json` merge.
|
||||
- Each agent's own config location — Cursor's `.cursor/rules/`, Windsurf's `.windsurf/rules/`, opencode's `~/.config/opencode/`, etc.
|
||||
- Your current working directory (only with `--with-init`) — repo-local rule files.
|
||||
- `~/.openclaw/workspace/` (only with `--only openclaw` or `--with-init` when OpenClaw is detected) — the one `--with-init` side-effect outside the cwd.
|
||||
|
||||
No telemetry. No analytics. Run from a clone or via npx, the installer's own code makes no network calls — files are copied locally. One exception: run detached from any checkout (the rare curl-fallback path), it downloads the hook files from raw.githubusercontent.com pinned to an immutable release tag and verifies each against a SHA-256 manifest before wiring anything. Network requests also happen indirectly through the per-agent CLIs it shells out to — `claude plugin marketplace add`, `claude plugin install`, `gemini extensions install`, `npm view caveman-shrink`, and `npx -y skills add`. Each fetches from its own registry (Anthropic / GitHub / npm). Source: [`cli/install.js`](cli/install.js). After install: zero network calls, ever — full statement in [SECURITY.md](./SECURITY.md#privacy--telemetry).
|
||||
|
||||
---
|
||||
|
||||
Stuck? Open an issue: <https://github.com/JuliusBrussee/caveman/issues>
|
||||
@@ -1,110 +1,164 @@
|
||||
<p align="center">
|
||||
<img src="https://em-content.zobj.net/source/apple/391/rock_1faa8.png" width="120" />
|
||||
<img src="docs/assets/caveman-logo-banner.png" alt="Caveman" width="720">
|
||||
</p>
|
||||
|
||||
<h1 align="center">caveman</h1>
|
||||
|
||||
<p align="center">
|
||||
<strong>why use many token when few do trick</strong>
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://github.com/JuliusBrussee/caveman/stargazers"><img src="https://img.shields.io/github/stars/JuliusBrussee/caveman?style=flat&color=yellow" alt="Stars"></a>
|
||||
<a href="https://github.com/JuliusBrussee/caveman/commits/main"><img src="https://img.shields.io/github/last-commit/JuliusBrussee/caveman?style=flat" alt="Last Commit"></a>
|
||||
<a href="LICENSE"><img src="https://img.shields.io/github/license/JuliusBrussee/caveman?style=flat" alt="License"></a>
|
||||
Make your AI coding agent talk like a caveman.<br>
|
||||
Same answers. <strong>65% fewer output tokens</strong> on prose,<br>
|
||||
<strong>8.5%</strong> on <a href="#independently-measured-jetbrains-86-tasks">long-horizon agentic coding runs</a>. Brain still big. Mouth small.
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="#install">Install</a> •
|
||||
<a href="#benchmarks">Benchmarks</a> •
|
||||
<a href="#before--after">Before/After</a> •
|
||||
<a href="#intensity-levels">Intensity Levels</a> •
|
||||
<a href="#caveman-compress">Compress</a> •
|
||||
<a href="#why">Why</a>
|
||||
<a href="https://trendshift.io/repositories/25391?utm_source=repository-badge&utm_medium=badge&utm_campaign=badge-repository-25391" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/repositories/25391" alt="JuliusBrussee%2Fcaveman | Trendshift" width="250" height="55"/></a>
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://github.com/JuliusBrussee/caveman/stargazers"><img src="https://img.shields.io/github/stars/JuliusBrussee/caveman?style=flat&color=yellow" alt="Stars"></a>
|
||||
<a href="./INSTALL.md"><img src="https://img.shields.io/badge/works_with-30%2B_agents-orange?style=flat" alt="30+ agents"></a>
|
||||
<a href="https://github.com/JuliusBrussee/caveman/commits/main"><img src="https://img.shields.io/github/last-commit/JuliusBrussee/caveman?style=flat" alt="Last commit"></a>
|
||||
<a href="LICENSE"><img src="https://img.shields.io/github/license/JuliusBrussee/caveman?style=flat" alt="License"></a>
|
||||
<a href="https://skills.sh/JuliusBrussee/caveman"><img src="https://skills.sh/b/JuliusBrussee/caveman"></a>
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="#before--after">See it</a> ·
|
||||
<a href="#install">Install</a> ·
|
||||
<a href="#pick-your-grunt">Levels</a> ·
|
||||
<a href="#what-you-get">What you get</a> ·
|
||||
<a href="#benchmarks">Benchmarks</a> ·
|
||||
<a href="#the-whole-cave">Ecosystem</a> ·
|
||||
<a href="#caveman-2">Caveman 2</a>
|
||||
</p>
|
||||
|
||||
---
|
||||
|
||||
A [Claude Code](https://docs.anthropic.com/en/docs/claude-code) skill/plugin and Codex plugin that makes agent talk like caveman — cutting **~75% of output tokens** while keeping full technical accuracy. Plus a companion tool that compresses your memory files to cut **~45% of input tokens** every session.
|
||||
|
||||
Based on the viral observation that caveman-speak dramatically reduces LLM token usage without losing technical substance. So we made it a one-line install.
|
||||
Caveman is a skill/plugin for [Claude Code](https://docs.anthropic.com/en/docs/claude-code), Codex, Gemini, Cursor, Windsurf, Cline, Copilot, and 30+ other agents. Install once. Agent drops the filler and answers in tight caveman-speak, keeping code, commands, and errors byte-for-byte exact. You save output tokens on every reply, forever.
|
||||
|
||||
## Before / After
|
||||
|
||||
<table>
|
||||
<tr>
|
||||
<td width="50%">
|
||||
<th width="50%">🗣️ Normal agent — 69 tokens</th>
|
||||
<th width="50%"><img src="docs/assets/dancing-rock.svg" width="18" height="18" alt=""> Caveman agent — 19 tokens</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<td valign="top">
|
||||
|
||||
### 🗣️ Normal Claude (69 tokens)
|
||||
|
||||
> "The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object."
|
||||
> The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object.
|
||||
|
||||
</td>
|
||||
<td width="50%">
|
||||
<td valign="top">
|
||||
|
||||
### 🪨 Caveman Claude (19 tokens)
|
||||
|
||||
> "New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`."
|
||||
> New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`.
|
||||
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<td valign="top">
|
||||
|
||||
### 🗣️ Normal Claude
|
||||
|
||||
> "Sure! I'd be happy to help you with that. The issue you're experiencing is most likely caused by your authentication middleware not properly validating the token expiry. Let me take a look and suggest a fix."
|
||||
> Sure! I'd be happy to help you with that. The issue you're experiencing is most likely caused by your authentication middleware not properly validating the token expiry. Let me take a look and suggest a fix.
|
||||
|
||||
</td>
|
||||
<td>
|
||||
<td valign="top">
|
||||
|
||||
### 🪨 Caveman Claude
|
||||
|
||||
> "Bug in auth middleware. Token expiry check use `<` not `<=`. Fix:"
|
||||
> Bug in auth middleware. Token expiry check use `<` not `<=`. Fix:
|
||||
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
**Same fix. 75% less word. Brain still big.**
|
||||
Same fix. Third of the words. Nothing technical lost.
|
||||
|
||||
**Sometimes too much caveman. Sometimes not enough:**
|
||||
```
|
||||
┌────────────────────────────────────────────┐
|
||||
│ output tokens saved █████████ 65% │
|
||||
│ input tokens saved ░░░░░░░░░ 0% │
|
||||
│ technical accuracy █████████ 100% │
|
||||
│ vibes █████████ OOG │
|
||||
└────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
<table>
|
||||
<tr>
|
||||
<td width="33%">
|
||||
Caveman no make brain smaller. Caveman make *mouth* smaller. Shrinks what the agent **says**, not what it knows.
|
||||
|
||||
#### 🪶 Lite
|
||||
That 65% is the prose number, measured on replies like the ones above. On a full agentic coding run, where most of the output is code and tool calls, it's [8.5%](#independently-measured-jetbrains-86-tasks). Same skill, different workload — mechanism below.
|
||||
|
||||
> "Your component re-renders because you create a new object reference each render. Inline object props fail shallow comparison every time. Wrap it in `useMemo`."
|
||||
## Install
|
||||
|
||||
</td>
|
||||
<td width="33%">
|
||||
**One command. Finds every agent on your machine. Installs for each.**
|
||||
|
||||
#### 🪨 Full
|
||||
```bash
|
||||
# macOS · Linux · WSL · Git Bash
|
||||
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash
|
||||
```
|
||||
|
||||
> "New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`."
|
||||
```powershell
|
||||
# Windows · PowerShell 5.1+
|
||||
irm https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.ps1 | iex
|
||||
```
|
||||
|
||||
</td>
|
||||
<td width="33%">
|
||||
~30 seconds. Needs Node ≥18. Skips agents you no have. Safe to re-run.
|
||||
|
||||
#### 🔥 Ultra
|
||||
Prefer one agent at a time? Each has its own path:
|
||||
|
||||
> "Inline obj prop → new ref → re-render. `useMemo`."
|
||||
```bash
|
||||
# Claude Code plugin
|
||||
claude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@caveman
|
||||
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
# Gemini CLI extension
|
||||
gemini extensions install https://github.com/JuliusBrussee/caveman --consent
|
||||
|
||||
**Same answer. You pick how many word.**
|
||||
# Cursor / Windsurf / Cline / Codex / 30+ more, via the skills registry
|
||||
npx skills add JuliusBrussee/caveman -a cursor
|
||||
```
|
||||
|
||||
The full per-agent matrix, all flags, dry-run, and uninstall live in **[INSTALL.md](./INSTALL.md)**.
|
||||
|
||||
> [!TIP]
|
||||
> **Turn it on:** type `/caveman` or say *"talk like caveman"*. **Turn it off:** say *"normal mode"*. On Claude Code, Codex, and Gemini it's already on from message one. No command needed.
|
||||
|
||||
**Install broke?** Open your agent in this repo and say: *"Read CLAUDE.md and INSTALL.md, install caveman for me."* Agent read repo, agent fix own brain. Snake eat tail.
|
||||
|
||||
## Pick your grunt
|
||||
|
||||
Six levels. Switch anytime with `/caveman <level>`. Level sticks until you change it or the session ends.
|
||||
|
||||
| Level | Same sentence, shrunk |
|
||||
|---|---|
|
||||
| *normal agent* | You should wrap the object in `useMemo`, since a new reference is created on every render. |
|
||||
| `lite` | Wrap object in `useMemo`. New ref created every render. |
|
||||
| `full` *(default)* | New ref each render. Wrap object in `useMemo`. |
|
||||
| `ultra` | New ref/render. `useMemo` it. |
|
||||
| `wenyan` | New ref every render, so wrap in `useMemo` — rendered in classical Chinese, shorter still. |
|
||||
|
||||
> [!NOTE]
|
||||
> **Speak your tongue.** Caveman keeps your language. Write Portuguese, caveman grunt Portuguese. Spanish, French, same. It compresses the *style*, never translates. `wenyan` mode is the exception on purpose: classical Chinese packs the most meaning per token.
|
||||
|
||||
## What you get
|
||||
|
||||
| Command | What it does |
|
||||
|---|---|
|
||||
| `/caveman [lite\|full\|ultra\|wenyan]` | Compress every reply. Level sticks for the session. |
|
||||
| `/caveman-commit` | Conventional Commit messages, ≤50-char subject. Why over what. |
|
||||
| `/caveman-review` | One-line PR comments: `L42: 🔴 bug: user null. Add guard.` |
|
||||
| `/caveman-stats` | Real session token usage, lifetime savings, USD. Tweetable line with `--share`. |
|
||||
| `/caveman-compress <file>` | Rewrite a memory file (like `CLAUDE.md`) into caveman-speak. Cuts ~46% input tokens **every session after**. Code, URLs, paths byte-preserved. |
|
||||
| `caveman-shrink` | MCP middleware. Wraps any MCP server, compresses its tool descriptions. [npm](https://www.npmjs.com/package/caveman-shrink). |
|
||||
| `cavecrew-*` | Caveman subagents (investigator, builder, reviewer). ~60% fewer tokens than vanilla, so main context lasts longer. |
|
||||
|
||||
> [!TIP]
|
||||
> On Claude Code the statusline shows `[CAVEMAN] ⛏ 12.4k` — that's your lifetime tokens saved, updated on every `/caveman-stats`. Silence it with `CAVEMAN_STATUSLINE_SAVINGS=0`.
|
||||
|
||||
## Benchmarks
|
||||
|
||||
Real token counts from the Claude API ([reproduce it yourself](benchmarks/)):
|
||||
Real token counts from the Claude API. Average **65% output reduction** across 10 **chat-style prompts** (range 22–87%), measured against default verbose replies. Output tokens only, committed and reproducible in [`benchmarks/`](./benchmarks/) and [`evals/`](./evals/). This is one-question-one-answer, not a full agentic coding run — for that number, see [JetBrains](#independently-measured-jetbrains-86-tasks) below.
|
||||
|
||||
<!-- BENCHMARK-TABLE-START -->
|
||||
| Task | Normal (tokens) | Caveman (tokens) | Saved |
|
||||
|------|---------------:|----------------:|------:|
|
||||
| Task | Normal | Caveman | Saved |
|
||||
|------|-------:|--------:|------:|
|
||||
| Explain React re-render bug | 1180 | 159 | 87% |
|
||||
| Fix auth middleware token expiry | 704 | 121 | 83% |
|
||||
| Set up PostgreSQL connection pool | 2347 | 380 | 84% |
|
||||
@@ -116,178 +170,183 @@ Real token counts from the Claude API ([reproduce it yourself](benchmarks/)):
|
||||
| Debug PostgreSQL race condition | 1200 | 232 | 81% |
|
||||
| Implement React error boundary | 3454 | 456 | 87% |
|
||||
| **Average** | **1214** | **294** | **65%** |
|
||||
|
||||
*Range: 22%–87% savings across prompts.*
|
||||
<!-- BENCHMARK-TABLE-END -->
|
||||
|
||||
> [!IMPORTANT]
|
||||
> Caveman only affects output tokens — thinking/reasoning tokens are untouched. Caveman no make brain smaller. Caveman make *mouth* smaller. Biggest win is **readability and speed**, cost savings are a bonus.
|
||||
> **Honest number warning.** Caveman only shrinks **output** tokens. Input and reasoning tokens are untouched, and the skill itself adds ~1–1.5k input tokens per turn. So whole-session savings run smaller than the output number, and on already-terse workloads they can go net-negative. The real win is **readability and speed**. Cost savings are the bonus. When caveman wins, when it loses, and how to measure it yourself: **[docs/HONEST-NUMBERS.md](./docs/HONEST-NUMBERS.md)**.
|
||||
|
||||
### Science back caveman up
|
||||
### Independently measured: JetBrains, 86 tasks
|
||||
|
||||
A March 2026 paper ["Brevity Constraints Reverse Performance Hierarchies in Language Models"](https://arxiv.org/abs/2604.00025) found that constraining large models to brief responses **improved accuracy by 26 percentage points** on certain benchmarks and completely reversed performance hierarchies. Verbose not always better. Sometimes less word = more correct.
|
||||
JetBrains ran the skill against [86 tasks from SkillsBench](https://blog.jetbrains.com/ai/2026/07/speak-to-ai-agents-like-cavemen-tosave-tokens/) in July 2026 — real coding work, auto-graded by each task's own tests, Claude Code on `claude-sonnet-5`, skill forced on for every reply.
|
||||
|
||||
## Install
|
||||
| Workload | Output tokens saved | Measured by |
|
||||
|---|---:|---|
|
||||
| Chat-style prose | **65%** | us, table above |
|
||||
| Agentic coding run | **8.5%** | JetBrains, 86 tasks |
|
||||
|
||||
```bash
|
||||
npx skills add JuliusBrussee/caveman
|
||||
```
|
||||
Both numbers are real. They measure different workloads, and the gap is mechanical: caveman compresses narration and leaves code, diffs, tool calls, and error strings byte-exact. In a chat answer, narration is the whole reply. In an agentic run it's the thin layer between tool calls, so that's all there is to squeeze. An output-only skill has a low ceiling on work that is mostly not prose.
|
||||
|
||||
`npx skills` supports 40+ agents — Claude Code, GitHub Copilot, Cursor, Windsurf, Cline, and more. To install for a specific agent:
|
||||
Pick the number that matches your workload:
|
||||
|
||||
```bash
|
||||
npx skills add JuliusBrussee/caveman -a cursor
|
||||
npx skills add JuliusBrussee/caveman -a copilot
|
||||
npx skills add JuliusBrussee/caveman -a cline
|
||||
npx skills add JuliusBrussee/caveman -a windsurf
|
||||
```
|
||||
- **Agent writes you prose** — explanations, review, docs, debugging walkthroughs → 65% territory.
|
||||
- **Agent works a repo unattended** → single digits. Not zero, not 65%.
|
||||
|
||||
Or with Claude Code plugin system:
|
||||
Quality was unaffected: across 86 auto-graded tasks the two arms were statistically indistinguishable. Small mouth, same brain — checked by someone who didn't ship it.
|
||||
|
||||
```bash
|
||||
claude plugin marketplace add JuliusBrussee/caveman
|
||||
claude plugin install caveman@caveman
|
||||
```
|
||||
Two things follow:
|
||||
|
||||
Codex:
|
||||
- **Agentic bills are mostly input tokens**, which an output-only skill cannot touch by construction. `/caveman-compress` and `caveman-shrink` chip at that side; the skill alone never will.
|
||||
- **The right number is your number.** JetBrains had to run a full paid benchmark to find out what caveman does on their stack. That's the job [Caveman 2](#caveman-2) exists to do — for yours, continuously.
|
||||
|
||||
1. Clone repo
|
||||
2. Open Codex in repo
|
||||
3. Run `/plugins`
|
||||
4. Search `Caveman`
|
||||
5. Install plugin
|
||||
Turns out short isn't just cheaper. A March 2026 paper, [*Brevity Constraints Reverse Performance Hierarchies in Language Models*](https://arxiv.org/abs/2604.00025), tested 31 models and found that constraining large models to brief answers **improved accuracy by ~26 points** on some benchmarks. Sometimes less word = more correct.
|
||||
|
||||
Install once. Use in all sessions after that.
|
||||
<details>
|
||||
<summary><strong>caveman-compress receipts</strong> — real memory files, cutting input tokens forever</summary>
|
||||
|
||||
One rock. That it.
|
||||
|
||||
## Usage
|
||||
|
||||
Trigger with:
|
||||
- `/caveman` or Codex `$caveman`
|
||||
- "talk like caveman"
|
||||
- "caveman mode"
|
||||
- "less tokens please"
|
||||
|
||||
Stop with: "stop caveman" or "normal mode"
|
||||
|
||||
### Intensity Levels
|
||||
|
||||
Sometimes full caveman too much. Sometimes not enough. Now you pick:
|
||||
|
||||
| Level | Trigger | What it do |
|
||||
|-------|---------|------------|
|
||||
| **Lite** | `/caveman lite` or `$caveman lite` | Drop filler, keep grammar. Professional but no fluff |
|
||||
| **Full** | `/caveman full` or `$caveman full` | Default caveman. Drop articles, fragments, full grunt |
|
||||
| **Ultra** | `/caveman ultra` or `$caveman ultra` | Maximum compression. Telegraphic. Abbreviate everything |
|
||||
|
||||
Level stick until you change it or session end.
|
||||
|
||||
## What Caveman Do
|
||||
|
||||
| Thing | Caveman Do? |
|
||||
|-------|------------|
|
||||
| English explanation | 🪨 Caveman smash filler words |
|
||||
| Code blocks | ✍️ Write normal (caveman not stupid) |
|
||||
| Technical terms | 🧠 Keep exact (polymorphism stay polymorphism) |
|
||||
| Error messages | 📋 Quote exact |
|
||||
| Git commits & PRs | ✍️ Write normal |
|
||||
| Articles (a, an, the) | 💀 Gone |
|
||||
| Pleasantries | 💀 "Sure I'd be happy to" is dead |
|
||||
| Hedging | 💀 "It might be worth considering" extinct |
|
||||
|
||||
## Why
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────┐
|
||||
│ TOKENS SAVED ████████ 75% │
|
||||
│ TECHNICAL ACCURACY ████████ 100%│
|
||||
│ SPEED INCREASE ████████ ~3x │
|
||||
│ VIBES ████████ OOG │
|
||||
└─────────────────────────────────────┘
|
||||
```
|
||||
|
||||
- **Faster response** — less token to generate = speed go brrr
|
||||
- **Easier to read** — no wall of text, just the answer
|
||||
- **Same accuracy** — all technical info kept, only fluff removed ([science say so](https://arxiv.org/abs/2604.00025))
|
||||
- **Save money** — ~71% less output token = less cost
|
||||
- **Fun** — every code review become comedy
|
||||
|
||||
## How It Work
|
||||
|
||||
Caveman not dumb. Caveman **efficient**.
|
||||
|
||||
Normal LLM waste token on:
|
||||
- "I'd be happy to help you with that" (8 wasted tokens)
|
||||
- "The reason this is happening is because" (7 wasted tokens)
|
||||
- "I would recommend that you consider" (7 wasted tokens)
|
||||
- "Sure, let me take a look at that for you" (10 wasted tokens)
|
||||
|
||||
Caveman say what need saying. Then stop.
|
||||
|
||||
## Caveman Compress
|
||||
|
||||
Caveman makes Claude *speak* with fewer tokens. **Caveman Compress** makes Claude *read* fewer tokens.
|
||||
|
||||
Your `CLAUDE.md` loads on **every session start**. A 1000-token project memory file costs you tokens every single time you open a project. Caveman Compress rewrites those files into caveman-speak so Claude reads less — without you losing the human-readable original.
|
||||
|
||||
```
|
||||
/caveman-compress CLAUDE.md
|
||||
```
|
||||
|
||||
```
|
||||
CLAUDE.md ← compressed (Claude reads this every session — fewer tokens)
|
||||
CLAUDE.original.md ← human-readable backup (you read and edit this)
|
||||
```
|
||||
|
||||
### How it works
|
||||
|
||||
A Python pipeline that shells out to `claude --print` for the actual compression, then validates the result locally — no tokens wasted on checking.
|
||||
|
||||
```
|
||||
detect file type (local) → compress with Claude (1 call) → validate (local)
|
||||
↓
|
||||
if errors: targeted fix (1 call, cherry-pick only)
|
||||
↓
|
||||
retry up to 2×, restore original on failure
|
||||
```
|
||||
|
||||
### What's preserved exactly
|
||||
|
||||
Code blocks, inline code, URLs, file paths, commands, headings, table structure, dates, version numbers — anything technical passes through untouched. Only natural language prose gets compressed.
|
||||
|
||||
### Compress benchmarks
|
||||
<br>
|
||||
|
||||
| File | Original | Compressed | Saved |
|
||||
|------|----------:|----------:|------:|
|
||||
|---|---:|---:|---:|
|
||||
| `claude-md-preferences.md` | 706 | 285 | **59.6%** |
|
||||
| `project-notes.md` | 1145 | 535 | **53.3%** |
|
||||
| `claude-md-project.md` | 1122 | 687 | **38.8%** |
|
||||
| `claude-md-project.md` | 1122 | 636 | **43.3%** |
|
||||
| `todo-list.md` | 627 | 388 | **38.1%** |
|
||||
| `mixed-with-code.md` | 888 | 574 | **35.4%** |
|
||||
| **Average** | **898** | **494** | **45%** |
|
||||
| `mixed-with-code.md` | 888 | 560 | **36.9%** |
|
||||
| **Average** | **898** | **481** | **46%** |
|
||||
|
||||
### Full-circle token savings
|
||||
Every session after, that file loads ~46% smaller. Input tokens saved forever, not just one reply.
|
||||
|
||||
| Tool | What it cuts | Savings |
|
||||
|------|-------------|---------|
|
||||
| **caveman** | Output tokens (Claude's responses) | ~65% |
|
||||
| **caveman-compress** | Input tokens (memory files loaded per session) | ~45% |
|
||||
| **Both together** | The whole conversation | Output + input both shrunk |
|
||||
</details>
|
||||
|
||||
See the full [caveman-compress README](caveman-compress/README.md) for install, usage, and validation details.
|
||||
## The whole cave
|
||||
|
||||
## Star This Repo
|
||||
<table>
|
||||
<tr><td>
|
||||
|
||||
If caveman save you mass token, mass money — leave mass star. ⭐
|
||||
### <img src="docs/assets/dancing-rock.svg" width="20" height="20" alt=""> Want the whole agent, not just its mouth? → caveman-code
|
||||
|
||||
[](https://star-history.com/#JuliusBrussee/caveman&Date)
|
||||
This skill shrinks what an agent **says**. **[caveman-code](https://github.com/JuliusBrussee/caveman-code)** shrinks **everything** — a full terminal coding agent, caveman top to bottom. **~2× fewer tokens than Codex** on identical tasks. 20+ providers, plan mode, autopilot goal loop, MIT.
|
||||
|
||||
## Also by Julius Brussee
|
||||
```bash
|
||||
npm install -g @juliusbrussee/caveman-code
|
||||
```
|
||||
|
||||
- **[Blueprint](https://github.com/JuliusBrussee/blueprint)** — specification-driven development for Claude Code. Natural language → blueprints → parallel builds → working software.
|
||||
- **[Revu](https://github.com/JuliusBrussee/revu-swift)** — local-first macOS study app with FSRS spaced repetition, decks, exams, and study guides. [revu.cards](https://revu.cards)
|
||||
[**▶ Try caveman-code →**](https://github.com/JuliusBrussee/caveman-code)
|
||||
|
||||
## License
|
||||
</td></tr>
|
||||
</table>
|
||||
|
||||
Five tools, one idea: **agent do more with less.**
|
||||
|
||||
| Repo | What it shrinks |
|
||||
|------|------|
|
||||
| [**caveman**](https://github.com/JuliusBrussee/caveman) *(you here)* | What the agent **says** |
|
||||
| [**caveman-code**](https://github.com/JuliusBrussee/caveman-code) | The **whole agent**, end to end |
|
||||
| [**cavemem**](https://github.com/JuliusBrussee/cavemem) | What the agent **remembers**, across sessions |
|
||||
| [**cavekit**](https://github.com/JuliusBrussee/cavekit) | The **build loop** — spec-driven, no guessing |
|
||||
| [**cavegemma**](https://github.com/JuliusBrussee/finetune-caveman) | The compression **baked into weights** (Gemma fine-tune) |
|
||||
|
||||
<details>
|
||||
<summary><strong>Also: five sibling skills, one install</strong></summary>
|
||||
|
||||
<br>
|
||||
|
||||
[**JuliusBrussee/skills**](https://github.com/JuliusBrussee/skills) — works in Claude Code, Cursor, Gemini, Cline, Copilot, 40+ agents:
|
||||
|
||||
| Skill | What |
|
||||
|------|------|
|
||||
| [**caveman**](https://github.com/JuliusBrussee/skills/tree/main/skills/caveman) | This one. Speak less, say more. |
|
||||
| [**grill-me**](https://github.com/JuliusBrussee/skills/tree/main/skills/grill-me) | Agent grills your plan *before* you build the wrong thing. |
|
||||
| [**interface-kit**](https://github.com/JuliusBrussee/skills/tree/main/skills/interface-kit) | Build UI that looks good, loads fast, works for everyone. |
|
||||
| [**junior-to-senior**](https://github.com/JuliusBrussee/skills/tree/main/skills/junior-to-senior) | Adversarial review pass. Junior output in, senior output out. |
|
||||
| [**loop-factory**](https://github.com/JuliusBrussee/skills/tree/main/skills/loop-factory) | Spec-driven task loop — inbox → active → archive. |
|
||||
|
||||
```bash
|
||||
npx skills@latest add JuliusBrussee/skills
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><strong>🦞 Teach the lobster brevity — OpenClaw integration</strong></summary>
|
||||
|
||||
<br>
|
||||
|
||||
[**OpenClaw**](https://openclaw.ai) is a self-host gateway: one box, many agents inside, wired to Slack / Discord / iMessage / Telegram. Lobster strong. Lobster smart. Lobster also talk a lot.
|
||||
|
||||
Same installer, scoped to one agent:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash -s -- --only openclaw
|
||||
```
|
||||
|
||||
Two things happen, no more: a caveman skill lands in the workspace, and a tiny marker-fenced block is appended to `SOUL.md` (OpenClaw injects it every turn, so the lobster is terse from message one — no `/caveman` per session). Custom path? `OPENCLAW_WORKSPACE=/your/path`. Uninstall with the same line plus `--uninstall`; your other workspace content stays untouched. Lobster claw still sharp. Lobster mouth now small.
|
||||
|
||||
</details>
|
||||
|
||||
## Caveman 2
|
||||
|
||||
**Caveman make token small. Caveman 2 make it _provable_.**
|
||||
|
||||
Today's savings numbers (including `/caveman-stats`) are local estimates. Caveman 2 measures and verifies them across a whole team — real receipts, real dashboard, real proof the tokens went down. Building it now.
|
||||
|
||||
[The JetBrains result](#independently-measured-jetbrains-86-tasks) is the argument for it. 65% and 8.5% are both correct, and neither one is *your* number — one harness, one model, one task set, and your stack is none of those. The fix is not a better README claim, ours or anyone's. It's a baseline on your own traffic and a receipt at the end of the month.
|
||||
|
||||
[**Join the waitlist → caveman.so**](https://caveman.so)
|
||||
|
||||
## How it works
|
||||
|
||||
1. Install drops a skill file into your agent.
|
||||
2. Skill tells agent: drop filler, keep substance, use fragments — but never touch code, commands, or errors.
|
||||
3. On Claude Code, a hook writes a tiny flag file each session, so the agent talks caveman from message one without `/caveman`.
|
||||
4. `/caveman-stats` reads your session log, counts tokens saved, writes the number to your statusline.
|
||||
5. `/caveman-compress` rewrites memory files (like `CLAUDE.md`) so every future session starts with a smaller context. Save tokens forever, not just once.
|
||||
|
||||
Hook architecture, file ownership, and CI sync are documented for maintainers in [CLAUDE.md](./CLAUDE.md).
|
||||
|
||||
## Privacy
|
||||
|
||||
Caveman no phone home. No telemetry, no analytics, no accounts, no backend. After install, zero network calls — the skill is a prompt, the hooks are local scripts, and `/caveman-stats` reads a log already on your disk. Install-time fetches (GitHub plus your agents' own registries) are spelled out in [SECURITY.md](./SECURITY.md#privacy--telemetry).
|
||||
|
||||
## Sponsors
|
||||
|
||||
Caveman free forever. Sponsors keep the rock sharp.
|
||||
|
||||
<p align="center">
|
||||
<a href="https://www.atlascloud.ai">
|
||||
<picture>
|
||||
<source media="(prefers-color-scheme: dark)" srcset="docs/assets/atlas-cloud-dark.svg">
|
||||
<img src="docs/assets/atlas-cloud.svg" alt="Atlas Cloud" height="32">
|
||||
</picture>
|
||||
</a>
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://www.atlascloud.ai"><strong>Atlas Cloud</strong></a> — full-modal AI inference platform, one API.
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://github.com/sponsors/JuliusBrussee"><strong>Want your rock here? → Sponsor caveman</strong></a>
|
||||
</p>
|
||||
|
||||
## Star this repo
|
||||
|
||||
Caveman save you token, save you money. Star cost zero. Fair trade. ⭐
|
||||
|
||||
[](https://star-history.com/#JuliusBrussee/caveman&Date)
|
||||
|
||||
---
|
||||
|
||||
<sub>
|
||||
<strong>Docs:</strong>
|
||||
<a href="./INSTALL.md">Install matrix</a> ·
|
||||
<a href="./docs/HONEST-NUMBERS.md">Honest numbers</a> ·
|
||||
<a href="./CONTRIBUTING.md">Contributing</a> ·
|
||||
<a href="./CLAUDE.md">Maintainer guide</a> ·
|
||||
<a href="https://github.com/JuliusBrussee/caveman/issues">Issues</a>
|
||||
<br>
|
||||
<strong>Also by Julius Brussee:</strong>
|
||||
<a href="https://github.com/JuliusBrussee/revu-swift">Revu</a> — local-first macOS study app with FSRS spaced repetition (<a href="https://revu.cards">revu.cards</a>)
|
||||
<br><br>
|
||||
MIT — free like mass mammoth on open plain.
|
||||
</sub>
|
||||
|
||||
@@ -0,0 +1,46 @@
|
||||
# Security Policy
|
||||
|
||||
## Supported Versions
|
||||
|
||||
Only the latest stable release builds are supported with security patches.
|
||||
|
||||
## Reporting a Vulnerability
|
||||
|
||||
If you identify a security vulnerability in caveman (such as arbitrary shell execution, workspace folder escapes, token/credentials hijack via prompts, or malicious JSON parsing flaws in extension settings), please do **not** open a public issue.
|
||||
|
||||
Please report vulnerabilities privately by emailing the maintainers or using [GitHub's private vulnerability reporting](https://github.com/JuliusBrussee/caveman/security/advisories/new).
|
||||
|
||||
## Privacy & Telemetry
|
||||
|
||||
**Caveman has no telemetry. Zero.** No analytics, no crash reporting, no phone-home, no accounts, no API keys collected. There is no caveman backend — nothing to send data to.
|
||||
|
||||
### After install: zero network calls
|
||||
|
||||
Once installed, nothing in caveman touches the network. Verified against the code (audit it yourself — every file is in this repo):
|
||||
|
||||
- **The skill itself** (`skills/caveman/SKILL.md`) is a markdown prompt. It contains no code.
|
||||
- **The hooks** (`src/hooks/*.js`, statusline scripts) are local Node/shell scripts. They read and write local files only (flag file, session log, statusline savings file). No `http`/`https`/`fetch` anywhere in them.
|
||||
- **`/caveman-stats`** reads Claude Code's session JSONL from your local disk and prints counts. USD figures come from pricing constants hardcoded in the script. Nothing leaves your machine.
|
||||
- **`caveman-shrink`** (MCP middleware) spawns the MCP server *you* configure, locally, and compresses its output in-process. It makes no network calls of its own; any network activity belongs to the server you wrapped.
|
||||
- **`/caveman-compress`** rewrites a local file you name and saves a `.original.md` backup next to it. Local file I/O only.
|
||||
|
||||
### At install time: exactly these network requests, nothing else
|
||||
|
||||
- `curl … install.sh | bash` (or `irm … install.ps1 | iex`) fetches the shim from raw.githubusercontent.com, which delegates to `npx -y github:JuliusBrussee/caveman` — npm fetches this repo from GitHub.
|
||||
- The installer shells out to per-agent CLIs which fetch from their own registries: `claude plugin marketplace add` / `claude plugin install` (Anthropic/GitHub), `gemini extensions install`, `npm view caveman-shrink`, `npx -y skills add` (npm).
|
||||
- **Rare fallback:** if the installer runs detached from a repo checkout, it downloads the hook files from raw.githubusercontent.com **pinned to an immutable release tag** and verifies each against a published SHA-256 manifest before wiring anything (a mismatch aborts). From a normal clone or npx run, files are copied locally — offline installs work.
|
||||
|
||||
Nothing is uploaded in any of these steps. Details and the full list of paths written: [INSTALL.md → Privacy](./INSTALL.md#privacy).
|
||||
|
||||
### What stays on your machine
|
||||
|
||||
Everything. Skill/rule files in your agents' config dirs, the mode flag file and merged `settings.json` under `~/.claude/` (or `$CLAUDE_CONFIG_DIR`), the lifetime-savings statusline file, and `.original.md` backups from `/caveman-compress`. Uninstall removes what the installer wrote: `npx -y github:JuliusBrussee/caveman -- --uninstall`.
|
||||
|
||||
### Enterprise / air-gapped use
|
||||
|
||||
Caveman is self-contained after install and fully functional offline. There is no license server, no external backend, and no data flow to audit beyond the install-time fetches above. For air-gapped environments, clone the repo internally and run the installer from the clone — no network needed.
|
||||
|
||||
## About scanner warnings
|
||||
|
||||
- **Windows Defender / SmartScreen on `install.ps1` (#383):** piping a script from the internet into `iex` and writing into agent config directories matches generic dropper heuristics, so AV tools may warn. The script is short and readable in this repo; the hook files it installs are SHA-256-verified against the pinned release manifest. If you'd rather not pipe-to-shell, clone the repo and run `node cli/install.js` — same result, fully inspectable first.
|
||||
- **Snyk "High Risk" on `caveman-compress` (#28):** the compress skill instructs the agent to read a file you name, rewrite it in place, and save a backup. In-place file rewriting is exactly what generic risk scoring flags. It is a real capability, not hidden — but there is no network access, no shell execution beyond what's documented in [`skills/caveman-compress/`](./skills/caveman-compress/), and it never touches files you didn't name.
|
||||
@@ -0,0 +1,47 @@
|
||||
---
|
||||
name: cavecrew-builder
|
||||
description: >
|
||||
Surgical 1-2 file edit. Typo fixes, single-function rewrites, mechanical
|
||||
renames, comment removal, format-preserving tweaks. Hard refuses 3+ file
|
||||
scope. Returns caveman diff receipt. Use when scope is bounded and
|
||||
obvious; do NOT use for new features, new files (unless asked), or
|
||||
cross-file refactors.
|
||||
tools: [Read, Edit, Write, Grep, Glob]
|
||||
---
|
||||
|
||||
Caveman-ultra. Drop articles/filler. Code/paths exact, backticked. No narration.
|
||||
|
||||
## Scope
|
||||
|
||||
1 file ideal. 2 OK. 3+ → refuse.
|
||||
Edit existing only (new file iff user asked).
|
||||
No new abstractions. No drive-by refactors. No comment additions.
|
||||
No `Bash` available — cannot shell out, cannot push, cannot delete.
|
||||
|
||||
## Workflow
|
||||
|
||||
1. `Read` target(s). Never edit blind.
|
||||
2. `Edit` smallest diff that work.
|
||||
3. Re-`Read` to verify.
|
||||
4. Return receipt.
|
||||
|
||||
## Output (receipt)
|
||||
|
||||
```
|
||||
<path:line-range> — <change ≤10 words>.
|
||||
<path:line-range> — <change ≤10 words>.
|
||||
verified: <re-read OK | mismatch @ path:line>.
|
||||
```
|
||||
|
||||
Diff is the artifact. Receipt is the proof. No exploration story.
|
||||
|
||||
## Refusals (terminal lines)
|
||||
|
||||
3+ files → `too-big. split: <n one-line tasks>.`
|
||||
Destructive needed → `needs-confirm. op: <command>.`
|
||||
Spec ambiguous → `ambiguous. ask: <one question>.`
|
||||
Tests fail post-edit, can't fix in scope → `regressed. revert path:line. cause: <fragment>.`
|
||||
|
||||
## Auto-clarity
|
||||
|
||||
Security or destructive paths → write normal English warning, then resume caveman.
|
||||
@@ -0,0 +1,57 @@
|
||||
---
|
||||
name: cavecrew-investigator
|
||||
description: >
|
||||
Read-only code locator. Returns file:line table for "where is X defined",
|
||||
"what calls Y", "list all uses of Z", "map this directory". Output is
|
||||
caveman-compressed so the main thread eats ~60% fewer tokens than
|
||||
vanilla Explore. Refuses to suggest fixes.
|
||||
tools: [Read, Grep, Glob, Bash]
|
||||
model: haiku
|
||||
---
|
||||
|
||||
Caveman-ultra. Drop articles/filler/hedging. Code/symbols/paths exact, backticked. Lead with answer.
|
||||
|
||||
## Job
|
||||
|
||||
Locate. Report. Stop. Never edit, never propose fix.
|
||||
|
||||
## Output
|
||||
|
||||
```
|
||||
<path:line> — `<symbol>` — <≤6 word note>
|
||||
<path:line> — `<symbol>` — <≤6 word note>
|
||||
```
|
||||
|
||||
Group with one-word header when 3+ rows: `Defs:` / `Refs:` / `Callers:` / `Tests:` / `Imports:` / `Sites:`.
|
||||
Single hit → one line, no header.
|
||||
Zero hits → `No match.`
|
||||
Last line → totals: `2 defs, 5 refs.` (omit if 0 or 1).
|
||||
|
||||
## Tools
|
||||
|
||||
`Grep` for symbols/strings. `Glob` for paths. `Read` only specific ranges. `Bash` for `git log -S`/`git grep`/`find` when faster.
|
||||
|
||||
## Refusals
|
||||
|
||||
Asked to fix → `Read-only. Spawn cavecrew-builder.`
|
||||
Asked to design → `Read-only. Spawn cavecrew-builder or use main thread.`
|
||||
|
||||
## Auto-clarity
|
||||
|
||||
Security warnings, destructive ops → write normal English. Resume after.
|
||||
|
||||
## Example
|
||||
|
||||
Q: "where symlink-safe flag write?"
|
||||
|
||||
```
|
||||
Defs:
|
||||
- hooks/caveman-config.js:81 — `safeWriteFlag` — atomic write w/ O_NOFOLLOW
|
||||
- hooks/caveman-config.js:160 — `readFlag` — paired reader
|
||||
Callers:
|
||||
- hooks/caveman-mode-tracker.js:33,87
|
||||
- hooks/caveman-activate.js:40
|
||||
Tests:
|
||||
- tests/test_symlink_flag.js — 12 cases
|
||||
2 defs, 3 callers, 1 test file.
|
||||
```
|
||||
@@ -0,0 +1,48 @@
|
||||
---
|
||||
name: cavecrew-reviewer
|
||||
description: >
|
||||
Diff/branch/file reviewer. One line per finding, severity-tagged, no praise,
|
||||
no scope creep. Output format `path:line: <emoji> <severity>: <problem>. <fix>.`
|
||||
Use for "review this PR", "review my diff", "audit this file". Skips
|
||||
formatting nits unless they change meaning.
|
||||
tools: [Read, Grep, Bash]
|
||||
model: haiku
|
||||
---
|
||||
|
||||
Caveman-ultra. Findings only. No "looks good", no "I'd suggest", no preamble.
|
||||
|
||||
## Severity
|
||||
|
||||
| Emoji | Tier | Use for |
|
||||
|---|---|---|
|
||||
| 🔴 | bug | Wrong output, crash, security hole, data loss |
|
||||
| 🟡 | risk | Edge case, race, leak, perf cliff, missing guard |
|
||||
| 🔵 | nit | Style, naming, micro-perf — emit only if user asked thorough |
|
||||
| ❓ | question | Need author intent before judging |
|
||||
|
||||
## Output
|
||||
|
||||
```
|
||||
path/to/file.ts:42: 🔴 bug: token expiry uses `<` not `<=`. Off-by-one allows expired tokens 1 tick.
|
||||
path/to/file.ts:118: 🟡 risk: pool not closed on error path. Add `try/finally`.
|
||||
src/utils.ts:7: ❓ question: why duplicate `.trim()` here?
|
||||
totals: 1🔴 1🟡 1❓
|
||||
```
|
||||
|
||||
Zero findings → `No issues.`
|
||||
File order, ascending line numbers within file.
|
||||
|
||||
## Boundaries
|
||||
|
||||
- Review only what's in front of you. No "while we're here".
|
||||
- No big-refactor proposals.
|
||||
- Need more context → append `(see L<n> in <file>)`. Don't guess.
|
||||
- Formatting nits skipped unless they change meaning.
|
||||
|
||||
## Tools
|
||||
|
||||
`Bash` only for `git diff`/`git log -p`/`git show`. No mutating commands.
|
||||
|
||||
## Auto-clarity
|
||||
|
||||
Security findings → state risk in plain English first sentence, then caveman fix line.
|
||||
@@ -13,14 +13,23 @@ from pathlib import Path
|
||||
|
||||
import anthropic
|
||||
|
||||
# Load .env.local from repo root if it exists
|
||||
# The only env var this benchmark needs: the anthropic SDK reads it in
|
||||
# anthropic.Anthropic(). Read it — and ONLY it — from repo-root .env.local.
|
||||
# Deliberately narrow (issue #528): the old loader setdefault'ed EVERY key in
|
||||
# .env.local into os.environ, which security scanners rightly flag as an
|
||||
# exfiltration surface. Nothing else from the file is ever read or exported.
|
||||
_API_KEY_VAR = "ANTHROPIC_API_KEY"
|
||||
|
||||
_env_file = Path(__file__).parent.parent / ".env.local"
|
||||
if _env_file.exists():
|
||||
if _API_KEY_VAR not in os.environ and _env_file.exists():
|
||||
for line in _env_file.read_text().splitlines():
|
||||
line = line.strip()
|
||||
if line and not line.startswith("#") and "=" in line:
|
||||
key, _, value = line.partition("=")
|
||||
os.environ.setdefault(key.strip(), value.strip())
|
||||
if line.startswith("#") or "=" not in line:
|
||||
continue
|
||||
key, _, value = line.partition("=")
|
||||
if key.strip() == _API_KEY_VAR:
|
||||
os.environ.setdefault(_API_KEY_VAR, value.strip())
|
||||
break
|
||||
|
||||
SCRIPT_VERSION = "1.0.0"
|
||||
SCRIPT_DIR = Path(__file__).parent
|
||||
|
||||
@@ -1,158 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Caveman Memory Orchestrator
|
||||
|
||||
Usage:
|
||||
python memory/compress.py <filepath>
|
||||
"""
|
||||
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import List
|
||||
|
||||
from .detect import should_compress
|
||||
from .validate import validate
|
||||
|
||||
MAX_RETRIES = 2
|
||||
|
||||
|
||||
# ---------- Claude Calls ----------
|
||||
|
||||
|
||||
def call_claude(prompt: str) -> str:
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["claude", "--print"],
|
||||
input=prompt,
|
||||
text=True,
|
||||
capture_output=True,
|
||||
check=True,
|
||||
)
|
||||
return result.stdout.strip()
|
||||
except subprocess.CalledProcessError as e:
|
||||
raise RuntimeError(f"Claude call failed:\n{e.stderr}")
|
||||
|
||||
|
||||
def build_compress_prompt(original: str) -> str:
|
||||
return f"""
|
||||
Compress this markdown into caveman format.
|
||||
|
||||
STRICT RULES:
|
||||
- Do NOT modify anything inside ``` code blocks
|
||||
- Do NOT modify anything inside inline backticks
|
||||
- Preserve ALL URLs exactly
|
||||
- Preserve ALL headings exactly
|
||||
- Preserve file paths and commands
|
||||
|
||||
Only compress natural language.
|
||||
|
||||
TEXT:
|
||||
{original}
|
||||
"""
|
||||
|
||||
|
||||
def build_fix_prompt(original: str, compressed: str, errors: List[str]) -> str:
|
||||
errors_str = "\n".join(f"- {e}" for e in errors)
|
||||
return f"""You are fixing a caveman-compressed markdown file. Specific validation errors were found.
|
||||
|
||||
CRITICAL RULES:
|
||||
- DO NOT recompress or rephrase the file
|
||||
- ONLY fix the listed errors — leave everything else exactly as-is
|
||||
- The ORIGINAL is provided as reference only (to restore missing content)
|
||||
- Preserve caveman style in all untouched sections
|
||||
|
||||
ERRORS TO FIX:
|
||||
{errors_str}
|
||||
|
||||
HOW TO FIX:
|
||||
- Missing URL: find it in ORIGINAL, restore it exactly where it belongs in COMPRESSED
|
||||
- Code block mismatch: find the exact code block in ORIGINAL, restore it in COMPRESSED
|
||||
- Heading mismatch: restore the exact heading text from ORIGINAL into COMPRESSED
|
||||
- Do not touch any section not mentioned in the errors
|
||||
|
||||
ORIGINAL (reference only):
|
||||
{original}
|
||||
|
||||
COMPRESSED (fix this):
|
||||
{compressed}
|
||||
|
||||
Return ONLY the fixed compressed file. No explanation.
|
||||
"""
|
||||
|
||||
|
||||
# ---------- Core Logic ----------
|
||||
|
||||
|
||||
def compress_file(filepath: Path) -> bool:
|
||||
print(f"📄 Processing: {filepath}")
|
||||
|
||||
if not should_compress(filepath):
|
||||
print("⚠️ Skipping (not natural language)")
|
||||
return False
|
||||
|
||||
original_text = filepath.read_text(errors="ignore")
|
||||
backup_path = filepath.with_name(filepath.stem + ".original.md")
|
||||
|
||||
# Step 1: Compress
|
||||
print("🧠 Compressing with Claude...")
|
||||
compressed = call_claude(build_compress_prompt(original_text))
|
||||
|
||||
# Save original as backup, write compressed to original path
|
||||
backup_path.write_text(original_text)
|
||||
filepath.write_text(compressed)
|
||||
|
||||
# Step 2: Validate + Retry
|
||||
for attempt in range(MAX_RETRIES):
|
||||
print(f"\n🔍 Validation attempt {attempt + 1}")
|
||||
|
||||
result = validate(backup_path, filepath)
|
||||
|
||||
if result.is_valid:
|
||||
print("✅ Validation passed")
|
||||
break
|
||||
|
||||
print("❌ Validation failed:")
|
||||
for err in result.errors:
|
||||
print(f" - {err}")
|
||||
|
||||
if attempt == MAX_RETRIES - 1:
|
||||
# Restore original on failure
|
||||
filepath.write_text(original_text)
|
||||
backup_path.unlink(missing_ok=True)
|
||||
print("❌ Failed after retries — original restored")
|
||||
return False
|
||||
|
||||
print("🛠 Fixing with Claude...")
|
||||
compressed = call_claude(
|
||||
build_fix_prompt(original_text, compressed, result.errors)
|
||||
)
|
||||
filepath.write_text(compressed)
|
||||
|
||||
return True
|
||||
|
||||
|
||||
# ---------- Main ----------
|
||||
|
||||
|
||||
def main():
|
||||
if len(sys.argv) != 2:
|
||||
print("Usage: python memory/compress.py <filepath>")
|
||||
sys.exit(1)
|
||||
|
||||
filepath = Path(sys.argv[1])
|
||||
|
||||
if not filepath.exists():
|
||||
print(f"❌ File not found: {filepath}")
|
||||
sys.exit(1)
|
||||
|
||||
success = compress_file(filepath)
|
||||
|
||||
if success:
|
||||
sys.exit(0)
|
||||
else:
|
||||
sys.exit(2)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,72 +0,0 @@
|
||||
---
|
||||
name: caveman
|
||||
description: >
|
||||
Ultra-compressed communication mode. Slash token usage ~75% by speaking like caveman
|
||||
while keeping full technical accuracy. Use when user says "caveman mode", "talk like caveman",
|
||||
"use caveman", "less tokens", "be brief", or invokes /caveman. Also auto-triggers
|
||||
when token efficiency is requested.
|
||||
---
|
||||
|
||||
# Caveman Mode
|
||||
|
||||
## Core Rule
|
||||
|
||||
Respond like smart caveman. Cut articles, filler, pleasantries. Keep all technical substance.
|
||||
|
||||
## Grammar
|
||||
|
||||
- Drop articles (a, an, the)
|
||||
- Drop filler (just, really, basically, actually, simply)
|
||||
- Drop pleasantries (sure, certainly, of course, happy to)
|
||||
- Short synonyms (big not extensive, fix not "implement a solution for")
|
||||
- No hedging (skip "it might be worth considering")
|
||||
- Fragments fine. No need full sentence
|
||||
- Technical terms stay exact. "Polymorphism" stays "polymorphism"
|
||||
- Code blocks unchanged. Caveman speak around code, not in code
|
||||
- Error messages quoted exact. Caveman only for explanation
|
||||
|
||||
## Pattern
|
||||
|
||||
```
|
||||
[thing] [action] [reason]. [next step].
|
||||
```
|
||||
|
||||
Not:
|
||||
> Sure! I'd be happy to help you with that. The issue you're experiencing is likely caused by...
|
||||
|
||||
Yes:
|
||||
> Bug in auth middleware. Token expiry check use `<` not `<=`. Fix:
|
||||
|
||||
## Examples
|
||||
|
||||
**User:** Why is my React component re-rendering?
|
||||
|
||||
**Normal (69 tokens):** "The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object."
|
||||
|
||||
**Caveman (19 tokens):** "New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`."
|
||||
|
||||
---
|
||||
|
||||
**User:** How do I set up a PostgreSQL connection pool?
|
||||
|
||||
**Caveman:**
|
||||
```
|
||||
Use `pg` pool:
|
||||
```
|
||||
```js
|
||||
const pool = new Pool({
|
||||
max: 20,
|
||||
idleTimeoutMillis: 30000,
|
||||
connectionTimeoutMillis: 2000,
|
||||
})
|
||||
```
|
||||
```
|
||||
max = concurrent connections. Keep under DB limit. idleTimeout kill stale conn.
|
||||
```
|
||||
|
||||
## Boundaries
|
||||
|
||||
- Code: write normal. Caveman English only
|
||||
- Git commits: normal
|
||||
- PR descriptions: normal
|
||||
- User say "stop caveman" or "normal mode": revert immediately
|
||||
@@ -0,0 +1,318 @@
|
||||
// caveman → OpenClaw install / uninstall helper.
|
||||
//
|
||||
// OpenClaw is a self-hosted gateway that orchestrates Claude Code, Codex,
|
||||
// Pi, OpenCode, and others. It has its own workspace + skills system at
|
||||
// ~/.openclaw/workspace/. Skills there appear in a compact list and are
|
||||
// loaded on-demand by the model — they are NOT injected as system prompt
|
||||
// each turn. The bootstrap files (AGENTS.md, SOUL.md, TOOLS.md, MEMORY.md)
|
||||
// ARE injected each turn under "Project Context", subject to a 12K-per-file
|
||||
// and 60K-total cap.
|
||||
//
|
||||
// To make caveman always-on through OpenClaw, we do two writes:
|
||||
// 1. Drop a copy of skills/caveman/SKILL.md into <workspace>/skills/caveman/
|
||||
// with OpenClaw-required frontmatter (`version`, `always: true`) merged
|
||||
// in. Makes the skill discoverable via `openclaw skills list` and lets
|
||||
// the orchestrated agent `read` it on demand.
|
||||
// 2. Append a tiny marker-fenced bootstrap snippet to <workspace>/SOUL.md
|
||||
// pointing the agent at the skill. SOUL.md is auto-injected each turn,
|
||||
// so this is what actually drives always-on behavior.
|
||||
//
|
||||
// Idempotent on both writes. Uninstall removes the skill folder and strips
|
||||
// the marker block from SOUL.md while preserving any user-authored content.
|
||||
|
||||
'use strict';
|
||||
|
||||
const fs = require('fs');
|
||||
const os = require('os');
|
||||
const path = require('path');
|
||||
|
||||
const SKILL_NAME = 'caveman';
|
||||
const SKILL_VERSION = '1.0.0';
|
||||
const MARK_BEGIN = '<!-- caveman-begin -->';
|
||||
const MARK_END = '<!-- caveman-end -->';
|
||||
const SOUL_FILE = 'SOUL.md';
|
||||
|
||||
function resolveWorkspace(env = process.env) {
|
||||
if (env.OPENCLAW_WORKSPACE) return path.resolve(env.OPENCLAW_WORKSPACE);
|
||||
return path.join(os.homedir(), '.openclaw', 'workspace');
|
||||
}
|
||||
|
||||
function readIfExists(p) {
|
||||
try { return fs.readFileSync(p, 'utf8'); } catch (_) { return null; }
|
||||
}
|
||||
|
||||
// ── Frontmatter helpers ───────────────────────────────────────────────────
|
||||
// Lightweight YAML merge — we only need to insert `version` and `always` if
|
||||
// they're absent. Avoids pulling in a YAML dep for a job this small. The
|
||||
// caveman SKILL.md uses block-scalar `description: >`, which a naive split
|
||||
// would mangle — but since we're only ever appending top-level keys (never
|
||||
// editing existing ones), a string-prepend after the leading `---\n` is safe.
|
||||
|
||||
function splitFrontmatter(src) {
|
||||
if (!src.startsWith('---\n') && !src.startsWith('---\r\n')) {
|
||||
return { frontmatter: '', body: src };
|
||||
}
|
||||
const after = src.slice(src.indexOf('\n') + 1);
|
||||
const endRe = /(^|\n)---\s*(\r?\n|$)/;
|
||||
const m = endRe.exec(after);
|
||||
if (!m) return { frontmatter: '', body: src };
|
||||
const fmEnd = m.index + (m[1] ? 1 : 0);
|
||||
const fm = after.slice(0, fmEnd);
|
||||
const rest = after.slice(m.index + m[0].length);
|
||||
return { frontmatter: fm, body: rest };
|
||||
}
|
||||
|
||||
function frontmatterHasKey(fm, key) {
|
||||
const re = new RegExp('(^|\\n)' + key + '\\s*:', 'i');
|
||||
return re.test(fm);
|
||||
}
|
||||
|
||||
// `opts.version` defaults to SKILL_VERSION (the '1.0.0' fallback) when the
|
||||
// caller doesn't have a better one on hand — cli/install.js threads through
|
||||
// PINNED_REF (its release-tag source of truth) instead so the two never
|
||||
// drift. `opts.always` defaults to true (existing behavior); pass `false`
|
||||
// (from --no-always) to omit the `always: true` key entirely — the skill
|
||||
// then loads on demand instead of always-on.
|
||||
function mergeOpenclawFrontmatter(src, opts = {}) {
|
||||
const version = opts.version || SKILL_VERSION;
|
||||
const always = opts.always !== false;
|
||||
const { frontmatter, body } = splitFrontmatter(src);
|
||||
const additions = [];
|
||||
if (!frontmatterHasKey(frontmatter, 'name')) additions.push(`name: ${SKILL_NAME}`);
|
||||
if (!frontmatterHasKey(frontmatter, 'version')) additions.push(`version: ${version}`);
|
||||
if (always && !frontmatterHasKey(frontmatter, 'always')) additions.push('always: true');
|
||||
if (additions.length === 0 && frontmatter) return src;
|
||||
const fmBody = (frontmatter ? frontmatter.trimEnd() + '\n' : '') + additions.join('\n') + (additions.length ? '\n' : '');
|
||||
return '---\n' + fmBody + '---\n' + body;
|
||||
}
|
||||
|
||||
// ── Bootstrap snippet load ────────────────────────────────────────────────
|
||||
function loadBootstrapSnippet(repoRoot) {
|
||||
if (repoRoot) {
|
||||
const p = path.join(repoRoot, 'src', 'rules', 'caveman-openclaw-bootstrap.md');
|
||||
const body = readIfExists(p);
|
||||
if (body) return body.endsWith('\n') ? body : body + '\n';
|
||||
}
|
||||
// Standalone fallback (curl|node case where there's no repo on disk).
|
||||
// Keep this in sync with src/rules/caveman-openclaw-bootstrap.md.
|
||||
return [
|
||||
MARK_BEGIN,
|
||||
'## Caveman mode (always on)',
|
||||
'',
|
||||
'Respond terse like smart caveman. All technical substance stay. Only fluff die.',
|
||||
'',
|
||||
"The full ruleset and intensity levels live in this workspace's caveman skill:",
|
||||
'',
|
||||
' skills/caveman/SKILL.md',
|
||||
'',
|
||||
'Default intensity: `full`. Switch with `/caveman lite|full|ultra|wenyan`.',
|
||||
'Stop with: "stop caveman" / "normal mode" / "deactivate caveman".',
|
||||
'',
|
||||
'Auto-Clarity: drop caveman for security warnings, irreversible action',
|
||||
'confirmations, multi-step sequences where fragments risk misread, or when',
|
||||
'user is confused or repeating. Resume after.',
|
||||
'',
|
||||
'Boundaries: code, commit messages, and PR descriptions stay normal prose.',
|
||||
MARK_END,
|
||||
'',
|
||||
].join('\n');
|
||||
}
|
||||
|
||||
function loadSkillBody(repoRoot) {
|
||||
if (!repoRoot) return null;
|
||||
return readIfExists(path.join(repoRoot, 'skills', 'caveman', 'SKILL.md'));
|
||||
}
|
||||
|
||||
// ── SOUL.md marker-block append/strip ─────────────────────────────────────
|
||||
//
|
||||
// Damage tolerance (#596): a stray or truncated marker (interrupted write,
|
||||
// partial user edit) used to chain into data loss — append saw "no complete
|
||||
// block" and added a SECOND block; strip then cut from the FIRST begin to the
|
||||
// FIRST end, which spanned all user content between the stray marker and the
|
||||
// appended block. The scan below pairs each begin with the nearest end BEFORE
|
||||
// the next begin; an unpaired marker is removed as just the marker itself,
|
||||
// never as a span over user content.
|
||||
|
||||
function stripAllBootstrapBlocks(text) {
|
||||
let result = '';
|
||||
let found = false;
|
||||
let i = 0;
|
||||
while (i < text.length) {
|
||||
const b = text.indexOf(MARK_BEGIN, i);
|
||||
if (b === -1) { result += text.slice(i); break; }
|
||||
result += text.slice(i, b);
|
||||
found = true;
|
||||
const nextB = text.indexOf(MARK_BEGIN, b + MARK_BEGIN.length);
|
||||
const e = text.indexOf(MARK_END, b + MARK_BEGIN.length);
|
||||
if (e !== -1 && (nextB === -1 || e < nextB)) {
|
||||
i = e + MARK_END.length; // well-formed block — drop begin..end inclusive
|
||||
} else {
|
||||
i = b + MARK_BEGIN.length; // orphan begin — drop only the marker itself
|
||||
}
|
||||
// Collapse the blank-line scar around the cut (same cosmetic rule the
|
||||
// old single-cut code applied): keep at most one newline on each side.
|
||||
result = result.replace(/\n+$/, '\n');
|
||||
const lead = /^\n+/.exec(text.slice(i));
|
||||
if (lead) i += lead[0].length - (result ? 1 : 0);
|
||||
}
|
||||
// Orphan end markers (begin already gone or never written) — drop marker only.
|
||||
while (result.includes(MARK_END)) { found = true; result = result.replace(MARK_END, ''); }
|
||||
return { next: result, found };
|
||||
}
|
||||
|
||||
function appendBootstrapToSoul(soulPath, snippet) {
|
||||
const existing = readIfExists(soulPath);
|
||||
const count = (s, sub) => s.split(sub).length - 1;
|
||||
let base = existing;
|
||||
let repaired = false;
|
||||
if (existing) {
|
||||
const nb = count(existing, MARK_BEGIN);
|
||||
const ne = count(existing, MARK_END);
|
||||
if (nb === 1 && ne === 1 && existing.indexOf(MARK_END) > existing.indexOf(MARK_BEGIN)) {
|
||||
return { changed: false, reason: 'already present' };
|
||||
}
|
||||
if (nb > 0 || ne > 0) {
|
||||
// Damaged markers — strip them safely first, then append one clean block.
|
||||
base = stripAllBootstrapBlocks(existing).next;
|
||||
repaired = true;
|
||||
}
|
||||
}
|
||||
let next;
|
||||
if (base && base.length) {
|
||||
const sep = base.endsWith('\n\n') ? '' : (base.endsWith('\n') ? '\n' : '\n\n');
|
||||
next = base + sep + snippet;
|
||||
} else {
|
||||
next = snippet;
|
||||
}
|
||||
fs.writeFileSync(soulPath, next, { mode: 0o644 });
|
||||
return repaired ? { changed: true, repaired: true } : { changed: true };
|
||||
}
|
||||
|
||||
function stripBootstrapFromSoul(soulPath) {
|
||||
const existing = readIfExists(soulPath);
|
||||
if (!existing) return { changed: false, reason: 'no SOUL.md' };
|
||||
const { next: stripped, found } = stripAllBootstrapBlocks(existing);
|
||||
if (!found) return { changed: false, reason: 'no marker block' };
|
||||
let next = stripped.trimEnd();
|
||||
next = next ? next + '\n' : '';
|
||||
if (next === '') {
|
||||
// SOUL.md only contained our block — remove the file so OpenClaw doesn't
|
||||
// bootstrap an empty section every turn.
|
||||
try { fs.unlinkSync(soulPath); } catch (_) {}
|
||||
return { changed: true, removed: true };
|
||||
}
|
||||
fs.writeFileSync(soulPath, next, { mode: 0o644 });
|
||||
return { changed: true };
|
||||
}
|
||||
|
||||
// ── Public API ────────────────────────────────────────────────────────────
|
||||
// `version` — bare semver stamped into the skill frontmatter; defaults to
|
||||
// SKILL_VERSION if the caller doesn't pass one (see mergeOpenclawFrontmatter).
|
||||
// `always` — default true (existing behavior). Pass false (--no-always) to
|
||||
// skip the `always: true` frontmatter key AND the SOUL.md bootstrap append,
|
||||
// so the skill installs load-on-demand instead of always-on.
|
||||
function installOpenclaw({ workspace, repoRoot, dryRun = false, force = false, log = noopLog(), version, always = true } = {}) {
|
||||
const ws = workspace || resolveWorkspace();
|
||||
const skillBody = loadSkillBody(repoRoot);
|
||||
if (!skillBody) {
|
||||
log.warn(' openclaw install requires the caveman repo on disk (skills/caveman/SKILL.md missing).');
|
||||
log.note(' Re-run from a clone or via `npx -y github:JuliusBrussee/caveman -- --only openclaw`.');
|
||||
return { ok: false, reason: 'repo not available' };
|
||||
}
|
||||
const snippet = loadBootstrapSnippet(repoRoot);
|
||||
|
||||
if (!fs.existsSync(ws)) {
|
||||
if (!force) {
|
||||
log.warn(` openclaw workspace not found at ${ws}.`);
|
||||
log.note(' Either install OpenClaw (https://openclaw.ai) and re-run, or pass --force to mkdir.');
|
||||
return { ok: false, reason: 'workspace missing' };
|
||||
}
|
||||
if (!dryRun) fs.mkdirSync(ws, { recursive: true });
|
||||
}
|
||||
|
||||
const skillDir = path.join(ws, 'skills', SKILL_NAME);
|
||||
const skillFile = path.join(skillDir, 'SKILL.md');
|
||||
const soulFile = path.join(ws, SOUL_FILE);
|
||||
|
||||
if (dryRun) {
|
||||
log.note(` would write ${skillFile} (with version${always ? '/always' : ''} frontmatter)`);
|
||||
if (always) {
|
||||
log.note(` would ${fs.existsSync(soulFile) ? 'append to' : 'create'} ${soulFile} (caveman bootstrap block)`);
|
||||
} else {
|
||||
log.note(' --no-always: would skip SOUL.md bootstrap append (skill loads on demand)');
|
||||
}
|
||||
return { ok: true, dryRun: true };
|
||||
}
|
||||
|
||||
fs.mkdirSync(skillDir, { recursive: true });
|
||||
const merged = mergeOpenclawFrontmatter(skillBody, { version, always });
|
||||
fs.writeFileSync(skillFile, merged, { mode: 0o644 });
|
||||
log.write(` installed: ${skillFile}\n`);
|
||||
|
||||
if (always) {
|
||||
const soul = appendBootstrapToSoul(soulFile, snippet);
|
||||
if (soul.changed) log.write(` wrote bootstrap block to ${soulFile}\n`);
|
||||
else log.note(` ${soulFile} already contains caveman bootstrap`);
|
||||
} else {
|
||||
log.note(' --no-always: skipped SOUL.md bootstrap append (skill loads on demand via `openclaw skills list`)');
|
||||
}
|
||||
|
||||
return { ok: true };
|
||||
}
|
||||
|
||||
function uninstallOpenclaw({ workspace, dryRun = false, log = noopLog() } = {}) {
|
||||
const ws = workspace || resolveWorkspace();
|
||||
const skillDir = path.join(ws, 'skills', SKILL_NAME);
|
||||
const soulFile = path.join(ws, SOUL_FILE);
|
||||
|
||||
let touched = false;
|
||||
|
||||
if (fs.existsSync(skillDir)) {
|
||||
if (dryRun) {
|
||||
log.note(` would remove ${skillDir}/`);
|
||||
} else {
|
||||
try { fs.rmSync(skillDir, { recursive: true, force: true }); } catch (_) {}
|
||||
log.note(` removed ${skillDir}`);
|
||||
}
|
||||
touched = true;
|
||||
}
|
||||
|
||||
if (fs.existsSync(soulFile)) {
|
||||
if (dryRun) {
|
||||
log.note(` would strip caveman block from ${soulFile}`);
|
||||
touched = true;
|
||||
} else {
|
||||
const r = stripBootstrapFromSoul(soulFile);
|
||||
if (r.changed) {
|
||||
log.note(r.removed ? ` removed ${soulFile}` : ` stripped caveman block from ${soulFile}`);
|
||||
touched = true;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return { ok: true, touched };
|
||||
}
|
||||
|
||||
function noopLog() {
|
||||
return {
|
||||
write: (_) => {},
|
||||
note: (_) => {},
|
||||
warn: (_) => {},
|
||||
};
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
installOpenclaw,
|
||||
uninstallOpenclaw,
|
||||
resolveWorkspace,
|
||||
// exported for tests
|
||||
mergeOpenclawFrontmatter,
|
||||
splitFrontmatter,
|
||||
appendBootstrapToSoul,
|
||||
stripBootstrapFromSoul,
|
||||
loadBootstrapSnippet,
|
||||
MARK_BEGIN,
|
||||
MARK_END,
|
||||
SKILL_NAME,
|
||||
SKILL_VERSION,
|
||||
};
|
||||
@@ -0,0 +1,42 @@
|
||||
'use strict';
|
||||
|
||||
// Strip the `tools:` field from a Claude-Code-style subagent frontmatter so
|
||||
// the file is valid for opencode, whose schema rejects the YAML array form
|
||||
// (`tools: [Read, Grep, Bash]`) with:
|
||||
//
|
||||
// Configuration is invalid at .../agents/cavecrew-reviewer.md
|
||||
// ↳ Expected object | undefined, got ["Read","Grep","Bash"] tools
|
||||
//
|
||||
// opencode allows `tools` to be a map (`{read: true, grep: true}`) or
|
||||
// omitted entirely. Omitting falls back to opencode's default tool set,
|
||||
// which is what the cavecrew subagent prompts already self-restrict against
|
||||
// in their body ("Read-only locator", "No `Bash` available", etc.), so
|
||||
// dropping the array form is safe.
|
||||
|
||||
const TOOLS_FIELD_RE = /^tools[ \t]*:/;
|
||||
const CONTINUATION_RE = /^[ \t]/;
|
||||
const FRONTMATTER_FENCE = '---\n';
|
||||
|
||||
function stripOpencodeAgentTools(content) {
|
||||
if (typeof content !== 'string' || !content.startsWith(FRONTMATTER_FENCE)) return content;
|
||||
const fmEnd = content.indexOf('\n---', FRONTMATTER_FENCE.length);
|
||||
if (fmEnd < 0) return content;
|
||||
|
||||
const fm = content.slice(FRONTMATTER_FENCE.length, fmEnd);
|
||||
const rest = content.slice(fmEnd);
|
||||
|
||||
const out = [];
|
||||
let dropping = false;
|
||||
for (const line of fm.split('\n')) {
|
||||
if (dropping) {
|
||||
if (CONTINUATION_RE.test(line)) continue;
|
||||
dropping = false;
|
||||
}
|
||||
if (TOOLS_FIELD_RE.test(line)) { dropping = true; continue; }
|
||||
out.push(line);
|
||||
}
|
||||
|
||||
return FRONTMATTER_FENCE + out.join('\n') + rest;
|
||||
}
|
||||
|
||||
module.exports = { stripOpencodeAgentTools };
|
||||
@@ -0,0 +1,368 @@
|
||||
// caveman — JSONC-tolerant settings.json read/write + defensive hook validation.
|
||||
//
|
||||
// Lifted in spirit from gsd-build/get-shit-done's stripJsonComments + readSettings.
|
||||
// Reused by cli/install.js and (optionally) by hooks/caveman-activate.js so a
|
||||
// commented settings.json no longer crashes the installer or the runtime hooks.
|
||||
//
|
||||
// Public API:
|
||||
// readSettings(path) → object, {}, or null on hard parse failure
|
||||
// writeSettings(path, obj) → atomic write with newline
|
||||
// stripJsonComments(src) → string with // and /* */ stripped (string-aware)
|
||||
// validateHookFields(settings) → mutates: drops malformed hook entries
|
||||
// hasCavemanHook(settings, ev) → idempotency probe
|
||||
// addCommandHook(settings, ev, opts) → no-op if substring marker already present
|
||||
// removeCavemanHooks(settings) → uninstall helper
|
||||
//
|
||||
// Pure stdlib, CommonJS, Node ≥14.
|
||||
|
||||
'use strict';
|
||||
|
||||
const fs = require('fs');
|
||||
const os = require('os');
|
||||
const path = require('path');
|
||||
const crypto = require('crypto');
|
||||
|
||||
// ── stripJsonComments ──────────────────────────────────────────────────────
|
||||
// Hand-rolled state machine. Tracks string state + backslash escape so a
|
||||
// comment-looking sequence inside a quoted string is left alone. Removes
|
||||
// trailing commas in a final pass — JSONC tolerates those, JSON.parse does not.
|
||||
function stripJsonComments(src) {
|
||||
if (typeof src !== 'string') return src;
|
||||
let out = '';
|
||||
let i = 0;
|
||||
const n = src.length;
|
||||
let inString = false;
|
||||
let stringChar = '';
|
||||
let inLine = false;
|
||||
let inBlock = false;
|
||||
while (i < n) {
|
||||
const c = src[i];
|
||||
const next = i + 1 < n ? src[i + 1] : '';
|
||||
if (inLine) {
|
||||
if (c === '\n') { inLine = false; out += c; }
|
||||
i++; continue;
|
||||
}
|
||||
if (inBlock) {
|
||||
if (c === '*' && next === '/') { inBlock = false; i += 2; continue; }
|
||||
i++; continue;
|
||||
}
|
||||
if (inString) {
|
||||
out += c;
|
||||
if (c === '\\') { if (i + 1 < n) { out += src[i + 1]; i += 2; continue; } }
|
||||
if (c === stringChar) { inString = false; }
|
||||
i++; continue;
|
||||
}
|
||||
if (c === '"' || c === "'") { inString = true; stringChar = c; out += c; i++; continue; }
|
||||
if (c === '/' && next === '/') { inLine = true; i += 2; continue; }
|
||||
if (c === '/' && next === '*') { inBlock = true; i += 2; continue; }
|
||||
out += c; i++;
|
||||
}
|
||||
return stripTrailingCommas(out);
|
||||
}
|
||||
|
||||
// ── stripTrailingCommas ────────────────────────────────────────────────────
|
||||
// Remove `,` when the next non-whitespace char is `}` or `]` — but only
|
||||
// OUTSIDE strings. The old global regex ran over string contents too and
|
||||
// silently corrupted values like `"echo ,}"` → `"echo }"` (issue #595);
|
||||
// comment-stripping does not sanitize string bodies, so a string-aware scan
|
||||
// is required here as well.
|
||||
function stripTrailingCommas(src) {
|
||||
let out = '';
|
||||
let i = 0;
|
||||
const n = src.length;
|
||||
let inString = false;
|
||||
let stringChar = '';
|
||||
while (i < n) {
|
||||
const c = src[i];
|
||||
if (inString) {
|
||||
out += c;
|
||||
if (c === '\\') { if (i + 1 < n) { out += src[i + 1]; i += 2; continue; } }
|
||||
if (c === stringChar) inString = false;
|
||||
i++; continue;
|
||||
}
|
||||
if (c === '"' || c === "'") { inString = true; stringChar = c; out += c; i++; continue; }
|
||||
if (c === ',') {
|
||||
let j = i + 1;
|
||||
while (j < n && /\s/.test(src[j])) j++;
|
||||
if (j < n && (src[j] === '}' || src[j] === ']')) { i++; continue; } // drop the comma
|
||||
}
|
||||
out += c; i++;
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// ── readSettings ───────────────────────────────────────────────────────────
|
||||
// Try strict JSON first (fast path). On failure, strip comments and retry.
|
||||
// On total failure return `null` and warn — never silently overwrite a
|
||||
// malformed-but-recoverable file with `{}`.
|
||||
function readSettings(p) {
|
||||
if (!fs.existsSync(p)) return {};
|
||||
let raw;
|
||||
try { raw = fs.readFileSync(p, 'utf8'); }
|
||||
catch (e) {
|
||||
process.stderr.write(`caveman: cannot read ${p}: ${e.message}\n`);
|
||||
return null;
|
||||
}
|
||||
if (!raw.trim()) return {};
|
||||
try { return JSON.parse(raw); } catch (_) { /* fall through to JSONC */ }
|
||||
try { return JSON.parse(stripJsonComments(raw)); }
|
||||
catch (e) {
|
||||
process.stderr.write(`caveman: warning — ${p} is not valid JSON or JSONC: ${e.message}\n`);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
// ── writeSettings ──────────────────────────────────────────────────────────
|
||||
// Atomic write: temp file + rename. mode 0600 (settings often contains tokens).
|
||||
function writeSettings(p, obj) {
|
||||
const dir = path.dirname(p);
|
||||
fs.mkdirSync(dir, { recursive: true });
|
||||
const tmp = path.join(dir, `.${path.basename(p)}.${process.pid}.${crypto.randomBytes(4).toString('hex')}.tmp`);
|
||||
fs.writeFileSync(tmp, JSON.stringify(obj, null, 2) + '\n', { mode: 0o600 });
|
||||
fs.renameSync(tmp, p);
|
||||
}
|
||||
|
||||
// ── validateHookFields ────────────────────────────────────────────────────
|
||||
// Claude Code uses strict Zod on settings.json — a single malformed hook
|
||||
// silently discards the entire file. Mutate-to-valid before write.
|
||||
//
|
||||
// Required shape (per Claude Code docs):
|
||||
// settings.hooks[event] = [{ hooks: [{ type:'command', command:'…', timeout?:n }, ...] }, ...]
|
||||
// settings.hooks[event] = [{ matcher?:'…', hooks: [...] }, ...] // also valid
|
||||
function validateHookFields(settings) {
|
||||
if (!settings || typeof settings !== 'object') return settings;
|
||||
if (!settings.hooks || typeof settings.hooks !== 'object') return settings;
|
||||
for (const ev of Object.keys(settings.hooks)) {
|
||||
const arr = settings.hooks[ev];
|
||||
if (!Array.isArray(arr)) { delete settings.hooks[ev]; continue; }
|
||||
settings.hooks[ev] = arr.filter(entry => {
|
||||
if (!entry || typeof entry !== 'object') return false;
|
||||
if (!Array.isArray(entry.hooks)) return false;
|
||||
entry.hooks = entry.hooks.filter(h => {
|
||||
if (!h || typeof h !== 'object') return false;
|
||||
if (h.type === 'command') return typeof h.command === 'string' && h.command.length > 0;
|
||||
if (h.type === 'agent') return typeof h.prompt === 'string' && h.prompt.length > 0;
|
||||
return false;
|
||||
});
|
||||
return entry.hooks.length > 0;
|
||||
});
|
||||
if (settings.hooks[ev].length === 0) delete settings.hooks[ev];
|
||||
}
|
||||
if (Object.keys(settings.hooks).length === 0) delete settings.hooks;
|
||||
return settings;
|
||||
}
|
||||
|
||||
// ── Idempotency probe ──────────────────────────────────────────────────────
|
||||
function hasCavemanHook(settings, event, marker = 'caveman') {
|
||||
const arr = settings && settings.hooks && settings.hooks[event];
|
||||
if (!Array.isArray(arr)) return false;
|
||||
return arr.some(e =>
|
||||
e && Array.isArray(e.hooks) &&
|
||||
e.hooks.some(h => h && typeof h.command === 'string' && h.command.includes(marker))
|
||||
);
|
||||
}
|
||||
|
||||
// ── addCommandHook ────────────────────────────────────────────────────────
|
||||
// Idempotent push. `marker` defaults to opts.command — pass an explicit
|
||||
// shorter substring (e.g. the script basename) when the full command path
|
||||
// might rotate across reinstalls.
|
||||
function addCommandHook(settings, event, opts) {
|
||||
if (!settings.hooks) settings.hooks = {};
|
||||
if (!Array.isArray(settings.hooks[event])) settings.hooks[event] = [];
|
||||
const marker = opts.marker || opts.command;
|
||||
if (hasCavemanHook(settings, event, marker)) return false;
|
||||
const hook = { type: 'command', command: opts.command };
|
||||
if (typeof opts.timeout === 'number') hook.timeout = opts.timeout;
|
||||
if (typeof opts.statusMessage === 'string') hook.statusMessage = opts.statusMessage;
|
||||
settings.hooks[event].push({ hooks: [hook] });
|
||||
return true;
|
||||
}
|
||||
|
||||
// ── Managed hook scripts ──────────────────────────────────────────────────
|
||||
// The exact script basenames this installer wires into settings.json. Every
|
||||
// helper that decides "is this hook ours?" must match against these — never
|
||||
// against a bare "caveman" substring, which also matches user-authored hooks
|
||||
// that merely mention the word in a path (issue #593).
|
||||
const MANAGED_HOOK_BASENAMES = new Set([
|
||||
'caveman-activate.js',
|
||||
'caveman-mode-tracker.js',
|
||||
'caveman-stats.js',
|
||||
'caveman-statusline.sh',
|
||||
'caveman-statusline.ps1',
|
||||
]);
|
||||
|
||||
// Split a command into shell-ish tokens, honoring single/double quotes so a
|
||||
// path containing spaces survives intact. Good enough for hook commands we
|
||||
// generate (`node "/a/x.js"`, `"/abs/node" "/a/x.js"`, `bash /a/x.sh`); not
|
||||
// a full shell parser.
|
||||
function tokenizeCommand(command) {
|
||||
const out = [];
|
||||
const re = /"([^"]*)"|'([^']*)'|(\S+)/g;
|
||||
let m;
|
||||
while ((m = re.exec(command)) !== null) out.push(m[1] ?? m[2] ?? m[3]);
|
||||
return out;
|
||||
}
|
||||
|
||||
// True iff some token's BASENAME exactly equals a managed script name. Exact
|
||||
// match — not substring — so `mycaveman-activate.js` or a user hook living
|
||||
// under a `caveman-notes/` directory is never treated as ours. win32.basename
|
||||
// splits on both / and \ so a settings.json written on Windows still matches
|
||||
// when processed elsewhere.
|
||||
function referencesManagedScript(command) {
|
||||
try {
|
||||
for (const tok of tokenizeCommand(command)) {
|
||||
if (tok && typeof tok === 'string' && MANAGED_HOOK_BASENAMES.has(path.win32.basename(tok))) return true;
|
||||
}
|
||||
} catch (_) { /* malformed command — treat as not ours */ }
|
||||
return false;
|
||||
}
|
||||
|
||||
// ── removeCavemanHooks ────────────────────────────────────────────────────
|
||||
// Strip every entry whose any hook command targets one of our managed hook
|
||||
// scripts (exact basename match, see above). Empties events. Tolerates
|
||||
// malformed pre-existing settings (non-array hook lists, foreign shapes) —
|
||||
// those get dropped by validateHookFields first so we never call .length /
|
||||
// .filter on a non-array.
|
||||
function removeCavemanHooks(settings) {
|
||||
if (!settings || !settings.hooks) return 0;
|
||||
validateHookFields(settings);
|
||||
if (!settings.hooks) return 0; // validate may have deleted the whole tree
|
||||
let removed = 0;
|
||||
for (const ev of Object.keys(settings.hooks)) {
|
||||
if (!Array.isArray(settings.hooks[ev])) { delete settings.hooks[ev]; continue; }
|
||||
const before = settings.hooks[ev].length;
|
||||
settings.hooks[ev] = settings.hooks[ev].filter(entry => {
|
||||
if (!entry || !Array.isArray(entry.hooks)) return true;
|
||||
return !entry.hooks.some(h => h && typeof h.command === 'string' && referencesManagedScript(h.command));
|
||||
});
|
||||
removed += before - settings.hooks[ev].length;
|
||||
if (settings.hooks[ev].length === 0) delete settings.hooks[ev];
|
||||
}
|
||||
if (Object.keys(settings.hooks).length === 0) delete settings.hooks;
|
||||
return removed;
|
||||
}
|
||||
|
||||
// ── rewriteLegacyManagedHookCommands ──────────────────────────────────────
|
||||
// Walk every hook command. If it's a bare `node /path/to/<managed>.js` (no
|
||||
// absolute node path) and the basename is one of ours, rewrite to use
|
||||
// `absoluteNode` so GUI launchers with minimal PATH still find Node. Only
|
||||
// touches commands matching the exact bare-node shape — won't false-positive
|
||||
// on user-authored hooks that just happen to mention "caveman".
|
||||
function rewriteLegacyManagedHookCommands(settings, absoluteNode) {
|
||||
if (!settings || !settings.hooks || !absoluteNode) return 0;
|
||||
let rewritten = 0;
|
||||
const reBare = /^node\s+("([^"]+)"|'([^']+)'|(\S+))\s*$/;
|
||||
for (const ev of Object.keys(settings.hooks)) {
|
||||
// A hook event value that is an object/string (not an array) survives
|
||||
// JSONC parse untouched — this runs BEFORE validateHookFields in
|
||||
// installHooks, so it must tolerate malformed input itself rather than
|
||||
// assume the array shape. Pre-fix, `for...of` on a non-iterable object
|
||||
// threw a TypeError here and killed the installer mid-run (mirrors the
|
||||
// guard removeCavemanHooks already has).
|
||||
if (!Array.isArray(settings.hooks[ev])) continue;
|
||||
for (const entry of settings.hooks[ev]) {
|
||||
if (!entry || !Array.isArray(entry.hooks)) continue;
|
||||
for (const h of entry.hooks) {
|
||||
if (!h || typeof h.command !== 'string') continue;
|
||||
const m = reBare.exec(h.command);
|
||||
if (!m) continue;
|
||||
const scriptPath = m[2] || m[3] || m[4];
|
||||
const basename = path.basename(scriptPath);
|
||||
if (!MANAGED_HOOK_BASENAMES.has(basename)) continue;
|
||||
h.command = `"${absoluteNode}" "${scriptPath}"`;
|
||||
rewritten++;
|
||||
}
|
||||
}
|
||||
}
|
||||
return rewritten;
|
||||
}
|
||||
|
||||
// ── pruneOrphanedManagedHooks ─────────────────────────────────────────────
|
||||
// Remove managed hook entries whose target script no longer exists on disk.
|
||||
//
|
||||
// Migrating an old manual install (settings.json hooks → ~/.claude/hooks/
|
||||
// caveman-*.js) to the Claude Code plugin disables/renames those local
|
||||
// scripts but leaves the settings.json entries pointing at the now-missing
|
||||
// file. Claude Code then runs `node <missing>` every SessionStart /
|
||||
// UserPromptSubmit and crashes with `node:…/loader:1478 — Cannot find module
|
||||
// …caveman-activate.js` (issue #471). rewriteLegacyManagedHookCommands can't
|
||||
// help — it only matches the bare-node shape and these orphans are usually
|
||||
// absolute-node — and removeCavemanHooks runs only on uninstall.
|
||||
//
|
||||
// We extract the script path from any managed-looking command (bare- or
|
||||
// absolute-node, quoted or not), resolve it relative to dir if not absolute,
|
||||
// and drop the hook only when its target is genuinely absent. A managed hook
|
||||
// whose script still exists is left untouched, so this is safe to run on
|
||||
// every install.
|
||||
function pruneOrphanedManagedHooks(settings, configDir) {
|
||||
if (!settings || typeof settings !== 'object') return 0;
|
||||
const baseDir = configDir || claudeConfigDir();
|
||||
let removed = 0;
|
||||
|
||||
// A command is a missing managed target iff some token's BASENAME exactly
|
||||
// equals a managed script (exact match — not substring — so a user hook like
|
||||
// `mycaveman-activate.js` is never touched) and that resolved path is absent.
|
||||
// Relative paths resolve against configDir; honors CLAUDE_CONFIG_DIR. Wrapped
|
||||
// so a malformed command or fs error never throws out of the prune pass.
|
||||
const targetMissing = (command) => {
|
||||
try {
|
||||
for (const tok of tokenizeCommand(command)) {
|
||||
if (!tok || typeof tok !== 'string') continue;
|
||||
if (!MANAGED_HOOK_BASENAMES.has(path.basename(tok))) continue;
|
||||
const scriptPath = path.isAbsolute(tok) ? tok : path.join(baseDir, tok);
|
||||
return !fs.existsSync(scriptPath);
|
||||
}
|
||||
} catch (_) { /* silent-fail: never block install on a parse/fs hiccup */ }
|
||||
return false;
|
||||
};
|
||||
|
||||
if (settings.hooks && typeof settings.hooks === 'object') {
|
||||
// Normalize malformed shapes first so the filter below only sees valid
|
||||
// entries (and a poisoned settings.json can't survive the rewrite).
|
||||
validateHookFields(settings);
|
||||
}
|
||||
if (settings.hooks && typeof settings.hooks === 'object') {
|
||||
for (const ev of Object.keys(settings.hooks)) {
|
||||
if (!Array.isArray(settings.hooks[ev])) { delete settings.hooks[ev]; continue; }
|
||||
const before = settings.hooks[ev].length;
|
||||
settings.hooks[ev] = settings.hooks[ev].filter(entry => {
|
||||
if (!entry || typeof entry !== 'object' || !Array.isArray(entry.hooks)) return true;
|
||||
return !entry.hooks.some(h => h && typeof h.command === 'string' && targetMissing(h.command));
|
||||
});
|
||||
removed += before - settings.hooks[ev].length;
|
||||
if (settings.hooks[ev].length === 0) delete settings.hooks[ev];
|
||||
}
|
||||
if (Object.keys(settings.hooks).length === 0) delete settings.hooks;
|
||||
}
|
||||
|
||||
// statusLine lives outside settings.hooks. A managed statusline command
|
||||
// pointing at a missing script leaves a blank statusline (cosmetic, exits
|
||||
// clean) but is still stale — drop it so Claude Code falls back to default.
|
||||
if (settings.statusLine && typeof settings.statusLine.command === 'string'
|
||||
&& targetMissing(settings.statusLine.command)) {
|
||||
delete settings.statusLine;
|
||||
removed++;
|
||||
}
|
||||
|
||||
return removed;
|
||||
}
|
||||
|
||||
// ── claudeConfigDir ───────────────────────────────────────────────────────
|
||||
function claudeConfigDir() {
|
||||
if (process.env.CLAUDE_CONFIG_DIR) return process.env.CLAUDE_CONFIG_DIR;
|
||||
return path.join(os.homedir(), '.claude');
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
stripJsonComments,
|
||||
readSettings,
|
||||
writeSettings,
|
||||
validateHookFields,
|
||||
hasCavemanHook,
|
||||
addCommandHook,
|
||||
removeCavemanHooks,
|
||||
rewriteLegacyManagedHookCommands,
|
||||
pruneOrphanedManagedHooks,
|
||||
claudeConfigDir,
|
||||
MANAGED_HOOK_BASENAMES,
|
||||
};
|
||||
@@ -0,0 +1,5 @@
|
||||
---
|
||||
description: Generate terse caveman-style commit message
|
||||
---
|
||||
|
||||
Generate a terse commit message for the current staged changes. Conventional Commits format. Subject: ≤50 chars, imperative, lowercase after type. Body: only when 'why' isn't obvious from subject. Why over what. No period on subject.
|
||||
@@ -0,0 +1,2 @@
|
||||
description = "Generate terse caveman-style commit message"
|
||||
prompt = "Generate a terse commit message for the current staged changes. Conventional Commits format. Subject: ≤50 chars, imperative, lowercase after type. Body: only when 'why' isn't obvious from subject. Why over what. No period on subject."
|
||||
@@ -0,0 +1,13 @@
|
||||
---
|
||||
description: Drop the always-on caveman activation rule into the current repo for every IDE agent
|
||||
argument-hint: "[--dry-run|--force] [--only <agent>]"
|
||||
---
|
||||
|
||||
Write the per-repo caveman rule files (Cursor, Windsurf, Cline, Copilot, AGENTS.md) into the current repo, then report the result.
|
||||
|
||||
How to run the init script — pick the first that applies:
|
||||
|
||||
1. If `src/tools/caveman-init.js` exists in the current repo (you are inside a caveman checkout), run: `node src/tools/caveman-init.js $ARGUMENTS`
|
||||
2. Otherwise download and run the standalone script (it is self-contained and supports stdin execution): `curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/src/tools/caveman-init.js | node - $ARGUMENTS`
|
||||
|
||||
Use `--dry-run` first if the user did not pass `--force`, so we never silently overwrite an existing rule file.
|
||||
@@ -0,0 +1,2 @@
|
||||
description = "Drop the always-on caveman activation rule into the current repo for every IDE agent"
|
||||
prompt = "Write the per-repo caveman rule files into the current repo and report the result. If `src/tools/caveman-init.js` exists in the current repo (a caveman checkout), run `node src/tools/caveman-init.js {{args}}`. Otherwise run the standalone script (self-contained, supports stdin execution): `curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/src/tools/caveman-init.js | node - {{args}}`. Use --dry-run first if the user did not pass --force, so we never silently overwrite an existing rule file."
|
||||
@@ -0,0 +1,5 @@
|
||||
---
|
||||
description: One-line code review comments
|
||||
---
|
||||
|
||||
Review the current code changes. One-line per finding. Format: L<line>: <severity> <problem>. <fix>. Severity: bug, risk, nit, q. Skip praise. Skip obvious. If code look good, say 'LGTM' and stop.
|
||||
@@ -0,0 +1,2 @@
|
||||
description = "One-line code review comments"
|
||||
prompt = "Review the current code changes. One-line per finding. Format: L<line>: <severity> <problem>. <fix>. Severity: bug, risk, nit, q. Skip praise. Skip obvious. If code look good, say 'LGTM' and stop."
|
||||
@@ -0,0 +1,6 @@
|
||||
---
|
||||
description: Real session token usage + lifetime savings + USD. Tweetable line via --share.
|
||||
argument-hint: "[--share|--all|--since 7d]"
|
||||
---
|
||||
|
||||
/caveman-stats $ARGUMENTS
|
||||
@@ -0,0 +1,2 @@
|
||||
description = "Real session token usage + lifetime savings + USD. Tweetable line via --share."
|
||||
prompt = "/caveman-stats {{args}}"
|
||||
@@ -0,0 +1,6 @@
|
||||
---
|
||||
description: Switch caveman intensity level (lite/full/ultra/wenyan)
|
||||
argument-hint: "[lite|full|ultra|wenyan]"
|
||||
---
|
||||
|
||||
Switch to caveman $ARGUMENTS mode. If no level specified, use full. Respond terse like smart caveman — drop articles, filler, pleasantries. Fragments OK. Technical terms exact. Code unchanged. Pattern: [thing] [action] [reason]. [next step].
|
||||
@@ -0,0 +1,2 @@
|
||||
description = "Switch caveman intensity level (lite/full/ultra/wenyan)"
|
||||
prompt = "Switch to caveman {{args}} mode. If no level specified, use full. Respond terse like smart caveman — drop articles, filler, pleasantries. Fragments OK. Technical terms exact. Code unchanged. Pattern: [thing] [action] [reason]. [next step]."
|
||||
@@ -0,0 +1,47 @@
|
||||
# Honest Numbers
|
||||
|
||||
Caveman save tokens sometimes. Caveman cost tokens sometimes. This page say which is which, with the real numbers. No marketing. If caveman lose for your workload, this page tell you to turn it off.
|
||||
|
||||
## What caveman actually does
|
||||
|
||||
Caveman is a system-prompt skill. It makes the model **write shorter output**. That is the whole mechanism. It does not compress your input, your context, your files, or the model's thinking tokens.
|
||||
|
||||
## The measured numbers
|
||||
|
||||
| What | Number | How measured | Source |
|
||||
|---|---|---|---|
|
||||
| Output reduction vs default verbose replies | **65% average** (range 22–87%) | Real Claude API token counts, 10 prompts | [`benchmarks/`](../benchmarks/) |
|
||||
| Input reduction from the skill | **0%** | It's an output-style instruction | — |
|
||||
| Input cost the skill *adds* | **~1–1.5k tokens per turn** | SKILL.md rules (~5 KB) injected into context, plus skill-list entries | [`skills/caveman/SKILL.md`](../skills/caveman/SKILL.md) |
|
||||
| `/caveman-compress` on memory files | ~46% average input reduction, per session, for those files only | Real files, token counts in README table | [README](../README.md#benchmarks) |
|
||||
|
||||
These figures are output tokens only — the skill does not compress your input, your context, your files, or the model's thinking tokens. The full eval harness and its correction history are documented in [`evals/README.md`](../evals/README.md).
|
||||
|
||||
## When caveman wins
|
||||
|
||||
- **Long chatty outputs.** Explanations, architecture discussions, code review, docs, debugging walkthroughs — anywhere the model would write 1k+ output tokens per reply. This is where the 50–87% cuts happen.
|
||||
- **Long sessions with verbose agents.** The per-reply savings compound; the fixed ~1–1.5k/turn rule cost stays flat.
|
||||
- **Reading speed.** Shorter replies finish sooner and you read them faster. For many users this, not cost, is the real win.
|
||||
|
||||
## When caveman loses (net-negative)
|
||||
|
||||
Plainly: **the skill costs ~1–1.5k input tokens every turn. If it saves less output than that, you are paying to use it.**
|
||||
|
||||
- **Terse coding Q&A** ([#145](https://github.com/JuliusBrussee/caveman/issues/145)). If your normal replies are ~150 output tokens, caveman saves maybe 70–100 of them and costs ~1k+ of input overhead per turn. Net loss. The user in #145 measured exactly this. They were right.
|
||||
- **Agents that bill by request or credit, not tokens** ([#506](https://github.com/JuliusBrussee/caveman/issues/506)). GitHub Copilot charges premium *requests*. A shorter answer is the same request. Caveman cannot lower your Copilot credit use. Same logic for any per-message pricing.
|
||||
- **Session-level totals** are always smaller than the output-reduction headline, because input tokens (your prompts, your context, your files, the injected rules) dwarf output tokens in agentic coding. Independent session-level measurements land around **14–21% total savings** on output-heavy workloads — and below zero on terse ones.
|
||||
- **Some tool-side counters go the wrong way** ([#550](https://github.com/JuliusBrussee/caveman/issues/550)). One Cursor A/B showed 4.3M tokens with caveman vs 1M without, and double the wall-clock time. We could not reproduce the exact run, but the honest reading is: rule re-injection, retries, and cache/context accounting can swamp output savings in some agents. If your A/B looks like that, caveman is net-negative for you. Turn it off. Wanting the rock to work does not make the rock work.
|
||||
|
||||
## Measure it yourself
|
||||
|
||||
1. **`/caveman-stats`** (Claude Code) reads your real session log and prints actual input/output token counts. The "saved" line is an **estimate**: it extrapolates what the output would have been without caveman using the benchmark ratio. Real usage, estimated baseline — the output labels it `est.` for exactly that reason.
|
||||
2. **The only fully honest test is an A/B**: run the same task with and without caveman and compare your provider's own usage/billing page. That number outranks anything this repo prints.
|
||||
3. **Reproduce our numbers**: `benchmarks/run.py` (needs an Anthropic key) and `evals/measure.py` (offline, reads the committed snapshot).
|
||||
|
||||
## Rule of thumb
|
||||
|
||||
> Normal reply longer than ~1.5–2k output tokens → caveman probably saves you money.
|
||||
> Normal reply shorter than that, or you pay per request → caveman probably costs you money.
|
||||
> Either way, caveman replies faster to read. That part is free.
|
||||
|
||||
Found a workload where our numbers are wrong? [Open an issue](https://github.com/JuliusBrussee/caveman/issues) with the A/B. We will put it on this page.
|
||||
@@ -0,0 +1,12 @@
|
||||
<svg width="163" height="26" viewBox="0 0 163 26" fill="none" xmlns="http://www.w3.org/2000/svg">
|
||||
<path d="M32.9475 25.7973C30.9995 25.7973 29.4793 25.2568 28.3984 24.1871C27.3174 23.1174 26.7769 21.6085 26.7769 19.6492V12.0148H23.8154V8.28766H24.1307C24.9752 8.28766 25.6283 8.06245 26.09 7.6233C26.5404 7.17289 26.7769 6.53106 26.7769 5.68654V4.34657H30.9769V8.28766H34.9518V12.0148H30.9769V19.4353C30.9769 20.0096 31.0783 20.4938 31.281 20.8991C31.4837 21.3045 31.7989 21.6085 32.2381 21.8225C32.6772 22.0364 33.229 22.1378 33.9046 22.1378C34.051 22.1378 34.2312 22.1378 34.4338 22.104C34.6365 22.0814 34.828 22.0589 35.0194 22.0364V25.6059C34.7266 25.6509 34.3775 25.696 34.006 25.7298C33.6231 25.7748 33.274 25.7973 32.9588 25.7973H32.9475Z" fill="#FFFFFF"/>
|
||||
<path d="M36.5732 25.6059V1.52028H40.7733V25.6059H36.5732Z" fill="#FFFFFF"/>
|
||||
<path d="M48.2839 25.9888C47.079 25.9888 46.0206 25.7861 45.1197 25.3807C44.2189 24.9753 43.5208 24.4011 43.0366 23.6466C42.5524 22.8922 42.3047 22.0139 42.3047 21.023C42.3047 20.0321 42.5186 19.2101 42.9578 18.4557C43.3969 17.7012 44.05 17.0706 44.9508 16.5639C45.8404 16.0572 46.9664 15.6969 48.3289 15.483L53.959 14.5596V17.7463L49.1171 18.602C48.2951 18.7484 47.6758 19.0074 47.2705 19.379C46.8651 19.7506 46.6624 20.246 46.6624 20.8541C46.6624 21.4621 46.8876 21.9238 47.3493 22.2729C47.7997 22.6219 48.374 22.8021 49.0496 22.8021C49.9166 22.8021 50.6936 22.6219 51.3579 22.2504C52.0223 21.8788 52.5403 21.3608 52.9006 20.7077C53.2609 20.0546 53.4411 19.334 53.4411 18.5795V14.0867C53.4411 13.3435 53.1596 12.7242 52.5853 12.2287C52.011 11.7333 51.2453 11.4856 50.2995 11.4856C49.4099 11.4856 48.6217 11.7333 47.9236 12.2175C47.2367 12.7017 46.73 13.3322 46.4147 14.0979L43.0141 12.4427C43.3519 11.5306 43.8924 10.7424 44.6243 10.0668C45.3562 9.39116 46.2233 8.87319 47.2142 8.49034C48.2163 8.10749 49.2973 7.91607 50.4571 7.91607C51.8759 7.91607 53.1258 8.17505 54.2068 8.69303C55.2878 9.211 56.1323 9.94291 56.7403 10.8775C57.3484 11.8121 57.6524 12.8818 57.6524 14.0867V25.6059H53.7113V22.6445L54.6009 22.6107C54.1505 23.3313 53.6212 23.9507 52.9907 24.4574C52.3601 24.9641 51.662 25.3469 50.885 25.6059C50.108 25.8649 49.241 25.9888 48.2951 25.9888H48.2839Z" fill="#FFFFFF"/>
|
||||
<path d="M66.4802 25.9888C64.6336 25.9888 63.0233 25.5496 61.6609 24.6713C60.2871 23.793 59.3412 22.5994 58.812 21.0906L61.9649 19.5929C62.4153 20.5726 63.0346 21.3383 63.8228 21.89C64.6223 22.4418 65.5006 22.712 66.4802 22.712C67.2234 22.712 67.8202 22.5431 68.2594 22.2053C68.7098 21.8675 68.9237 21.4171 68.9237 20.8653C68.9237 20.5275 68.8336 20.246 68.6535 20.0208C68.4733 19.7956 68.2368 19.6042 67.9328 19.4466C67.64 19.2889 67.291 19.1538 66.9194 19.0524L64.0818 18.253C62.6405 17.8476 61.537 17.2058 60.7826 16.3275C60.0281 15.4492 59.6565 14.4132 59.6565 13.2196C59.6565 12.1612 59.9268 11.2266 60.4673 10.4384C61.0078 9.65015 61.7622 9.01957 62.7306 8.58042C63.699 8.13001 64.8025 7.91607 66.0524 7.91607C67.6851 7.91607 69.1264 8.31018 70.3763 9.09839C71.6262 9.88661 72.5157 10.9901 73.045 12.4089L69.8583 13.9065C69.5656 13.1183 69.0588 12.499 68.3607 12.0486C67.6626 11.5982 66.8743 11.3617 66.0073 11.3617C65.3092 11.3617 64.7574 11.5193 64.3521 11.8234C63.9467 12.1274 63.744 12.5553 63.744 13.0845C63.744 13.3773 63.8341 13.6475 64.003 13.884C64.1719 14.1205 64.4084 14.3119 64.7236 14.4583C65.0277 14.6047 65.388 14.7398 65.7934 14.8749L68.5634 15.6969C69.9822 16.1248 71.0857 16.7554 71.8626 17.6111C72.6396 18.4557 73.0224 19.5029 73.0224 20.7302C73.0224 21.7662 72.7409 22.6895 72.2005 23.4777C71.6487 24.2772 70.883 24.8965 69.9146 25.3357C68.935 25.7861 67.7977 26 66.4802 26V25.9888Z" fill="#FFFFFF"/>
|
||||
<path d="M90.2282 25.9889C88.5279 25.9889 86.9627 25.6849 85.5214 25.0656C84.0801 24.4463 82.8302 23.5905 81.7717 22.487C80.7133 21.3835 79.88 20.0886 79.2719 18.6022C78.6639 17.1159 78.3599 15.4944 78.3599 13.7378C78.3599 11.9812 78.6526 10.3485 79.2494 8.85085C79.8462 7.35324 80.6795 6.05831 81.7492 4.96606C82.8189 3.87382 84.0688 3.0293 85.4989 2.42125C86.9289 1.8132 88.5053 1.50917 90.2282 1.50917C91.951 1.50917 93.4486 1.79068 94.7998 2.36495C96.1511 2.93922 97.2884 3.69366 98.223 4.63952C99.1576 5.58538 99.8219 6.62132 100.227 7.74735L96.3425 9.59403C95.8921 8.38918 95.1489 7.39828 94.0792 6.62132C93.0207 5.84436 91.737 5.46152 90.2282 5.46152C88.7193 5.46152 87.4356 5.81058 86.2983 6.50872C85.1611 7.20685 84.2828 8.17523 83.6522 9.4026C83.0216 10.63 82.7176 12.0713 82.7176 13.7265C82.7176 15.3818 83.0329 16.8344 83.6522 18.073C84.2828 19.3116 85.1611 20.28 86.2983 20.9894C87.4356 21.6875 88.7418 22.0366 90.2282 22.0366C91.7145 22.0366 93.0207 21.6537 94.0792 20.8768C95.1376 20.0998 95.8921 19.1202 96.3425 17.9379L100.227 19.7508C99.8219 20.8768 99.1576 21.9127 98.223 22.8586C97.2884 23.8045 96.1511 24.5589 94.7998 25.1332C93.4486 25.7074 91.9285 25.9889 90.2282 25.9889Z" fill="#FFFFFF"/>
|
||||
<path d="M101.748 25.6059V1.52028H105.948V25.6059H101.748Z" fill="#FFFFFF"/>
|
||||
<path d="M116.645 25.9884C114.968 25.9884 113.436 25.5943 112.051 24.8061C110.666 24.0178 109.551 22.9481 108.729 21.5969C107.907 20.2344 107.491 18.6917 107.491 16.9464C107.491 15.2011 107.907 13.6584 108.729 12.2959C109.551 10.9334 110.655 9.86371 112.04 9.08676C113.414 8.29854 114.956 7.90443 116.657 7.90443C118.357 7.90443 119.922 8.29854 121.307 9.08676C122.681 9.87497 123.784 10.9334 124.606 12.2847C125.417 13.6359 125.834 15.1898 125.834 16.9464C125.834 18.703 125.417 20.2344 124.595 21.5969C123.773 22.9594 122.67 24.0291 121.285 24.8061C119.911 25.5943 118.368 25.9884 116.668 25.9884H116.645ZM116.645 22.1712C117.602 22.1712 118.436 21.946 119.145 21.5068C119.854 21.0564 120.417 20.4371 120.834 19.6489C121.251 18.8494 121.453 17.9598 121.453 16.9577C121.453 15.9555 121.251 15.0659 120.834 14.289C120.417 13.5008 119.854 12.8927 119.145 12.4423C118.436 11.9919 117.602 11.778 116.645 11.778C115.688 11.778 114.889 12.0032 114.168 12.4423C113.447 12.8927 112.884 13.5008 112.468 14.289C112.051 15.0772 111.848 15.9668 111.848 16.9577C111.848 17.9486 112.051 18.8494 112.468 19.6489C112.884 20.4483 113.447 21.0677 114.168 21.5068C114.889 21.9572 115.722 22.1712 116.645 22.1712Z" fill="#FFFFFF"/>
|
||||
<path d="M133.535 25.9884C132.172 25.9884 131.013 25.6956 130.033 25.0988C129.053 24.502 128.299 23.68 127.77 22.6215C127.24 21.5631 126.97 20.3245 126.97 18.8944V8.29852H131.17V18.5453C131.17 19.266 131.317 19.8966 131.598 20.4371C131.88 20.9775 132.296 21.4054 132.837 21.7095C133.377 22.0135 133.985 22.1711 134.672 22.1711C135.359 22.1711 135.956 22.0135 136.485 21.7095C137.014 21.4054 137.431 20.9775 137.724 20.4258C138.017 19.874 138.174 19.221 138.174 18.4553V8.29852H142.34V25.6055H138.399V22.2049L138.715 22.813C138.309 23.8714 137.656 24.6709 136.744 25.2001C135.832 25.7294 134.762 25.9996 133.535 25.9996V25.9884Z" fill="#FFFFFF"/>
|
||||
<path d="M152.7 25.9888C151.022 25.9888 149.525 25.5947 148.196 24.7952C146.867 23.9957 145.82 22.9147 145.066 21.5297C144.3 20.156 143.917 18.6246 143.917 16.9468C143.917 15.269 144.3 13.7264 145.077 12.3639C145.854 11.0014 146.901 9.92042 148.207 9.12094C149.525 8.31021 151.011 7.9161 152.666 7.9161C153.984 7.9161 155.155 8.17508 156.179 8.69305C157.204 9.21103 158.015 9.94294 158.612 10.8775L157.97 11.7333V1.52028H162.136V25.6059H158.195V22.2616L158.645 23.0836C158.049 24.0408 157.227 24.7614 156.168 25.2456C155.11 25.7298 153.95 25.9775 152.7 25.9775V25.9888ZM153.139 22.1716C154.074 22.1716 154.907 21.9464 155.639 21.5072C156.371 21.0568 156.945 20.4487 157.362 19.6605C157.778 18.8723 157.981 17.9715 157.981 16.9581C157.981 15.9447 157.778 15.0664 157.362 14.2894C156.945 13.5012 156.371 12.8931 155.639 12.4427C154.907 11.9923 154.074 11.7784 153.139 11.7784C152.205 11.7784 151.371 12.0036 150.628 12.4427C149.885 12.8931 149.311 13.5012 148.894 14.2894C148.477 15.0776 148.275 15.9672 148.275 16.9581C148.275 17.949 148.477 18.8836 148.894 19.6605C149.311 20.4487 149.885 21.0568 150.628 21.5072C151.371 21.9576 152.205 22.1716 153.139 22.1716Z" fill="#FFFFFF"/>
|
||||
<path d="M13.4447 1.71661e-05L0 25.9887C6.22692 23.5227 11.249 23.1623 15.7643 23.3763L13.7037 18.8159C12.8029 18.7258 10.1905 18.7258 8.9519 19.0523L13.4447 9.06451C13.4447 9.06451 20.2009 23.7366 20.2121 23.7366C21.5183 23.9393 24.9977 25.1104 26.8895 25.9887L13.4447 1.71661e-05Z" fill="#FFFFFF"/>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 8.1 KiB |
@@ -0,0 +1,12 @@
|
||||
<svg width="163" height="26" viewBox="0 0 163 26" fill="none" xmlns="http://www.w3.org/2000/svg">
|
||||
<path d="M32.9475 25.7973C30.9995 25.7973 29.4793 25.2568 28.3984 24.1871C27.3174 23.1174 26.7769 21.6085 26.7769 19.6492V12.0148H23.8154V8.28766H24.1307C24.9752 8.28766 25.6283 8.06245 26.09 7.6233C26.5404 7.17289 26.7769 6.53106 26.7769 5.68654V4.34657H30.9769V8.28766H34.9518V12.0148H30.9769V19.4353C30.9769 20.0096 31.0783 20.4938 31.281 20.8991C31.4837 21.3045 31.7989 21.6085 32.2381 21.8225C32.6772 22.0364 33.229 22.1378 33.9046 22.1378C34.051 22.1378 34.2312 22.1378 34.4338 22.104C34.6365 22.0814 34.828 22.0589 35.0194 22.0364V25.6059C34.7266 25.6509 34.3775 25.696 34.006 25.7298C33.6231 25.7748 33.274 25.7973 32.9588 25.7973H32.9475Z" fill="#0F1111"/>
|
||||
<path d="M36.5732 25.6059V1.52028H40.7733V25.6059H36.5732Z" fill="#0F1111"/>
|
||||
<path d="M48.2839 25.9888C47.079 25.9888 46.0206 25.7861 45.1197 25.3807C44.2189 24.9753 43.5208 24.4011 43.0366 23.6466C42.5524 22.8922 42.3047 22.0139 42.3047 21.023C42.3047 20.0321 42.5186 19.2101 42.9578 18.4557C43.3969 17.7012 44.05 17.0706 44.9508 16.5639C45.8404 16.0572 46.9664 15.6969 48.3289 15.483L53.959 14.5596V17.7463L49.1171 18.602C48.2951 18.7484 47.6758 19.0074 47.2705 19.379C46.8651 19.7506 46.6624 20.246 46.6624 20.8541C46.6624 21.4621 46.8876 21.9238 47.3493 22.2729C47.7997 22.6219 48.374 22.8021 49.0496 22.8021C49.9166 22.8021 50.6936 22.6219 51.3579 22.2504C52.0223 21.8788 52.5403 21.3608 52.9006 20.7077C53.2609 20.0546 53.4411 19.334 53.4411 18.5795V14.0867C53.4411 13.3435 53.1596 12.7242 52.5853 12.2287C52.011 11.7333 51.2453 11.4856 50.2995 11.4856C49.4099 11.4856 48.6217 11.7333 47.9236 12.2175C47.2367 12.7017 46.73 13.3322 46.4147 14.0979L43.0141 12.4427C43.3519 11.5306 43.8924 10.7424 44.6243 10.0668C45.3562 9.39116 46.2233 8.87319 47.2142 8.49034C48.2163 8.10749 49.2973 7.91607 50.4571 7.91607C51.8759 7.91607 53.1258 8.17505 54.2068 8.69303C55.2878 9.211 56.1323 9.94291 56.7403 10.8775C57.3484 11.8121 57.6524 12.8818 57.6524 14.0867V25.6059H53.7113V22.6445L54.6009 22.6107C54.1505 23.3313 53.6212 23.9507 52.9907 24.4574C52.3601 24.9641 51.662 25.3469 50.885 25.6059C50.108 25.8649 49.241 25.9888 48.2951 25.9888H48.2839Z" fill="#0F1111"/>
|
||||
<path d="M66.4802 25.9888C64.6336 25.9888 63.0233 25.5496 61.6609 24.6713C60.2871 23.793 59.3412 22.5994 58.812 21.0906L61.9649 19.5929C62.4153 20.5726 63.0346 21.3383 63.8228 21.89C64.6223 22.4418 65.5006 22.712 66.4802 22.712C67.2234 22.712 67.8202 22.5431 68.2594 22.2053C68.7098 21.8675 68.9237 21.4171 68.9237 20.8653C68.9237 20.5275 68.8336 20.246 68.6535 20.0208C68.4733 19.7956 68.2368 19.6042 67.9328 19.4466C67.64 19.2889 67.291 19.1538 66.9194 19.0524L64.0818 18.253C62.6405 17.8476 61.537 17.2058 60.7826 16.3275C60.0281 15.4492 59.6565 14.4132 59.6565 13.2196C59.6565 12.1612 59.9268 11.2266 60.4673 10.4384C61.0078 9.65015 61.7622 9.01957 62.7306 8.58042C63.699 8.13001 64.8025 7.91607 66.0524 7.91607C67.6851 7.91607 69.1264 8.31018 70.3763 9.09839C71.6262 9.88661 72.5157 10.9901 73.045 12.4089L69.8583 13.9065C69.5656 13.1183 69.0588 12.499 68.3607 12.0486C67.6626 11.5982 66.8743 11.3617 66.0073 11.3617C65.3092 11.3617 64.7574 11.5193 64.3521 11.8234C63.9467 12.1274 63.744 12.5553 63.744 13.0845C63.744 13.3773 63.8341 13.6475 64.003 13.884C64.1719 14.1205 64.4084 14.3119 64.7236 14.4583C65.0277 14.6047 65.388 14.7398 65.7934 14.8749L68.5634 15.6969C69.9822 16.1248 71.0857 16.7554 71.8626 17.6111C72.6396 18.4557 73.0224 19.5029 73.0224 20.7302C73.0224 21.7662 72.7409 22.6895 72.2005 23.4777C71.6487 24.2772 70.883 24.8965 69.9146 25.3357C68.935 25.7861 67.7977 26 66.4802 26V25.9888Z" fill="#0F1111"/>
|
||||
<path d="M90.2282 25.9889C88.5279 25.9889 86.9627 25.6849 85.5214 25.0656C84.0801 24.4463 82.8302 23.5905 81.7717 22.487C80.7133 21.3835 79.88 20.0886 79.2719 18.6022C78.6639 17.1159 78.3599 15.4944 78.3599 13.7378C78.3599 11.9812 78.6526 10.3485 79.2494 8.85085C79.8462 7.35324 80.6795 6.05831 81.7492 4.96606C82.8189 3.87382 84.0688 3.0293 85.4989 2.42125C86.9289 1.8132 88.5053 1.50917 90.2282 1.50917C91.951 1.50917 93.4486 1.79068 94.7998 2.36495C96.1511 2.93922 97.2884 3.69366 98.223 4.63952C99.1576 5.58538 99.8219 6.62132 100.227 7.74735L96.3425 9.59403C95.8921 8.38918 95.1489 7.39828 94.0792 6.62132C93.0207 5.84436 91.737 5.46152 90.2282 5.46152C88.7193 5.46152 87.4356 5.81058 86.2983 6.50872C85.1611 7.20685 84.2828 8.17523 83.6522 9.4026C83.0216 10.63 82.7176 12.0713 82.7176 13.7265C82.7176 15.3818 83.0329 16.8344 83.6522 18.073C84.2828 19.3116 85.1611 20.28 86.2983 20.9894C87.4356 21.6875 88.7418 22.0366 90.2282 22.0366C91.7145 22.0366 93.0207 21.6537 94.0792 20.8768C95.1376 20.0998 95.8921 19.1202 96.3425 17.9379L100.227 19.7508C99.8219 20.8768 99.1576 21.9127 98.223 22.8586C97.2884 23.8045 96.1511 24.5589 94.7998 25.1332C93.4486 25.7074 91.9285 25.9889 90.2282 25.9889Z" fill="#0F1111"/>
|
||||
<path d="M101.748 25.6059V1.52028H105.948V25.6059H101.748Z" fill="#0F1111"/>
|
||||
<path d="M116.645 25.9884C114.968 25.9884 113.436 25.5943 112.051 24.8061C110.666 24.0178 109.551 22.9481 108.729 21.5969C107.907 20.2344 107.491 18.6917 107.491 16.9464C107.491 15.2011 107.907 13.6584 108.729 12.2959C109.551 10.9334 110.655 9.86371 112.04 9.08676C113.414 8.29854 114.956 7.90443 116.657 7.90443C118.357 7.90443 119.922 8.29854 121.307 9.08676C122.681 9.87497 123.784 10.9334 124.606 12.2847C125.417 13.6359 125.834 15.1898 125.834 16.9464C125.834 18.703 125.417 20.2344 124.595 21.5969C123.773 22.9594 122.67 24.0291 121.285 24.8061C119.911 25.5943 118.368 25.9884 116.668 25.9884H116.645ZM116.645 22.1712C117.602 22.1712 118.436 21.946 119.145 21.5068C119.854 21.0564 120.417 20.4371 120.834 19.6489C121.251 18.8494 121.453 17.9598 121.453 16.9577C121.453 15.9555 121.251 15.0659 120.834 14.289C120.417 13.5008 119.854 12.8927 119.145 12.4423C118.436 11.9919 117.602 11.778 116.645 11.778C115.688 11.778 114.889 12.0032 114.168 12.4423C113.447 12.8927 112.884 13.5008 112.468 14.289C112.051 15.0772 111.848 15.9668 111.848 16.9577C111.848 17.9486 112.051 18.8494 112.468 19.6489C112.884 20.4483 113.447 21.0677 114.168 21.5068C114.889 21.9572 115.722 22.1712 116.645 22.1712Z" fill="#0F1111"/>
|
||||
<path d="M133.535 25.9884C132.172 25.9884 131.013 25.6956 130.033 25.0988C129.053 24.502 128.299 23.68 127.77 22.6215C127.24 21.5631 126.97 20.3245 126.97 18.8944V8.29852H131.17V18.5453C131.17 19.266 131.317 19.8966 131.598 20.4371C131.88 20.9775 132.296 21.4054 132.837 21.7095C133.377 22.0135 133.985 22.1711 134.672 22.1711C135.359 22.1711 135.956 22.0135 136.485 21.7095C137.014 21.4054 137.431 20.9775 137.724 20.4258C138.017 19.874 138.174 19.221 138.174 18.4553V8.29852H142.34V25.6055H138.399V22.2049L138.715 22.813C138.309 23.8714 137.656 24.6709 136.744 25.2001C135.832 25.7294 134.762 25.9996 133.535 25.9996V25.9884Z" fill="#0F1111"/>
|
||||
<path d="M152.7 25.9888C151.022 25.9888 149.525 25.5947 148.196 24.7952C146.867 23.9957 145.82 22.9147 145.066 21.5297C144.3 20.156 143.917 18.6246 143.917 16.9468C143.917 15.269 144.3 13.7264 145.077 12.3639C145.854 11.0014 146.901 9.92042 148.207 9.12094C149.525 8.31021 151.011 7.9161 152.666 7.9161C153.984 7.9161 155.155 8.17508 156.179 8.69305C157.204 9.21103 158.015 9.94294 158.612 10.8775L157.97 11.7333V1.52028H162.136V25.6059H158.195V22.2616L158.645 23.0836C158.049 24.0408 157.227 24.7614 156.168 25.2456C155.11 25.7298 153.95 25.9775 152.7 25.9775V25.9888ZM153.139 22.1716C154.074 22.1716 154.907 21.9464 155.639 21.5072C156.371 21.0568 156.945 20.4487 157.362 19.6605C157.778 18.8723 157.981 17.9715 157.981 16.9581C157.981 15.9447 157.778 15.0664 157.362 14.2894C156.945 13.5012 156.371 12.8931 155.639 12.4427C154.907 11.9923 154.074 11.7784 153.139 11.7784C152.205 11.7784 151.371 12.0036 150.628 12.4427C149.885 12.8931 149.311 13.5012 148.894 14.2894C148.477 15.0776 148.275 15.9672 148.275 16.9581C148.275 17.949 148.477 18.8836 148.894 19.6605C149.311 20.4487 149.885 21.0568 150.628 21.5072C151.371 21.9576 152.205 22.1716 153.139 22.1716Z" fill="#0F1111"/>
|
||||
<path d="M13.4447 1.71661e-05L0 25.9887C6.22692 23.5227 11.249 23.1623 15.7643 23.3763L13.7037 18.8159C12.8029 18.7258 10.1905 18.7258 8.9519 19.0523L13.4447 9.06451C13.4447 9.06451 20.2009 23.7366 20.2121 23.7366C21.5183 23.9393 24.9977 25.1104 26.8895 25.9887L13.4447 1.71661e-05Z" fill="#0F1111"/>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 8.1 KiB |
|
After Width: | Height: | Size: 19 KiB |
|
After Width: | Height: | Size: 117 B |
|
After Width: | Height: | Size: 16 KiB |
|
After Width: | Height: | Size: 338 KiB |
@@ -33,7 +33,7 @@ body {
|
||||
background-color: var(--bg); color: var(--text-primary);
|
||||
font-family: 'Inter', -apple-system, BlinkMacSystemFont, sans-serif;
|
||||
overflow-x: hidden; -webkit-font-smoothing: antialiased; letter-spacing: -0.02em;
|
||||
cursor: url("data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' width='24' height='24'%3E%3Ctext y='20' font-size='20'%3E🪨%3C/text%3E%3C/svg%3E"), auto;
|
||||
cursor: url("assets/dancing-rock-32.png") 16 16, auto;
|
||||
}
|
||||
|
||||
::selection { background: var(--text-primary); color: var(--bg); }
|
||||
@@ -208,7 +208,7 @@ body {
|
||||
<div class="telemetry-widget">
|
||||
<div>STATUS: <span class="telemetry-val" style="color:#4ade80">OPTIMIZED</span></div>
|
||||
<div>PHYSICAL_IMPACTS: <span class="telemetry-val" id="bonkCount">0</span></div>
|
||||
<div>TOKENS_PURGED: <span class="telemetry-val">~75%</span></div>
|
||||
<div>TOKENS_PURGED: <span class="telemetry-val">65%</span></div>
|
||||
</div>
|
||||
|
||||
<section class="hero container">
|
||||
@@ -363,7 +363,7 @@ document.addEventListener('mousemove', (e) => {
|
||||
|
||||
// --- Marquee Population ---
|
||||
const items = [
|
||||
{ k: "TOKENS_SAVED", v: "75%" }, { k: "ACCURACY", v: "100%" },
|
||||
{ k: "TOKENS_SAVED", v: "65%" }, { k: "ACCURACY", v: "100%" },
|
||||
{ k: "LATENCY_DROP", v: "3x" }, { k: "VIBES", v: "OOG" },
|
||||
{ k: "PRICE", v: "$0.00" }, { k: "DEPENDENCIES", v: "NONE" }
|
||||
];
|
||||
@@ -417,7 +417,17 @@ cliInput.addEventListener('keydown', (e) => {
|
||||
const val = cliInput.value.trim();
|
||||
const row = document.createElement('div');
|
||||
row.className = 'term-line';
|
||||
row.innerHTML = `<span class="term-accent">❯</span> ${val}`;
|
||||
|
||||
// SECURITY FIX: Create DOM elements safely to prevent XSS
|
||||
const prompt = document.createElement('span');
|
||||
prompt.className = 'term-accent';
|
||||
prompt.textContent = '❯';
|
||||
|
||||
const commandText = document.createElement('span');
|
||||
commandText.textContent = ` ${val}`; // textContent automatically escapes HTML
|
||||
|
||||
row.appendChild(prompt);
|
||||
row.appendChild(commandText);
|
||||
cliInput.parentElement.before(row);
|
||||
|
||||
const res = document.createElement('div');
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
# Windows install fallback
|
||||
|
||||
If `irm https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.ps1 | iex` fails on Windows (issues #249, #199, #72), set up plugin-skill activation by hand. This does **not** install the standalone hooks or the statusline — for those, run the unified Node installer afterwards: `npx -y github:JuliusBrussee/caveman -- --only claude` (or `node cli/install.js --only claude` from a clone).
|
||||
|
||||
```powershell
|
||||
$ClaudeDir = if ($env:CLAUDE_CONFIG_DIR) { $env:CLAUDE_CONFIG_DIR } else { Join-Path $HOME ".claude" }
|
||||
$PluginSkillDir = Join-Path $ClaudeDir ".agents\plugins\caveman\skills\caveman"
|
||||
$MarketplaceDir = Join-Path $ClaudeDir ".agents\plugins"
|
||||
$MarketplaceFile = Join-Path $MarketplaceDir "marketplace.json"
|
||||
|
||||
# Copy SKILL.md into the plugin path (run from a clone of the repo)
|
||||
New-Item -ItemType Directory -Path $PluginSkillDir -Force | Out-Null
|
||||
Copy-Item ".\skills\caveman\SKILL.md" "$PluginSkillDir\SKILL.md" -Force
|
||||
|
||||
# Create or update marketplace.json with the caveman entry
|
||||
New-Item -ItemType Directory -Path $MarketplaceDir -Force | Out-Null
|
||||
if (Test-Path $MarketplaceFile) {
|
||||
$marketplace = Get-Content $MarketplaceFile -Raw | ConvertFrom-Json
|
||||
} else {
|
||||
$marketplace = [pscustomobject]@{}
|
||||
}
|
||||
if (-not ($marketplace.PSObject.Properties.Name -contains "plugins")) {
|
||||
$marketplace | Add-Member -NotePropertyName plugins -NotePropertyValue ([pscustomobject]@{})
|
||||
}
|
||||
$plugins = [ordered]@{}
|
||||
foreach ($p in $marketplace.plugins.PSObject.Properties) { $plugins[$p.Name] = $p.Value }
|
||||
$plugins["caveman"] = [ordered]@{ name = "caveman"; source = "JuliusBrussee/caveman"; version = "main" }
|
||||
$marketplace.plugins = [pscustomobject]$plugins
|
||||
$marketplace | ConvertTo-Json -Depth 10 | Set-Content -Path $MarketplaceFile -Encoding UTF8
|
||||
```
|
||||
|
||||
Verify: `Test-Path "$PluginSkillDir\SKILL.md"` should print `True`. Restart Claude Code, then run `/caveman` to confirm the skill loads.
|
||||
|
||||
## Codex on Windows
|
||||
|
||||
1. Enable symlinks first: `git config --global core.symlinks true` (requires Developer Mode or admin).
|
||||
2. Clone repo → Open VS Code → Codex Settings → Plugins → find "Caveman" under the local marketplace → Install → Reload Window.
|
||||
3. Codex hooks are currently disabled on Windows, so use `$caveman` to start the mode manually each session.
|
||||
|
||||
## `npx skills` symlink fallback
|
||||
|
||||
`npx skills` uses symlinks by default. If symlinks fail, add `--copy`:
|
||||
|
||||
```powershell
|
||||
npx skills add JuliusBrussee/caveman --copy
|
||||
```
|
||||
|
||||
## Want it always on (any agent)?
|
||||
|
||||
Paste this into the agent's system prompt or rules file:
|
||||
|
||||
```
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging.
|
||||
Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step].
|
||||
ACTIVE EVERY RESPONSE. No revert after many turns. No filler drift.
|
||||
Code/commits/PRs: normal. Off: "stop caveman" / "normal mode".
|
||||
```
|
||||
@@ -0,0 +1,84 @@
|
||||
# Evals
|
||||
|
||||
Measures real token compression of caveman skills by running the same
|
||||
prompts through Claude Code under three conditions and comparing the
|
||||
generated output token counts.
|
||||
|
||||
## The three arms
|
||||
|
||||
| Arm | System prompt |
|
||||
|-----|--------------|
|
||||
| `__baseline__` | none |
|
||||
| `__terse__` | `Answer concisely.` |
|
||||
| `<skill>` | `Answer concisely.\n\n{SKILL.md}` |
|
||||
|
||||
The honest delta for any skill is **`<skill>` vs `__terse__`** — i.e.
|
||||
how much the skill itself adds on top of a plain "be terse" instruction.
|
||||
Comparing a skill to the no-system-prompt baseline conflates the skill
|
||||
with the generic terseness ask, which is what an earlier version of
|
||||
this harness did and is why its numbers were inflated.
|
||||
|
||||
## Why this design
|
||||
|
||||
- **Real LLM output**, not hand-written examples (no circularity).
|
||||
- **Same Claude Code** the skills target — no separate API key.
|
||||
- **Snapshot committed to git** so CI runs are deterministic and free,
|
||||
and so any change to the numbers is reviewable as a diff.
|
||||
- **Control arm** isolates the skill's contribution from the generic
|
||||
"be terse" effect.
|
||||
|
||||
## Files
|
||||
|
||||
- `prompts/en.txt` — fixed list of dev questions, one per line.
|
||||
- `llm_run.py` — runs `claude -p --system-prompt …` per (prompt, arm),
|
||||
captures real LLM output, writes `snapshots/results.json` along with
|
||||
metadata (model, CLI version, generation timestamp).
|
||||
- `measure.py` — reads the snapshot, counts tokens with tiktoken
|
||||
`o200k_base`, prints a markdown table with median / mean / min / max /
|
||||
stdev across prompts.
|
||||
- `snapshots/results.json` — committed source of truth, regenerated only
|
||||
when SKILL.md files or prompts change.
|
||||
|
||||
## Refresh the snapshot (requires `claude` CLI logged in)
|
||||
|
||||
```bash
|
||||
uv run python evals/llm_run.py
|
||||
```
|
||||
|
||||
This calls Claude once per prompt × (N skills + 2 control arms). Use
|
||||
a small model to keep it cheap:
|
||||
|
||||
```bash
|
||||
CAVEMAN_EVAL_MODEL=claude-haiku-4-5 uv run python evals/llm_run.py
|
||||
```
|
||||
|
||||
## Read the snapshot (no LLM, no API key, runs in CI)
|
||||
|
||||
```bash
|
||||
uv run --with tiktoken python evals/measure.py
|
||||
```
|
||||
|
||||
## Adding a prompt
|
||||
|
||||
Append a line to `prompts/en.txt`, then refresh the snapshot.
|
||||
|
||||
## Adding a skill
|
||||
|
||||
Drop a `skills/<name>/SKILL.md`, then refresh the snapshot. `llm_run.py`
|
||||
picks up every skill directory automatically.
|
||||
|
||||
## What this does NOT measure
|
||||
|
||||
- **Fidelity** — does the compressed answer preserve the technical
|
||||
claims? A skill that replies `k` to everything would score −99% and
|
||||
"win". A future v2 could add a judge-model rubric.
|
||||
- **Latency or cost** — out of scope. Note that skills add input tokens
|
||||
on every call, so output savings are not the full economic picture.
|
||||
- **Cross-model behavior** — only the model used to generate the
|
||||
snapshot is measured.
|
||||
- **Exact Claude tokens** — `tiktoken o200k_base` is OpenAI's BPE and is
|
||||
only an approximation of Claude's tokenizer. Ratios between arms are
|
||||
meaningful; absolute numbers are approximate.
|
||||
- **Statistical significance** — single run per (prompt, arm) at default
|
||||
temperature. The min/max/stdev columns let you eyeball whether a
|
||||
number is solid or noisy, but this is not a powered experiment.
|
||||
@@ -0,0 +1,105 @@
|
||||
"""
|
||||
Run each prompt through Claude Code in three conditions and snapshot the
|
||||
real LLM outputs:
|
||||
|
||||
1. baseline — no extra system prompt at all
|
||||
2. terse — system prompt: "Answer concisely."
|
||||
3. terse+skill — system prompt: "Answer concisely.\n\n{SKILL.md}"
|
||||
|
||||
The honest delta is (3) vs (2): how much does the SKILL itself add on top
|
||||
of a plain "be terse" instruction? Comparing (3) vs (1) conflates the
|
||||
skill with the generic terseness ask, which is what the previous version
|
||||
of this harness did.
|
||||
|
||||
This is the source-of-truth generator. It calls a real LLM and produces
|
||||
evals/snapshots/results.json. Run it locally when SKILL.md files change.
|
||||
The CI-side `measure.py` only reads the snapshot and counts tokens.
|
||||
|
||||
Requires:
|
||||
- `claude` CLI on PATH (Claude Code), authenticated
|
||||
|
||||
Run: uv run python evals/llm_run.py
|
||||
|
||||
Environment:
|
||||
CAVEMAN_EVAL_MODEL optional --model flag value passed through to claude
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import datetime as dt
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
from pathlib import Path
|
||||
|
||||
EVALS = Path(__file__).parent
|
||||
SKILLS = EVALS.parent / "skills"
|
||||
PROMPTS = EVALS / "prompts" / "en.txt"
|
||||
SNAPSHOT = EVALS / "snapshots" / "results.json"
|
||||
|
||||
TERSE_PREFIX = "Answer concisely."
|
||||
|
||||
|
||||
def run_claude(prompt: str, system: str | None = None) -> str:
|
||||
cmd = ["claude", "-p"]
|
||||
if system:
|
||||
cmd += ["--system-prompt", system]
|
||||
if model := os.environ.get("CAVEMAN_EVAL_MODEL"):
|
||||
cmd += ["--model", model]
|
||||
cmd.append(prompt)
|
||||
out = subprocess.run(cmd, capture_output=True, text=True, check=True)
|
||||
return out.stdout.strip()
|
||||
|
||||
|
||||
def claude_version() -> str:
|
||||
try:
|
||||
out = subprocess.run(
|
||||
["claude", "--version"], capture_output=True, text=True, check=True
|
||||
)
|
||||
return out.stdout.strip()
|
||||
except Exception:
|
||||
return "unknown"
|
||||
|
||||
|
||||
def main() -> None:
|
||||
prompts = [p.strip() for p in PROMPTS.read_text().splitlines() if p.strip()]
|
||||
skills = sorted(p.name for p in SKILLS.iterdir() if (p / "SKILL.md").exists())
|
||||
|
||||
print(
|
||||
f"=== {len(prompts)} prompts × ({len(skills)} skills + 2 control arms) ===",
|
||||
flush=True,
|
||||
)
|
||||
|
||||
snapshot: dict = {
|
||||
"metadata": {
|
||||
"generated_at": dt.datetime.now(dt.timezone.utc).isoformat(),
|
||||
"claude_cli_version": claude_version(),
|
||||
"model": os.environ.get("CAVEMAN_EVAL_MODEL", "default"),
|
||||
"n_prompts": len(prompts),
|
||||
"terse_prefix": TERSE_PREFIX,
|
||||
},
|
||||
"prompts": prompts,
|
||||
"arms": {},
|
||||
}
|
||||
|
||||
print("baseline (no system prompt)", flush=True)
|
||||
snapshot["arms"]["__baseline__"] = [run_claude(p) for p in prompts]
|
||||
|
||||
print("terse (control: terse instruction only, no skill)", flush=True)
|
||||
snapshot["arms"]["__terse__"] = [
|
||||
run_claude(p, system=TERSE_PREFIX) for p in prompts
|
||||
]
|
||||
|
||||
for skill in skills:
|
||||
skill_md = (SKILLS / skill / "SKILL.md").read_text()
|
||||
system = f"{TERSE_PREFIX}\n\n{skill_md}"
|
||||
print(f" {skill}", flush=True)
|
||||
snapshot["arms"][skill] = [run_claude(p, system=system) for p in prompts]
|
||||
|
||||
SNAPSHOT.parent.mkdir(parents=True, exist_ok=True)
|
||||
SNAPSHOT.write_text(json.dumps(snapshot, ensure_ascii=False, indent=2))
|
||||
print(f"\nWrote {SNAPSHOT}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,107 @@
|
||||
"""
|
||||
Read evals/snapshots/results.json (produced by llm_run.py) and report
|
||||
real token compression per skill against the *terse control arm* — i.e.
|
||||
how much the skill adds on top of a plain "Answer concisely." instruction.
|
||||
|
||||
Reports median, min, max and stdev across prompts, not just the mean,
|
||||
so the reader can see whether a number is solid or noisy.
|
||||
|
||||
Tokenizer note: tiktoken o200k_base is OpenAI's tokenizer and is only an
|
||||
approximation of Claude's BPE. The ratios are still meaningful for
|
||||
comparing skills against each other, but the absolute numbers should be
|
||||
read as "approximate output-length reduction", not "exact Claude tokens".
|
||||
|
||||
Run: uv run --with tiktoken python evals/measure.py
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import statistics
|
||||
from pathlib import Path
|
||||
|
||||
import tiktoken
|
||||
|
||||
ENCODING = tiktoken.get_encoding("o200k_base")
|
||||
SNAPSHOT = Path(__file__).parent / "snapshots" / "results.json"
|
||||
|
||||
|
||||
def count(text: str) -> int:
|
||||
return len(ENCODING.encode(text))
|
||||
|
||||
|
||||
def stats(savings: list[float]) -> tuple[float, float, float, float, float]:
|
||||
return (
|
||||
statistics.median(savings),
|
||||
statistics.mean(savings),
|
||||
min(savings),
|
||||
max(savings),
|
||||
statistics.stdev(savings) if len(savings) > 1 else 0.0,
|
||||
)
|
||||
|
||||
|
||||
def fmt_pct(x: float) -> str:
|
||||
sign = "−" if x < 0 else "+"
|
||||
return f"{sign}{abs(x) * 100:.0f}%"
|
||||
|
||||
|
||||
def main() -> None:
|
||||
if not SNAPSHOT.exists():
|
||||
print(f"No snapshot at {SNAPSHOT}. Run `python evals/llm_run.py` first.")
|
||||
return
|
||||
|
||||
data = json.loads(SNAPSHOT.read_text())
|
||||
arms = data["arms"]
|
||||
meta = data.get("metadata", {})
|
||||
|
||||
baseline_tokens = [count(o) for o in arms["__baseline__"]]
|
||||
terse_tokens = [count(o) for o in arms["__terse__"]]
|
||||
|
||||
print(f"_Generated: {meta.get('generated_at', '?')}_")
|
||||
print(
|
||||
f"_Model: {meta.get('model', '?')} · CLI: {meta.get('claude_cli_version', '?')}_"
|
||||
)
|
||||
print(f"_Tokenizer: tiktoken o200k_base (approximation of Claude's BPE)_")
|
||||
print(
|
||||
f"_n = {meta.get('n_prompts', len(baseline_tokens))} prompts, single run per arm_"
|
||||
)
|
||||
print()
|
||||
print(f"**Reference arms (no skill):**")
|
||||
print(f"- baseline (no system prompt): {sum(baseline_tokens)} tokens total")
|
||||
print(
|
||||
f"- terse control (`Answer concisely.`): {sum(terse_tokens)} tokens total "
|
||||
f"({fmt_pct(1 - sum(terse_tokens) / sum(baseline_tokens))} vs baseline)"
|
||||
)
|
||||
print()
|
||||
print("**Skills, measured as additional reduction on top of the terse control:**")
|
||||
print()
|
||||
print("| Skill | Median | Mean | Min | Max | Stdev | Tokens (skill / terse) |")
|
||||
print("|-------|--------|------|-----|-----|-------|-------------------------|")
|
||||
|
||||
rows = []
|
||||
for skill, outputs in arms.items():
|
||||
if skill in ("__baseline__", "__terse__"):
|
||||
continue
|
||||
skill_tokens = [count(o) for o in outputs]
|
||||
savings = [
|
||||
1 - (s / t) if t else 0.0 for s, t in zip(skill_tokens, terse_tokens)
|
||||
]
|
||||
med, mean, lo, hi, sd = stats(savings)
|
||||
rows.append(
|
||||
(skill, med, mean, lo, hi, sd, sum(skill_tokens), sum(terse_tokens))
|
||||
)
|
||||
|
||||
for row in sorted(rows, key=lambda r: -r[1]):
|
||||
skill, med, mean, lo, hi, sd, st, tt = row
|
||||
print(
|
||||
f"| **{skill}** | {fmt_pct(med)} | {fmt_pct(mean)} | "
|
||||
f"{fmt_pct(lo)} | {fmt_pct(hi)} | {sd * 100:.0f}% | {st} / {tt} |"
|
||||
)
|
||||
|
||||
print()
|
||||
print("_Savings = `1 - skill_tokens / terse_tokens` per prompt._")
|
||||
print(f"_Source: {SNAPSHOT.name}. Refresh with `python evals/llm_run.py`._")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,150 @@
|
||||
"""
|
||||
Generate a boxplot showing the distribution of token compression per
|
||||
skill, compared against a plain "Answer concisely." control.
|
||||
|
||||
Reads evals/snapshots/results.json and writes:
|
||||
- evals/snapshots/results.html (interactive plotly)
|
||||
- evals/snapshots/results.png (static export for README/PR embed)
|
||||
|
||||
Run: uv run --with tiktoken --with plotly --with kaleido python evals/plot.py
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import statistics
|
||||
from pathlib import Path
|
||||
|
||||
import plotly.graph_objects as go
|
||||
import tiktoken
|
||||
|
||||
ENCODING = tiktoken.get_encoding("o200k_base")
|
||||
SNAPSHOT = Path(__file__).parent / "snapshots" / "results.json"
|
||||
HTML_OUT = Path(__file__).parent / "snapshots" / "results.html"
|
||||
PNG_OUT = Path(__file__).parent / "snapshots" / "results.png"
|
||||
|
||||
|
||||
def count(text: str) -> int:
|
||||
return len(ENCODING.encode(text))
|
||||
|
||||
|
||||
def main() -> None:
|
||||
data = json.loads(SNAPSHOT.read_text())
|
||||
arms = data["arms"]
|
||||
meta = data.get("metadata", {})
|
||||
|
||||
terse_tokens = [count(o) for o in arms["__terse__"]]
|
||||
|
||||
rows = []
|
||||
for skill, outputs in arms.items():
|
||||
if skill in ("__baseline__", "__terse__"):
|
||||
continue
|
||||
skill_tokens = [count(o) for o in outputs]
|
||||
savings = [
|
||||
(1 - (s / t)) * 100 if t else 0.0
|
||||
for s, t in zip(skill_tokens, terse_tokens)
|
||||
]
|
||||
rows.append(
|
||||
{"skill": skill, "savings": savings, "median": statistics.median(savings)}
|
||||
)
|
||||
|
||||
rows.sort(key=lambda r: -r["median"]) # best first
|
||||
|
||||
fig = go.Figure()
|
||||
|
||||
for row in rows:
|
||||
fig.add_trace(
|
||||
go.Box(
|
||||
y=row["savings"],
|
||||
name=row["skill"],
|
||||
boxpoints="all",
|
||||
jitter=0.4,
|
||||
pointpos=0,
|
||||
marker=dict(color="#2ca02c", size=7, opacity=0.7),
|
||||
line=dict(color="#2c3e50", width=2),
|
||||
fillcolor="rgba(76, 120, 168, 0.25)",
|
||||
boxmean=True,
|
||||
hovertemplate="<b>%{x}</b><br>%{y:.1f}%<extra></extra>",
|
||||
)
|
||||
)
|
||||
|
||||
# zero line — "no effect"
|
||||
fig.add_hline(
|
||||
y=0,
|
||||
line=dict(color="black", width=1.5, dash="dash"),
|
||||
annotation_text="no effect (= same length as control)",
|
||||
annotation_position="top right",
|
||||
annotation_font=dict(size=11, color="black"),
|
||||
)
|
||||
|
||||
# median labels above each box
|
||||
for row in rows:
|
||||
fig.add_annotation(
|
||||
x=row["skill"],
|
||||
y=max(row["savings"]),
|
||||
text=f"<b>{row['median']:+.0f}%</b>",
|
||||
showarrow=False,
|
||||
yshift=22,
|
||||
font=dict(size=16, color="#2c3e50"),
|
||||
)
|
||||
|
||||
fig.update_layout(
|
||||
title=dict(
|
||||
text=f"<b>How much shorter does each skill make Claude's answers?</b><br>"
|
||||
f"<sub>Distribution of per-prompt savings vs system prompt = "
|
||||
f"<i>'Answer concisely.'</i><br>"
|
||||
f"{meta.get('model', '?')} · n={meta.get('n_prompts', '?')} prompts · "
|
||||
f"single run per arm</sub>",
|
||||
x=0.5,
|
||||
xanchor="center",
|
||||
),
|
||||
xaxis=dict(title="", automargin=True),
|
||||
yaxis=dict(
|
||||
title="↑ shorter · vs control · longer ↓",
|
||||
ticksuffix="%",
|
||||
zeroline=False,
|
||||
gridcolor="rgba(0,0,0,0.08)",
|
||||
range=[-30, 115],
|
||||
),
|
||||
plot_bgcolor="white",
|
||||
height=560,
|
||||
width=980,
|
||||
margin=dict(l=140, r=80, t=120, b=120),
|
||||
showlegend=False,
|
||||
annotations=[
|
||||
dict(
|
||||
x=0.5,
|
||||
y=-0.22,
|
||||
xref="paper",
|
||||
yref="paper",
|
||||
showarrow=False,
|
||||
font=dict(size=11, color="#555"),
|
||||
text=(
|
||||
"<b>box</b> = IQR (middle 50%) · "
|
||||
"<b>line in box</b> = median · "
|
||||
"<b>dashed line</b> = mean · "
|
||||
"<b>green dots</b> = individual prompts"
|
||||
),
|
||||
)
|
||||
],
|
||||
)
|
||||
|
||||
# re-add labels after update_layout (which would otherwise wipe them)
|
||||
for row in rows:
|
||||
fig.add_annotation(
|
||||
x=row["skill"],
|
||||
y=max(row["savings"]),
|
||||
text=f"<b>{row['median']:+.0f}%</b>",
|
||||
showarrow=False,
|
||||
yshift=22,
|
||||
font=dict(size=16, color="#2c3e50"),
|
||||
)
|
||||
|
||||
fig.write_html(HTML_OUT)
|
||||
print(f"Wrote {HTML_OUT}")
|
||||
fig.write_image(PNG_OUT, scale=2)
|
||||
print(f"Wrote {PNG_OUT}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,10 @@
|
||||
Why does my React component re-render every time the parent updates?
|
||||
Explain database connection pooling.
|
||||
What's the difference between TCP and UDP?
|
||||
How do I fix a memory leak in a long-running Node.js process?
|
||||
What does the SQL EXPLAIN command tell me?
|
||||
How does a hash table handle collisions?
|
||||
Why am I getting CORS errors in my browser console?
|
||||
What's the point of using a debouncer on a search input?
|
||||
How does git rebase differ from git merge?
|
||||
When should I use a queue vs a topic in messaging systems?
|
||||
@@ -0,0 +1,95 @@
|
||||
{
|
||||
"metadata": {
|
||||
"generated_at": "2026-04-08T22:01:24.027359+00:00",
|
||||
"claude_cli_version": "2.1.97 (Claude Code)",
|
||||
"model": "claude-opus-4-6",
|
||||
"n_prompts": 10,
|
||||
"terse_prefix": "Answer concisely."
|
||||
},
|
||||
"prompts": [
|
||||
"Why does my React component re-render every time the parent updates?",
|
||||
"Explain database connection pooling.",
|
||||
"What's the difference between TCP and UDP?",
|
||||
"How do I fix a memory leak in a long-running Node.js process?",
|
||||
"What does the SQL EXPLAIN command tell me?",
|
||||
"How does a hash table handle collisions?",
|
||||
"Why am I getting CORS errors in my browser console?",
|
||||
"What's the point of using a debouncer on a search input?",
|
||||
"How does git rebase differ from git merge?",
|
||||
"When should I use a queue vs a topic in messaging systems?"
|
||||
],
|
||||
"arms": {
|
||||
"__baseline__": [
|
||||
"Because by default, React re-renders all children when a parent re-renders, regardless of whether their props changed.\n\nTo prevent it, wrap the component in `React.memo`:\n\n```jsx\nconst Child = React.memo(function Child(props) { ... });\n```\n\nCaveats:\n- `memo` does a shallow prop comparison. New object/array/function props created inline in the parent (`<Child onClick={() => ...} data={{...}} />`) break it — wrap those in `useCallback`/`useMemo`.\n- Children passed via `children` or other JSX props are new elements each render, so `memo` won't help there.\n- Context updates still re-render consumers regardless of `memo`.\n\nIf you share the component + parent code, I can point to the exact cause.",
|
||||
"Database connection pooling maintains a cache of reusable DB connections instead of opening/closing one per query.\n\n**Why:** Opening a TCP + auth handshake per query is expensive (often 10–100ms). Pools amortize that cost.\n\n**How it works:**\n1. Pool opens N connections at startup (or lazily).\n2. App \"checks out\" a connection, runs queries, \"returns\" it.\n3. Idle connections stay open, ready for reuse.\n4. If all are busy, requests wait or the pool grows up to `max_size`.\n\n**Key knobs:**\n- `min_size` / `max_size` — floor and ceiling of connections.\n- `max_idle_time` — close connections idle too long.\n- `max_lifetime` — recycle old connections (avoids stale TCP state).\n- `timeout` — how long a checkout waits before erroring.\n\n**Gotchas:**\n- **Sizing**: `max_size` × app instances must stay below DB's `max_connections`. Postgres typically handles ~100–200 well; beyond that use PgBouncer.\n- **Transaction leaks**: forgetting to release a connection starves the pool.\n- **Serverless**: short-lived functions defeat pooling — use an external pooler (PgBouncer, RDS Proxy, Neon pooler).\n- **Async frameworks** (FastAPI + asyncpg/SQLAlchemy async): pool is per-process; workers multiply it.\n\n**Common libraries:** HikariCP (Java), pgbouncer (external, Postgres), SQLAlchemy `QueuePool`, asyncpg `create_pool`, node-postgres `Pool`.",
|
||||
"TCP is connection-oriented, reliable, and ordered: it establishes a handshake, retransmits lost packets, guarantees delivery and sequence, and handles flow/congestion control. Used for HTTP, SSH, email.\n\nUDP is connectionless and best-effort: no handshake, no retransmission, no ordering, minimal overhead. Faster and lighter, but packets can be lost or arrive out of order. Used for DNS, video/voice streaming, games, QUIC.\n\nKey tradeoff: TCP = reliability, UDP = low latency.",
|
||||
"Common approaches:\n\n1. **Find the leak** — run with `--inspect`, take heap snapshots in Chrome DevTools at intervals, compare retained objects. Or use `clinic.js doctor` / `heapdump`.\n\n2. **Usual suspects**:\n - Unbounded caches/Maps — use `LRUCache` with a max size.\n - Event listeners not removed — check `emitter.listenerCount()`, use `once()` or `off()`.\n - Closures holding large scopes — especially in callbacks/timers.\n - Global arrays that only grow (logs, metrics buffers).\n - Timers (`setInterval`) keeping refs alive — `clearInterval` on shutdown, or `.unref()`.\n - Promises that never settle, holding their chain in memory.\n\n3. **Measure** — log `process.memoryUsage().heapUsed` over time; a steadily rising baseline across GCs confirms a leak (vs. normal sawtooth).\n\n4. **Bound it** — run with `--max-old-space-size=N` so it crashes loudly instead of thrashing, and use a supervisor (pm2, systemd) to restart.\n\nIf you share the code or a heap snapshot I can point at the specific culprit.",
|
||||
"`EXPLAIN` shows the query planner's execution plan for a SQL statement — how the database intends to run it, without actually executing it (use `EXPLAIN ANALYZE` to run it and get real timings).\n\nKey things it reveals:\n- **Scan type**: Seq Scan (full table) vs Index Scan / Index Only Scan / Bitmap Scan — tells you if indexes are being used.\n- **Join strategy**: Nested Loop, Hash Join, Merge Join.\n- **Estimated cost**: `cost=startup..total`, `rows=`, `width=` (bytes/row). Costs are in arbitrary planner units, not ms.\n- **Order of operations**: read bottom-up, inside-out.\n- **Filters vs Index Conds**: a `Filter:` after a Seq Scan means rows were read then discarded — often a missing index.\n\nWith `EXPLAIN ANALYZE` you also get `actual time`, `rows`, and `loops`. Big gaps between estimated and actual rows usually mean stale stats (`ANALYZE` the table). Add `BUFFERS` to see cache vs disk reads.",
|
||||
"Two main strategies:\n\n**1. Separate chaining** — each bucket stores a linked list (or tree) of entries. Colliding keys get appended. Lookup scans the chain. Used by Java's `HashMap`, Python's `dict` historically is different.\n\n**2. Open addressing** — on collision, probe for another empty slot in the array itself:\n- *Linear probing*: try `i+1, i+2, ...` (cache-friendly but clusters)\n- *Quadratic probing*: try `i+1, i+4, i+9, ...` (reduces clustering)\n- *Double hashing*: step size from a second hash function\n\nPython's `dict` and Ruby's `Hash` use open addressing; Go's `map` uses chaining with bucket arrays.\n\nBoth degrade as load factor rises, so tables resize (typically at 0.7–0.75) to keep average lookup O(1).",
|
||||
"CORS errors happen when your frontend (origin A) makes a request to a backend (origin B) and the server doesn't return the `Access-Control-Allow-Origin` header permitting origin A. The browser enforces this — the request often reaches the server, but the response is blocked.\n\nCommon causes:\n- **Missing CORS middleware** on the backend (e.g., FastAPI needs `CORSMiddleware`, Express needs `cors`).\n- **Origin mismatch**: `localhost:3000` ≠ `127.0.0.1:3000`, `http` ≠ `https`, wrong port.\n- **Preflight (OPTIONS) failing**: non-simple requests (custom headers, `PUT`/`DELETE`, `application/json` with auth) trigger an OPTIONS preflight the server must handle.\n- **Credentials**: if sending cookies/auth, you need `Access-Control-Allow-Credentials: true` AND an explicit origin (not `*`).\n- **Wildcard + credentials**: `Allow-Origin: *` is rejected when credentials are included.\n\nTo help more specifically: what's the exact error message, and what stack is your backend?",
|
||||
"A debouncer delays firing the search until the user pauses typing (e.g., 300ms of no keystrokes). Without it, every keystroke triggers a search — typing \"react\" fires 5 requests.\n\nBenefits:\n- **Fewer API calls** — cuts network load and backend cost\n- **Avoids race conditions** — stale responses arriving out of order\n- **Better UX** — no flickering results as the user types\n- **Less CPU work** — for expensive client-side filtering\n\nThrottle vs debounce: throttle fires at a fixed rate during activity; debounce waits for silence. Search inputs want debounce because you only care about the final query.",
|
||||
"**Merge**: Combines branches by creating a new \"merge commit\" that ties their histories together. Preserves full history, including the branch topology. Non-destructive.\n\n**Rebase**: Replays your branch's commits on top of another branch, rewriting them as new commits. Produces a linear history with no merge commit.\n\nKey differences:\n- **History shape**: merge = branching graph; rebase = straight line\n- **Commit hashes**: merge preserves them; rebase rewrites them\n- **Conflicts**: merge resolves once; rebase may resolve per-commit\n- **Safety**: merge is safe on shared branches; rebase should be avoided on already-pushed/shared branches (rewrites history)\n\nRule of thumb: rebase local work to clean it up before sharing; merge when integrating shared branches.",
|
||||
"**Queue**: point-to-point. One message → one consumer. Use for work distribution where each job must be processed exactly once (task queues, order processing, background jobs). Multiple workers compete; load balances across them.\n\n**Topic**: pub/sub. One message → many subscribers, each gets a copy. Use for broadcasting events to independent consumers (event notifications, cache invalidation, fan-out to analytics + audit + email).\n\nRule of thumb: \"who needs to react?\"\n- One of N workers → queue\n- All interested parties → topic\n\nModern brokers blur the line (Kafka topics with consumer groups act queue-like; SNS+SQS combines both)."
|
||||
],
|
||||
"__terse__": [
|
||||
"React components re-render by default whenever their parent re-renders, regardless of whether their props changed. Common causes and fixes:\n\n1. **Default behavior** — Wrap the child in `React.memo(MyComponent)` to skip re-renders when props are shallow-equal.\n\n2. **New object/array/function props each render** — Even with `memo`, passing `{}`, `[]`, or inline `() => {}` creates new references every render, breaking memoization. Fix with `useMemo` / `useCallback` in the parent.\n\n3. **Children prop / context** — `memo` doesn't help if you pass `children` that change, or if the component consumes a context whose value changes.\n\n4. **New reference from hooks** — e.g. `useState` setters are stable, but values derived inline aren't.\n\nQuick check: add `console.log` with the props and see which one changes identity. If you share the component + parent code I can point to the exact culprit.",
|
||||
"Database connection pooling maintains a cache of reusable database connections instead of opening/closing one per request.\n\n**Why it matters:** Opening a DB connection is expensive — TCP handshake, TLS, auth, session setup (often 50–200ms). Doing this per query kills throughput.\n\n**How it works:**\n1. Pool pre-opens N connections at startup.\n2. App \"checks out\" a connection, runs queries, \"returns\" it.\n3. Idle connections stay warm, ready for reuse.\n4. If all are busy, new requests wait (or the pool grows up to a max).\n\n**Key parameters:**\n- `min_size` / `max_size` — floor and ceiling of connections\n- `timeout` — max wait for a free connection\n- `max_lifetime` / `idle_timeout` — recycle stale connections\n- `max_overflow` — temporary burst capacity above max_size\n\n**Common tools:** SQLAlchemy's `QueuePool`, `asyncpg.create_pool`, HikariCP (Java), PgBouncer (external pooler sitting in front of Postgres).\n\n**Gotchas:**\n- Pool size > DB's `max_connections` → errors. Size per-process, multiply by workers.\n- Long transactions starve the pool.\n- Serverless/short-lived processes benefit from an external pooler (PgBouncer) in transaction mode.\n- Connections can go stale (network drops, DB restarts) — use health checks / `pre_ping`.\n\n**Rule of thumb:** pool size ≈ `(cores * 2) + effective_spindles`, tuned by load testing. Bigger isn't better — contention at the DB beats connection reuse gains.",
|
||||
"TCP is connection-oriented, reliable, and ordered: it establishes a handshake, guarantees delivery, retransmits lost packets, and preserves sequence — used for HTTP, SSH, email.\n\nUDP is connectionless and best-effort: no handshake, no delivery guarantee, no ordering, lower latency and overhead — used for DNS, video/voice streaming, games.\n\nKey tradeoff: TCP = reliability, UDP = speed.",
|
||||
"Common approaches:\n\n1. **Profile heap**: Run with `node --inspect` and use Chrome DevTools → Memory → take heap snapshots at intervals, compare retained size to find growing objects.\n2. **Usual culprits**:\n - Unbounded caches/Maps → use `lru-cache` or `WeakMap`.\n - Event listeners not removed → check `emitter.listenerCount()`, call `off()`/`removeListener()`.\n - Closures holding large scopes (e.g. in timers, promises).\n - Global arrays that only grow.\n - Unclosed DB/HTTP connections or streams.\n3. **Monitor**: log `process.memoryUsage().heapUsed` over time; use `--max-old-space-size` only as a bandaid.\n4. **Tools**: `clinic.js doctor`, `heapdump`, `0x`, or `--heap-prof` flag for sampling.\n5. **Reproduce in isolation**: load-test one endpoint/job at a time to localize the leak.\n\nStart with a heap snapshot diff — it usually points straight at the retainer.",
|
||||
"SQL `EXPLAIN` shows the query execution plan — how the database will run your query. Key info:\n\n- **Access method**: sequential scan vs index scan vs index-only scan\n- **Join strategy**: nested loop, hash join, merge join\n- **Row estimates**: how many rows the planner expects at each step\n- **Cost estimates**: relative startup/total cost units\n- **Order of operations**: which tables/filters/sorts happen first\n\nUse `EXPLAIN ANALYZE` to actually execute the query and get real timings + actual row counts (vs estimates), which is what you need to spot bad plans (e.g., estimate says 10 rows, reality is 10M → stale stats or bad selectivity).\n\nCommon things to look for:\n- Seq Scan on large tables → missing index\n- Big gap between estimated and actual rows → run `ANALYZE`\n- Nested loop over huge row counts → join/index problem\n- Sort spilling to disk → need `work_mem` or an index providing order",
|
||||
"Hash tables handle collisions with two main strategies:\n\n1. **Separate chaining**: Each bucket holds a linked list (or tree) of entries. Collisions append to the list. Lookup scans the chain. Java's `HashMap` uses this (converting to a tree after 8 entries).\n\n2. **Open addressing**: Store entries directly in the array; on collision, probe for another slot:\n - **Linear probing**: try `i+1, i+2, ...`\n - **Quadratic probing**: try `i+1², i+2², ...`\n - **Double hashing**: step size from a second hash function\n\nChaining is simpler and degrades gracefully; open addressing is cache-friendlier but needs resizing before load factor gets high (~0.7).",
|
||||
"CORS errors happen when your frontend (origin A) requests a resource from a server (origin B) that doesn't return the right `Access-Control-Allow-*` headers. The browser blocks the response — the request often reaches the server fine.\n\nCommon causes:\n1. **Missing header** — server doesn't send `Access-Control-Allow-Origin` matching your frontend origin.\n2. **Wildcard + credentials** — using `Access-Control-Allow-Origin: *` with `credentials: 'include'`. You must echo the exact origin and set `Access-Control-Allow-Credentials: true`.\n3. **Preflight failing** — non-simple requests (custom headers, `PUT`/`DELETE`, `application/json`) trigger an `OPTIONS` preflight. Server must handle `OPTIONS` and return `Access-Control-Allow-Methods` / `Access-Control-Allow-Headers`.\n4. **Protocol/port mismatch** — `http://localhost:3000` ≠ `http://localhost:8000` ≠ `https://...`. All count as different origins.\n5. **Redirects** — CORS headers must be on the final response; redirects can strip them.\n\nQuick fixes by stack:\n- **FastAPI**: `app.add_middleware(CORSMiddleware, allow_origins=[...], allow_credentials=True, allow_methods=[\"*\"], allow_headers=[\"*\"])`\n- **Express**: `app.use(cors({ origin: '...', credentials: true }))`\n- **Dev only**: use a Vite/Next proxy so requests are same-origin.\n\nShare the exact error message and your frontend/backend origins and I can pinpoint it.",
|
||||
"A debouncer delays running the search until the user stops typing for a short interval (e.g., 300ms). Benefits:\n\n- **Fewer API calls / DB queries**: avoids firing a request on every keystroke (\"a\", \"ap\", \"app\"...) — only one fires after typing pauses.\n- **Lower cost & server load**: especially important for expensive backends or rate-limited APIs.\n- **Better UX**: prevents flickering results and out-of-order responses racing each other.\n- **Saves client work**: less re-rendering and state churn.\n\nRelated: *throttle* caps frequency (e.g., 1/sec), while *debounce* waits for silence — debounce is usually what you want for search inputs.",
|
||||
"**Merge**: Combines branches by creating a new merge commit that ties both histories together. Preserves the full branch history and is non-destructive.\n\n**Rebase**: Replays your branch's commits on top of another branch, creating new commits with new hashes. Produces a linear history but rewrites commits.\n\nKey differences:\n- **History**: merge = branching graph; rebase = straight line\n- **Safety**: merge is safe on shared branches; rebase should be avoided on pushed/shared branches (rewrites history)\n- **Conflicts**: merge resolves once; rebase may require resolving per replayed commit\n- **Traceability**: merge keeps context of when/where branches diverged; rebase loses it\n\nRule of thumb: rebase local/private work to clean up, merge for integrating shared branches.",
|
||||
"**Queue** (point-to-point): one message → one consumer. Use for work distribution, task processing, load balancing across workers. Each job should be handled exactly once (e.g., order processing, email sending, image resizing).\n\n**Topic** (pub/sub): one message → many subscribers. Use for event broadcasting where multiple independent systems react to the same event (e.g., \"OrderPlaced\" → inventory, billing, analytics, notifications all consume it).\n\n**Rule of thumb:**\n- \"Who does this work?\" → queue\n- \"Who needs to know?\" → topic\n\nModern brokers blur the line (Kafka topics with consumer groups act like queues per group; SNS+SQS fan-out combines both), but the semantic distinction still drives the design choice."
|
||||
],
|
||||
"caveman": [
|
||||
"Parent re-render → child re-render by default. Props \"change\" each render if inline obj/array/fn → new ref → child sees new props.\n\nFix:\n- Wrap child: `React.memo(Child)`\n- Stabilize props: `useMemo` for obj/array, `useCallback` for fn\n- Lift state so parent re-render not needed\n\nCheck: if child memo'd but still re-renders, inline prop is culprit.",
|
||||
"Pool = reuse open DB conns. No new conn per req → skip TCP+auth handshake. Fixed max size cap load on DB. Idle conns recycled. Win: latency ↓, throughput ↑.",
|
||||
"TCP: connection-based. Handshake, ordered, reliable, retransmit lost packets, flow/congestion control. Slower. Use: web, SSH, files.\n\nUDP: connectionless. No handshake, no order, no retransmit. Fast, low overhead. Use: video, games, DNS, VoIP.",
|
||||
"Use caveman skill (requested mode by system).\n\n**Find leak:**\n- Run with `node --inspect`. Open `chrome://inspect` → Memory tab.\n- Take 3 heap snapshots over time. Compare. Look \"Retained Size\" growing.\n- Or `clinic doctor -- node app.js` for quick diagnosis.\n\n**Common causes:**\n- Global vars / caches no eviction → use `lru-cache` with max.\n- Event listeners not removed → `emitter.off()` on cleanup. Watch `MaxListenersExceededWarning`.\n- Closures holding big objects → null refs when done.\n- Timers (`setInterval`) never cleared → `clearInterval`.\n- Promises never resolve → pile up.\n\n**Fix pattern:**\n```js\nconst cache = new LRU({ max: 500 });\nemitter.on('x', handler);\n// cleanup\nemitter.off('x', handler);\nclearInterval(timer);\n```\n\n**Monitor prod:** `process.memoryUsage().heapUsed`. Alert on growth trend. Restart worker if threshold hit (pm2/k8s).\n\nRoot cause fix > restart band-aid.",
|
||||
"EXPLAIN show query plan. DB tell you *how* it run query, not result.\n\nReveal:\n- **Scan type**: seq scan (full table) vs index scan vs bitmap scan\n- **Join strategy**: nested loop, hash join, merge join\n- **Row estimates**: planner guess of rows per step\n- **Cost**: startup..total (arbitrary units)\n- **Order**: which step first, which last\n\n`EXPLAIN ANALYZE` = actually run query + show real time + real rows. Compare est vs actual → spot bad stats.\n\nUse for: find missing index (seq scan on big table), bad join order, wrong row estimate → `ANALYZE table`.",
|
||||
"Caveman mode on.\n\nTwo main way:\n\n1. **Chaining**: each bucket hold linked list (or tree). Collide → append to list. Lookup walk list.\n2. **Open addressing**: collide → probe next slot. Variants: linear probe (+1), quadratic (+i²), double hash (step = hash2(key)).\n\nLoad factor high → resize + rehash. Java `HashMap` chain→tree at 8. Python dict + Go map use open addressing.",
|
||||
"CORS = server no send `Access-Control-Allow-Origin` header match your origin. Browser block response.\n\nCommon cause:\n- Server miss CORS middleware\n- Origin not in allowlist\n- Preflight (OPTIONS) fail — need `Allow-Methods`, `Allow-Headers`\n- Credentials mode: need `Allow-Credentials: true` + explicit origin (no `*`)\n\nFix FastAPI:\n```python\nfrom fastapi.middleware.cors import CORSMiddleware\napp.add_middleware(CORSMiddleware, allow_origins=[\"http://localhost:3000\"], allow_credentials=True, allow_methods=[\"*\"], allow_headers=[\"*\"])\n```\n\nPaste exact error → me pinpoint.",
|
||||
"Debouncer delay fn call until user stop typing. Without: every keystroke → API hit → spam backend, race conditions, laggy UI. With: wait ~300ms idle → one request. Saves cost, cuts load, better UX.",
|
||||
"Merge = combine branches, make new merge commit. Keep history as-is, show branching.\n\nRebase = replay your commits on top of target branch. Rewrite history, linear shape.\n\nKey diff:\n- **Merge**: non-destructive, preserve context, extra merge commit, messy graph\n- **Rebase**: clean linear log, rewrite SHAs, dangerous on shared branches\n\nRule: merge for shared/public branches, rebase for local cleanup before push.",
|
||||
"Queue = point-to-point. One msg → one consumer. Work distribution, load balance. Ex: job processing.\n\nTopic = pub/sub. One msg → many subscribers. Event broadcast, fan-out. Ex: order placed → notify billing + inventory + email.\n\nRule: task need doing once? Queue. Event many care about? Topic."
|
||||
],
|
||||
"caveman-cn": [
|
||||
"父 render -> 子收到新 props 引用 -> 子重渲染。\n\n修:\n- `React.memo(Child)` 包子组件。\n- 父传的对象/数组/函数用 `useMemo` / `useCallback` 稳引用。\n- 否则 memo 浅比较次次 fail。",
|
||||
"池 = 复用 DB 连接。省握手 -> 更快。限并发 -> 保护 DB。",
|
||||
"TCP:有连接。三次握手。保序。重传。流控/拥塞控。慢但可靠。→ HTTP、SSH、DB。\n\nUDP:无连接。发了就忘。无序。可丢。无重传。快、开销小。→ DNS、视频、游戏、QUIC 底层。",
|
||||
"用 `--inspect` 启动。Chrome DevTools -> Memory -> 拍两张 heap snapshot(间隔负载后)。对比 \"Comparison\" 视图,看 Delta 正的对象。\n\n常见源头:\n- 闭包抓大对象\n- 全局 `Map`/`Set` 只加不删 -> 换 `WeakMap`/`WeakRef` 或加 TTL\n- 监听器泄漏 -> `emitter.setMaxListeners`,记得 `off()`\n- 定时器未 `clearInterval`\n- 缓存无上限 -> 用 `lru-cache`\n\n监控:`process.memoryUsage().heapUsed` 打点。或 `clinic heapprofiler` / `0x`。\n\n复现:`node --expose-gc`,循环后 `global.gc()`,看 RSS 是否回落。不回落 = 泄漏。",
|
||||
"查询执行计划。DB 怎么跑你的 SQL:用哪个 index、join 顺序、扫全表还是走索引、行数估算、成本。\n\nPostgres 用 `EXPLAIN ANALYZE` -> 真实执行时间 + 实际行数。对比 estimated vs actual 找烂计划。\n\n关键字段:\n- `Seq Scan` -> 全表扫,通常坏\n- `Index Scan` -> 走索引,好\n- `rows=` -> 优化器估算\n- `cost=` -> 相对开销\n- `Nested Loop` / `Hash Join` / `Merge Join` -> join 策略\n\n用途:debug 慢查询,决定加哪个 index。",
|
||||
"链地址法:桶存链表/树,冲突元素串一起。开放寻址:另找空槽——线性探测、二次探测、双重哈希。负载因子过高 -> 扩容 rehash。Java 8 `HashMap`:链表长度 >8 转红黑树。",
|
||||
"CORS = 浏览器拦。服务器没回 `Access-Control-Allow-Origin` 头,或回的 origin 不匹配。\n\n修:\n- 后端加头:`Access-Control-Allow-Origin: <你的前端域>`(或 `*`,但带 cookie 不行)\n- 预检 `OPTIONS` 要回 200 + `Allow-Methods` / `Allow-Headers`\n- 带 cookie:`Allow-Credentials: true` + 明确 origin(不能 `*`)\n- FastAPI:`CORSMiddleware`。Express:`cors` 包。\n\n贴报错原文 + 请求/响应头,能更准。",
|
||||
"防止每字触发请求。用户打字 -> 等停顿(如300ms) -> 才发请求。省 API 调用、减抖动、少无用渲染。",
|
||||
"merge: 保留两条历史,生成 merge commit。非破坏性。历史有分叉。\n\nrebase: 把你的 commit 摘下来,重放到目标分支顶端。线性历史。commit 变新 hash (重写)。\n\n选择:\n- 公共分支 -> merge (别重写共享历史)\n- 本地/私有分支整理 -> rebase\n- 想保留合并上下文 -> merge\n- 想干净线性 log -> rebase\n\n冲突:merge 解一次。rebase 可能每个 commit 解一次。",
|
||||
"队列 = 点对点。一消息一消费者。工作分发、任务处理。\n\n主题 = 发布/订阅。一消息多订阅者。事件广播、扇出。\n\n选队列:负载均衡 workers、订单处理、job queue。\n选主题:多系统响应同事件(如 `user.created` -> email + analytics + CRM)。"
|
||||
],
|
||||
"caveman-es": [
|
||||
"Padre re-render → hijos re-render por default. Props nuevas cada vez (obj/array/fn inline = ref nueva) rompen bail-out.\n\nFix:\n- `React.memo(Child)` para comparar props shallow\n- `useMemo`/`useCallback` en padre para props estables\n- Evitar `<Child data={{...}} onClick={() => ...} />` inline\n\nSin memo, React re-renderiza hijos siempre aunque props iguales.",
|
||||
"Pool = reusar conexiones BD abiertas. Evita handshake TCP + auth por cada request → lento bajo carga.\n\n**Cómo funciona:**\n- App pide conn → pool da una libre (o crea si < max)\n- App termina → conn vuelve al pool, no se cierra\n- Pool lleno → request espera o falla\n\n**Params clave:**\n- `min`: conns calientes siempre\n- `max`: techo (proteger BD de saturación)\n- `idle_timeout`: matar conns ociosas\n- `acquire_timeout`: cuánto esperar conn libre\n\n**Gotchas:**\n- Max muy alto → BD muere (Postgres ~100 conns default)\n- Serverless → usar pooler externo (PgBouncer, Supavisor). Lambdas efímeras rompen pools in-process\n- Transacciones largas bloquean pool → leaks\n\nLibs: HikariCP (Java), pgbouncer, SQLAlchemy pool, `pg` node.",
|
||||
"TCP: conexión, confiable, ordenado, handshake, retransmite pérdidas, control flujo/congestión. Lento pero seguro. Web/SSH/SQL.\n\nUDP: sin conexión, sin garantía, sin orden, cero handshake. Rápido, ligero. Pierde paquetes sin avisar. DNS/video/juegos/VoIP.\n\nClave: TCP = llega todo bien. UDP = llega rápido o no llega.",
|
||||
"Pasos:\n\n1. **Reproducir + medir**: `node --inspect` + Chrome DevTools → Memory tab. O `process.memoryUsage()` en loop.\n2. **Heap snapshots**: tomar 3 (baseline, medio, tarde). DevTools \"Comparison\" → ver qué objetos crecen.\n3. **Sospechosos comunes**:\n - Listeners no removidos (`emitter.on` sin `off`) → `EventEmitter` warning a 10+\n - Closures reteniendo refs grandes\n - Caches sin límite (usar `lru-cache`)\n - Globals/singletons acumulando\n - Timers (`setInterval`) sin `clearInterval`\n - Promesas colgadas reteniendo scope\n4. **Fix**: limpiar listeners en cleanup, TTL en caches, `WeakMap`/`WeakRef` para refs opcionales.\n5. **Verificar**: snapshot post-fix, memoria estable bajo carga (`autocannon`/`k6`).\n\nHerramientas: `clinic.js doctor`, `heapdump`, `--heap-prof`.",
|
||||
"EXPLAIN muestra plan ejecución query. Revela: orden joins, índices usados (o no), método scan (seq/index), filas estimadas, costo. `EXPLAIN ANALYZE` corre query real → tiempos + filas reales vs estimadas. Usar para diagnosticar queries lentas.",
|
||||
"Colisiones hash — dos claves → mismo bucket. Soluciones principales:\n\n**1. Chaining (encadenamiento):** Bucket guarda lista enlazada. Colisión → append a lista. Lookup = hash + scan lista. Java `HashMap` usa esto (lista → árbol si >8).\n\n**2. Open addressing (direccionamiento abierto):** Una entrada por bucket. Colisión → probar otro slot:\n- **Linear probing:** `i+1, i+2...` (rápido, pero clustering)\n- **Quadratic:** `i+1², i+2²...` (menos clusters)\n- **Double hashing:** segundo hash como paso\n\nPython `dict`, Go `map` usan open addressing.\n\n**Trade-offs:** Chaining tolera load factor >1, simple borrado. Open addressing mejor cache locality, pero requiere resize antes (~0.7 load factor) y borrado usa tombstones.",
|
||||
"Browser bloquea request cross-origin. Server no manda header `Access-Control-Allow-Origin` para tu origen.\n\nFix rápido:\n- **Backend**: agregar CORS middleware permitiendo tu origen frontend (ej FastAPI: `CORSMiddleware` con `allow_origins=[\"http://localhost:3000\"]`).\n- **Preflight `OPTIONS` falla** → permitir métodos/headers usados (`allow_methods`, `allow_headers`).\n- **Credentials (cookies)** → `allow_credentials=True` + origen explícito (no `*`).\n- **Dev hack**: proxy en dev server (Vite/Next) para evitar cross-origin.\n\nPegá el error exacto + stack (frontend/backend) para fix preciso.",
|
||||
"Debouncer espera usuario parar tecleo antes disparar búsqueda. Sin él: request por cada tecla → spam API, renders extra, resultados desordenados (race). Con él (~300ms): una sola query tras pausa → menos carga servidor, menos costo, UX suave.",
|
||||
"Rebase: reescribe commits encima de base nueva. Historia lineal, sin merge commit. Cambia SHAs → no usar en ramas compartidas.\n\nMerge: une ramas con commit nuevo. Preserva historia real, no destructivo. Seguro en ramas públicas.\n\nRegla: rebase local, merge público.",
|
||||
"Cola: 1 productor → 1 consumidor (de grupo). Mensaje entregado 1 vez. Work distribution, tareas, jobs.\n\nTópico: 1 productor → N suscriptores. Cada uno recibe copia. Pub/sub, eventos, broadcast.\n\nRegla: ¿trabajo a repartir? cola. ¿evento a anunciar? tópico.\n\nHíbrido (Kafka/SNS+SQS): tópico con consumer groups = ambos. Cada grupo = cola lógica."
|
||||
],
|
||||
"compress": [
|
||||
"React re-renders children by default when the parent renders. To prevent it:\n\n1. **Wrap in `React.memo`** — skips re-render if props are shallow-equal:\n ```jsx\n export default React.memo(MyComponent);\n ```\n\n2. **Stabilize prop references** — objects/arrays/functions created inline break memo. Use `useMemo`/`useCallback` in the parent:\n ```jsx\n const data = useMemo(() => ({...}), [deps]);\n const onClick = useCallback(() => {...}, [deps]);\n ```\n\n3. **Check context** — any component consuming a context re-renders when that context value changes.\n\nMost common cause: passing `{}`, `[]`, or `() => {}` inline as props defeats `memo`.",
|
||||
"Database connection pooling reuses a set of open DB connections instead of opening/closing one per query.\n\n**Why:** Opening a TCP + auth handshake is expensive (often 10–100ms). Pooling amortizes that cost.\n\n**How it works:**\n1. Pool opens N connections at startup.\n2. App \"checks out\" a connection, runs queries, \"returns\" it.\n3. Idle connections stay warm; pool enforces min/max size, timeouts, and health checks.\n\n**Key settings:**\n- `min_size` / `max_size` — floor/ceiling of connections\n- `max_idle_time` — recycle stale conns\n- `acquire_timeout` — fail fast when pool exhausted\n\n**Gotchas:**\n- Pool size > DB `max_connections` → errors. Size per-process; multiply by workers.\n- Long transactions starve the pool.\n- Serverless/short-lived functions → use an external pooler (PgBouncer, RDS Proxy).\n\n**Examples:** SQLAlchemy `QueuePool`, asyncpg `create_pool`, HikariCP (Java), PgBouncer (external).",
|
||||
"TCP is connection-oriented, reliable, ordered, with handshakes and retransmission (web, email, SSH). UDP is connectionless, unreliable, no ordering, low overhead (DNS, video, games).",
|
||||
"Common causes and fixes:\n\n1. **Profile first** — `node --inspect` + Chrome DevTools Memory tab, or `clinic doctor`. Take heap snapshots at intervals and diff them.\n\n2. **Frequent culprits**:\n - **Unbounded caches/Maps** → use `lru-cache` with max size\n - **Event listener leaks** → check `emitter.listenerCount()`; always `removeListener` on cleanup; raise/audit `setMaxListeners`\n - **Closures holding large objects** → avoid capturing big scopes in long-lived callbacks\n - **Global arrays that only grow** (logs, metrics buffers) → cap or flush\n - **Timers/intervals not cleared** → `clearInterval` on shutdown\n - **Promises never settling** → retained forever\n\n3. **Tools**: `heapdump`, `v8.writeHeapSnapshot()`, `process.memoryUsage()`, `--max-old-space-size` to delay OOM while debugging (not a fix).\n\n4. **Pattern**: snapshot → run workload → snapshot → compare retained size by constructor. Look for objects growing linearly with request count.\n\n5. **WeakMap/WeakRef** for caches keyed by objects you don't want to retain.\n\nStart with a heap diff — guessing wastes time.",
|
||||
"SQL `EXPLAIN` shows the query execution plan — how the database will run your query. Key info:\n\n- **Access method**: sequential scan vs index scan vs index-only scan\n- **Join strategy**: nested loop, hash join, merge join\n- **Row estimates**: how many rows the planner expects at each step\n- **Cost estimates**: relative cost units (startup..total)\n- **Order of operations**: which tables/filters/joins run first\n\nUse `EXPLAIN ANALYZE` to actually execute the query and get real timings + actual row counts, which reveals bad estimates and slow steps. Add `BUFFERS` (Postgres) to see cache hits/disk reads.\n\nMain use: finding missing indexes, bad join orders, and costly seq scans.",
|
||||
"Hash tables handle collisions mainly two ways:\n\n1. **Separate chaining**: each bucket stores a linked list (or tree) of entries that hash to the same index. On collision, append to the list. Lookup scans the list.\n\n2. **Open addressing**: store all entries in the array itself; on collision, probe for another slot:\n - *Linear probing*: try i+1, i+2, …\n - *Quadratic probing*: try i+1², i+2², …\n - *Double hashing*: step size from a second hash function.\n\nBoth degrade as load factor rises, so tables resize/rehash past a threshold (e.g., 0.75).",
|
||||
"CORS errors happen when your frontend (e.g. `http://localhost:3000`) makes a request to a backend on a different origin (different scheme/host/port) and the server doesn't return the expected `Access-Control-Allow-*` headers.\n\nCommon causes:\n- **Missing headers**: server doesn't send `Access-Control-Allow-Origin` matching your frontend origin.\n- **Preflight fails**: browser sends `OPTIONS` before `POST`/`PUT`/custom headers; server must respond 2xx with `Allow-Methods` / `Allow-Headers`.\n- **Credentials mismatch**: using `credentials: 'include'` requires `Access-Control-Allow-Credentials: true` AND a specific origin (not `*`).\n- **Wrong URL**: hitting `http://` vs `https://`, or `localhost` vs `127.0.0.1`.\n\nFix (FastAPI example):\n```python\nfrom fastapi.middleware.cors import CORSMiddleware\napp.add_middleware(\n CORSMiddleware,\n allow_origins=[\"http://localhost:3000\"],\n allow_credentials=True,\n allow_methods=[\"*\"],\n allow_headers=[\"*\"],\n)\n```\n\nShare the exact error message + your stack and I'll pinpoint it.",
|
||||
"A debouncer delays running the search until the user stops typing for a short interval (e.g., 300ms). Benefits:\n\n- **Fewer API calls**: avoids firing a request on every keystroke (\"r\", \"re\", \"rea\", \"reac\", \"react\" → 1 call instead of 5).\n- **Lower backend load & cost**: less traffic, fewer DB/LLM hits.\n- **Better UX**: results don't flicker/reorder mid-typing; avoids race conditions where a stale response overwrites a newer one.\n- **Saves client resources**: fewer re-renders and network round-trips.\n\nUse ~200–400ms typically. For instant local filtering, debouncing is less critical; for remote search, it's essentially mandatory.",
|
||||
"**Merge**: combines branches by creating a new merge commit that ties both histories together. Preserves full history, non-destructive.\n\n**Rebase**: replays your commits on top of another branch, creating new commits. Produces linear history, rewrites commits.\n\nRule of thumb: merge for shared/public branches, rebase for local cleanup before pushing.",
|
||||
"Queue = point-to-point, one consumer processes each message (work distribution, task processing). Topic = pub/sub, every subscriber gets a copy (event broadcasting, fan-out).\n\nUse a **queue** when: load-balancing work across workers, ensuring exactly one handler, order/retry matters per job.\n\nUse a **topic** when: multiple independent systems react to the same event, decoupling producers from N consumers, event sourcing/notifications.\n\nHybrid (e.g. Kafka consumer groups, SNS→SQS): topic for fan-out + per-subscriber queue for durability and load-balancing within each group."
|
||||
]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"name": "caveman",
|
||||
"description": "Ultra-compressed communication mode. Cuts 65% of output tokens (measured) while keeping full technical accuracy by speaking like a caveman.",
|
||||
"version": "1.0.1",
|
||||
"contextFileName": "GEMINI.md"
|
||||
}
|
||||
@@ -0,0 +1,81 @@
|
||||
# caveman — installer shim (Windows / PowerShell).
|
||||
#
|
||||
# Thin wrapper around cli/install.js (the unified Node installer). Every flag
|
||||
# you'd pass to cli/install.js can be passed here; we just forward them.
|
||||
#
|
||||
# One-line install:
|
||||
# irm https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.ps1 | iex
|
||||
#
|
||||
# Local clone:
|
||||
# pwsh install.ps1 [flags]
|
||||
#
|
||||
# Why a Node installer? install.sh + install.ps1 used to be parallel sources of
|
||||
# truth and constantly drifted (issue #249 was a `node -e "..."` quoting bug
|
||||
# that silently dropped the JSON merge step on every Windows install). One
|
||||
# Node script works everywhere without quoting bugs.
|
||||
#
|
||||
# Why no top-level param() and everything inside a function? `irm | iex`
|
||||
# executes this file as a string: script-path variables ($PSCommandPath,
|
||||
# $MyInvocation.MyCommand.Path) are $null and a top-level param block cannot
|
||||
# receive arguments through a pipe anyway (issue #565). Wrapping the logic in
|
||||
# a function and forwarding $args keeps one script working for both the pipe
|
||||
# path (no args, no script path) and the local-clone path.
|
||||
|
||||
function Install-Caveman {
|
||||
param(
|
||||
[string[]]$InstallerArgs = @()
|
||||
)
|
||||
|
||||
$ErrorActionPreference = "Stop"
|
||||
$Repo = "JuliusBrussee/caveman"
|
||||
|
||||
# Require Node ≥18.
|
||||
$node = Get-Command node -ErrorAction SilentlyContinue
|
||||
if (-not $node) {
|
||||
Write-Error @"
|
||||
caveman: Node.js (>=18) required. Install:
|
||||
- winget install OpenJS.NodeJS.LTS
|
||||
- or download from https://nodejs.org
|
||||
"@
|
||||
exit 1
|
||||
}
|
||||
|
||||
$nodeMajor = [int](& node -p "process.versions.node.split('.')[0]")
|
||||
if ($nodeMajor -lt 18) {
|
||||
Write-Error "caveman: Node $nodeMajor too old. Need Node >=18. Upgrade: https://nodejs.org"
|
||||
exit 1
|
||||
}
|
||||
|
||||
# If we're inside the repo clone, run the local installer directly.
|
||||
# $PSCommandPath is $null when piped to iex (#565) — the old unguarded
|
||||
# Split-Path on it was the "Cannot bind argument to parameter 'Path'
|
||||
# because it is null" crash.
|
||||
if ($PSCommandPath) {
|
||||
$here = Split-Path -Parent $PSCommandPath
|
||||
$local = Join-Path $here "cli/install.js"
|
||||
if (Test-Path $local) {
|
||||
& node $local @InstallerArgs
|
||||
exit $LASTEXITCODE
|
||||
}
|
||||
}
|
||||
|
||||
# Curl-pipe path: delegate to npx.
|
||||
$npx = Get-Command npx -ErrorAction SilentlyContinue
|
||||
if (-not $npx) {
|
||||
Write-Error "caveman: npx required (ships with Node >=18). Reinstall Node.js."
|
||||
exit 1
|
||||
}
|
||||
|
||||
# Do NOT pass `--` here — npm 7+ npx already forwards trailing args to the
|
||||
# package, and a literal `--` was tripping cli/install.js's parseArgs as an
|
||||
# unknown flag.
|
||||
# npm >=12 defaults allow-git to "none", failing github: specs with
|
||||
# EALLOWGIT (#698). Scope the override to this invocation.
|
||||
$env:NPM_CONFIG_ALLOW_GIT = "all"
|
||||
& npx -y "github:$Repo" @InstallerArgs
|
||||
exit $LASTEXITCODE
|
||||
}
|
||||
|
||||
# $args is the automatic variable: populated when run as a file
|
||||
# (`pwsh install.ps1 --force`), empty under `irm | iex`.
|
||||
Install-Caveman -InstallerArgs $args
|
||||
@@ -0,0 +1,56 @@
|
||||
#!/usr/bin/env bash
|
||||
# caveman — installer shim.
|
||||
#
|
||||
# Thin wrapper around cli/install.js (the unified Node installer). Every flag
|
||||
# you'd pass to cli/install.js can be passed here; we just forward them.
|
||||
#
|
||||
# One-line install:
|
||||
# curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash
|
||||
# curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash -s -- --all
|
||||
#
|
||||
# Local clone:
|
||||
# bash install.sh [flags]
|
||||
#
|
||||
# Why a Node installer? install.sh + install.ps1 used to be parallel sources
|
||||
# of truth and constantly drifted (issue #249, etc.). One Node script works
|
||||
# everywhere without bash/PowerShell quoting bugs.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
REPO="JuliusBrussee/caveman"
|
||||
|
||||
# Require Node ≥18. nvm is a common path; print a hint if missing.
|
||||
if ! command -v node >/dev/null 2>&1; then
|
||||
echo "caveman: Node.js (≥18) required. Install:" >&2
|
||||
echo " macOS: brew install node" >&2
|
||||
echo " Linux: see https://nodejs.org or use nvm (https://github.com/nvm-sh/nvm)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
NODE_MAJOR=$(node -p "process.versions.node.split('.')[0]")
|
||||
if [ "$NODE_MAJOR" -lt 18 ]; then
|
||||
echo "caveman: Node $NODE_MAJOR too old. Need Node ≥18." >&2
|
||||
echo " Upgrade: https://nodejs.org" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# If we're inside the repo clone, run the local installer directly — saves
|
||||
# the npx round-trip and keeps offline installs working. BASH_SOURCE is unset
|
||||
# when bash is invoked from stdin (curl | bash), and `set -u` would trip on a
|
||||
# bare reference — default to empty so the curl-pipe path falls through cleanly.
|
||||
here="$(cd "$(dirname "${BASH_SOURCE[0]:-}")" 2>/dev/null && pwd)" || here=""
|
||||
if [ -n "$here" ] && [ -f "$here/cli/install.js" ]; then
|
||||
exec node "$here/cli/install.js" "$@"
|
||||
fi
|
||||
|
||||
# Curl-pipe path: delegate to npx. We do NOT pass `--` here — npm 7+ npx
|
||||
# already forwards trailing args to the package, and a literal `--` tripped
|
||||
# cli/install.js's parseArgs as an unknown flag.
|
||||
if ! command -v npx >/dev/null 2>&1; then
|
||||
echo "caveman: npx required (ships with Node ≥18). Reinstall Node.js." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# npm >=12 defaults allow-git to "none", failing github: specs with EALLOWGIT
|
||||
# (#698). Scope the override to this one invocation.
|
||||
NPM_CONFIG_ALLOW_GIT=all exec npx -y "github:$REPO" "$@"
|
||||
@@ -0,0 +1,35 @@
|
||||
{
|
||||
"name": "caveman-installer",
|
||||
"version": "0.1.0",
|
||||
"description": "Caveman installer — detects your AI coding agents and installs caveman for each one.",
|
||||
"license": "MIT",
|
||||
"author": "Julius Brussee",
|
||||
"homepage": "https://github.com/JuliusBrussee/caveman",
|
||||
"repository": {
|
||||
"type": "git",
|
||||
"url": "git+https://github.com/JuliusBrussee/caveman.git"
|
||||
},
|
||||
"bugs": {
|
||||
"url": "https://github.com/JuliusBrussee/caveman/issues"
|
||||
},
|
||||
"bin": {
|
||||
"caveman": "./cli/install.js"
|
||||
},
|
||||
"engines": {
|
||||
"node": ">=18"
|
||||
},
|
||||
"scripts": {
|
||||
"test": "node --test tests/installer/*.test.mjs"
|
||||
},
|
||||
"files": [
|
||||
"cli/",
|
||||
"src/",
|
||||
"agents/",
|
||||
"skills/",
|
||||
"plugins/",
|
||||
"commands/",
|
||||
"dist/caveman.skill",
|
||||
"README.md",
|
||||
"LICENSE"
|
||||
]
|
||||
}
|
||||
@@ -26,9 +26,14 @@
|
||||
"Write"
|
||||
],
|
||||
"websiteURL": "https://github.com/JuliusBrussee/caveman",
|
||||
"privacyPolicyURL": "https://github.com/JuliusBrussee/caveman/blob/main/README.md",
|
||||
"termsOfServiceURL": "https://github.com/JuliusBrussee/caveman/blob/main/LICENSE",
|
||||
"defaultPrompt": [
|
||||
"Use caveman mode. Cut filler. Keep technical accuracy."
|
||||
],
|
||||
"composerIcon": "./assets/caveman-small.svg",
|
||||
"logo": "./assets/caveman.svg",
|
||||
"screenshots": [],
|
||||
"brandColor": "#6B7280"
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,47 @@
|
||||
---
|
||||
name: cavecrew-builder
|
||||
description: >
|
||||
Surgical 1-2 file edit. Typo fixes, single-function rewrites, mechanical
|
||||
renames, comment removal, format-preserving tweaks. Hard refuses 3+ file
|
||||
scope. Returns caveman diff receipt. Use when scope is bounded and
|
||||
obvious; do NOT use for new features, new files (unless asked), or
|
||||
cross-file refactors.
|
||||
tools: [Read, Edit, Write, Grep, Glob]
|
||||
---
|
||||
|
||||
Caveman-ultra. Drop articles/filler. Code/paths exact, backticked. No narration.
|
||||
|
||||
## Scope
|
||||
|
||||
1 file ideal. 2 OK. 3+ → refuse.
|
||||
Edit existing only (new file iff user asked).
|
||||
No new abstractions. No drive-by refactors. No comment additions.
|
||||
No `Bash` available — cannot shell out, cannot push, cannot delete.
|
||||
|
||||
## Workflow
|
||||
|
||||
1. `Read` target(s). Never edit blind.
|
||||
2. `Edit` smallest diff that work.
|
||||
3. Re-`Read` to verify.
|
||||
4. Return receipt.
|
||||
|
||||
## Output (receipt)
|
||||
|
||||
```
|
||||
<path:line-range> — <change ≤10 words>.
|
||||
<path:line-range> — <change ≤10 words>.
|
||||
verified: <re-read OK | mismatch @ path:line>.
|
||||
```
|
||||
|
||||
Diff is the artifact. Receipt is the proof. No exploration story.
|
||||
|
||||
## Refusals (terminal lines)
|
||||
|
||||
3+ files → `too-big. split: <n one-line tasks>.`
|
||||
Destructive needed → `needs-confirm. op: <command>.`
|
||||
Spec ambiguous → `ambiguous. ask: <one question>.`
|
||||
Tests fail post-edit, can't fix in scope → `regressed. revert path:line. cause: <fragment>.`
|
||||
|
||||
## Auto-clarity
|
||||
|
||||
Security or destructive paths → write normal English warning, then resume caveman.
|
||||
@@ -0,0 +1,57 @@
|
||||
---
|
||||
name: cavecrew-investigator
|
||||
description: >
|
||||
Read-only code locator. Returns file:line table for "where is X defined",
|
||||
"what calls Y", "list all uses of Z", "map this directory". Output is
|
||||
caveman-compressed so the main thread eats ~60% fewer tokens than
|
||||
vanilla Explore. Refuses to suggest fixes.
|
||||
tools: [Read, Grep, Glob, Bash]
|
||||
model: haiku
|
||||
---
|
||||
|
||||
Caveman-ultra. Drop articles/filler/hedging. Code/symbols/paths exact, backticked. Lead with answer.
|
||||
|
||||
## Job
|
||||
|
||||
Locate. Report. Stop. Never edit, never propose fix.
|
||||
|
||||
## Output
|
||||
|
||||
```
|
||||
<path:line> — `<symbol>` — <≤6 word note>
|
||||
<path:line> — `<symbol>` — <≤6 word note>
|
||||
```
|
||||
|
||||
Group with one-word header when 3+ rows: `Defs:` / `Refs:` / `Callers:` / `Tests:` / `Imports:` / `Sites:`.
|
||||
Single hit → one line, no header.
|
||||
Zero hits → `No match.`
|
||||
Last line → totals: `2 defs, 5 refs.` (omit if 0 or 1).
|
||||
|
||||
## Tools
|
||||
|
||||
`Grep` for symbols/strings. `Glob` for paths. `Read` only specific ranges. `Bash` for `git log -S`/`git grep`/`find` when faster.
|
||||
|
||||
## Refusals
|
||||
|
||||
Asked to fix → `Read-only. Spawn cavecrew-builder.`
|
||||
Asked to design → `Read-only. Spawn cavecrew-builder or use main thread.`
|
||||
|
||||
## Auto-clarity
|
||||
|
||||
Security warnings, destructive ops → write normal English. Resume after.
|
||||
|
||||
## Example
|
||||
|
||||
Q: "where symlink-safe flag write?"
|
||||
|
||||
```
|
||||
Defs:
|
||||
- hooks/caveman-config.js:81 — `safeWriteFlag` — atomic write w/ O_NOFOLLOW
|
||||
- hooks/caveman-config.js:160 — `readFlag` — paired reader
|
||||
Callers:
|
||||
- hooks/caveman-mode-tracker.js:33,87
|
||||
- hooks/caveman-activate.js:40
|
||||
Tests:
|
||||
- tests/test_symlink_flag.js — 12 cases
|
||||
2 defs, 3 callers, 1 test file.
|
||||
```
|
||||
@@ -0,0 +1,48 @@
|
||||
---
|
||||
name: cavecrew-reviewer
|
||||
description: >
|
||||
Diff/branch/file reviewer. One line per finding, severity-tagged, no praise,
|
||||
no scope creep. Output format `path:line: <emoji> <severity>: <problem>. <fix>.`
|
||||
Use for "review this PR", "review my diff", "audit this file". Skips
|
||||
formatting nits unless they change meaning.
|
||||
tools: [Read, Grep, Bash]
|
||||
model: haiku
|
||||
---
|
||||
|
||||
Caveman-ultra. Findings only. No "looks good", no "I'd suggest", no preamble.
|
||||
|
||||
## Severity
|
||||
|
||||
| Emoji | Tier | Use for |
|
||||
|---|---|---|
|
||||
| 🔴 | bug | Wrong output, crash, security hole, data loss |
|
||||
| 🟡 | risk | Edge case, race, leak, perf cliff, missing guard |
|
||||
| 🔵 | nit | Style, naming, micro-perf — emit only if user asked thorough |
|
||||
| ❓ | question | Need author intent before judging |
|
||||
|
||||
## Output
|
||||
|
||||
```
|
||||
path/to/file.ts:42: 🔴 bug: token expiry uses `<` not `<=`. Off-by-one allows expired tokens 1 tick.
|
||||
path/to/file.ts:118: 🟡 risk: pool not closed on error path. Add `try/finally`.
|
||||
src/utils.ts:7: ❓ question: why duplicate `.trim()` here?
|
||||
totals: 1🔴 1🟡 1❓
|
||||
```
|
||||
|
||||
Zero findings → `No issues.`
|
||||
File order, ascending line numbers within file.
|
||||
|
||||
## Boundaries
|
||||
|
||||
- Review only what's in front of you. No "while we're here".
|
||||
- No big-refactor proposals.
|
||||
- Need more context → append `(see L<n> in <file>)`. Don't guess.
|
||||
- Formatting nits skipped unless they change meaning.
|
||||
|
||||
## Tools
|
||||
|
||||
`Bash` only for `git diff`/`git log -p`/`git show`. No mutating commands.
|
||||
|
||||
## Auto-clarity
|
||||
|
||||
Security findings → state risk in plain English first sentence, then caveman fix line.
|
||||
@@ -0,0 +1,7 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 64 64" role="img" aria-label="Caveman">
|
||||
<circle cx="32" cy="32" r="28" fill="#6B7280"/>
|
||||
<path d="M18 40c4-10 24-10 28 0" fill="none" stroke="#F9FAFB" stroke-linecap="round" stroke-width="4"/>
|
||||
<circle cx="24" cy="27" r="4" fill="#F9FAFB"/>
|
||||
<circle cx="40" cy="27" r="4" fill="#F9FAFB"/>
|
||||
<path d="M21 18c3-5 8-8 15-8 5 0 10 2 14 7" fill="none" stroke="#D1D5DB" stroke-linecap="round" stroke-width="4"/>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 471 B |
@@ -0,0 +1,7 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 128 128" role="img" aria-label="Caveman">
|
||||
<rect width="128" height="128" rx="24" fill="#6B7280"/>
|
||||
<path d="M28 79c8-22 48-22 56 0" fill="none" stroke="#F9FAFB" stroke-linecap="round" stroke-width="8"/>
|
||||
<circle cx="48" cy="52" r="8" fill="#F9FAFB"/>
|
||||
<circle cx="80" cy="52" r="8" fill="#F9FAFB"/>
|
||||
<path d="M43 35c6-10 17-15 31-15 12 0 24 5 31 14" fill="none" stroke="#D1D5DB" stroke-linecap="round" stroke-width="8"/>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 487 B |
@@ -0,0 +1,82 @@
|
||||
---
|
||||
name: cavecrew
|
||||
description: >
|
||||
Decision guide for delegating to caveman-style subagents. Tells the main
|
||||
thread WHEN to spawn `cavecrew-investigator` (locate code), `cavecrew-builder`
|
||||
(1-2 file edit), or `cavecrew-reviewer` (diff review) instead of doing the
|
||||
work inline or using vanilla `Explore`. Subagent output is caveman-compressed
|
||||
so the tool-result injected back into main context is ~60% smaller — main
|
||||
context lasts longer across long sessions.
|
||||
Trigger: "delegate to subagent", "use cavecrew", "spawn investigator/builder/reviewer",
|
||||
"save context", "compressed agent output".
|
||||
---
|
||||
|
||||
Cavecrew = three subagent presets that emit caveman output. Same job as Anthropic defaults (`Explore`, edit-style agents, reviewer); difference is the tool-result they return is compressed, so main context shrinks per delegation.
|
||||
|
||||
## When to use cavecrew vs alternatives
|
||||
|
||||
| Task | Use |
|
||||
|---|---|
|
||||
| "Where is X defined / what calls Y / list uses of Z" | `cavecrew-investigator` |
|
||||
| Same but you also want suggestions/architecture commentary | `Explore` (vanilla) |
|
||||
| Surgical edit, ≤2 files, scope obvious | `cavecrew-builder` |
|
||||
| New feature / 3+ files / cross-cutting refactor | Main thread or `feature-dev:code-architect` |
|
||||
| Review diff, branch, or file for bugs | `cavecrew-reviewer` |
|
||||
| Deep code review with rationale + alternatives | `Code Reviewer` (vanilla) |
|
||||
| One-line answer you already know | Main thread, no subagent |
|
||||
|
||||
Rule of thumb: **if you'd want the subagent's output in 1/3 the tokens, pick cavecrew. If you'd want prose, pick vanilla.**
|
||||
|
||||
## Why this exists (the real win)
|
||||
|
||||
Subagent tool results get injected into main context verbatim. A vanilla `Explore` that returns 2k tokens of prose costs 2k tokens of main-context budget every time. The same finding from `cavecrew-investigator` returns ~700 tokens. Across 20 delegations in one session that's the difference between context exhaustion and finishing the task.
|
||||
|
||||
## Output contracts
|
||||
|
||||
What main thread can rely on per agent:
|
||||
|
||||
**`cavecrew-investigator`**
|
||||
```
|
||||
<Header>:
|
||||
- path:line — `symbol` — short note
|
||||
totals: <counts>.
|
||||
```
|
||||
Or `No match.` Always file-path-first, line-number-attached, backticked symbols. Safe to grep with `path:\d+`.
|
||||
|
||||
**`cavecrew-builder`**
|
||||
```
|
||||
<path:line-range> — <change ≤10 words>.
|
||||
verified: <re-read OK | mismatch @ path:line>.
|
||||
```
|
||||
Or one of: `too-big.` / `needs-confirm.` / `ambiguous.` / `regressed.` (terminal first token).
|
||||
|
||||
**`cavecrew-reviewer`**
|
||||
```
|
||||
path:line: <emoji> <severity>: <problem>. <fix>.
|
||||
totals: N🔴 N🟡 N🔵 N❓
|
||||
```
|
||||
Or `No issues.` Findings sorted file → line ascending.
|
||||
|
||||
## Chaining patterns
|
||||
|
||||
**Locate → fix → verify** (most common):
|
||||
1. `cavecrew-investigator` returns site list.
|
||||
2. Main thread picks 1-2 sites, hands paths to `cavecrew-builder`.
|
||||
3. `cavecrew-reviewer` audits the diff.
|
||||
|
||||
**Parallel scout** (when investigation is broad):
|
||||
Spawn 2-3 `cavecrew-investigator` calls in one message (different angles: defs vs callers vs tests). Aggregate in main thread.
|
||||
|
||||
**Single-shot edit** (when site is already known):
|
||||
Skip investigator. Hand exact path:line to `cavecrew-builder` directly.
|
||||
|
||||
## What NOT to do
|
||||
|
||||
- Don't use `cavecrew-builder` when you don't already know the file. Spawn investigator first or main thread will eat tokens passing context.
|
||||
- Don't chain `cavecrew-investigator → cavecrew-builder` for a 5-file refactor. Builder will return `too-big.` and you'll have wasted a turn.
|
||||
- Don't ask `cavecrew-reviewer` for "general feedback" — it returns findings only, no architecture opinions. Use `Code Reviewer` for that.
|
||||
- Don't expect prose. Cavecrew output is structured, sometimes terse to the point of cryptic. If a human will read it directly, paraphrase.
|
||||
|
||||
## Auto-clarity (inherited)
|
||||
|
||||
Subagents drop caveman → normal English for security warnings, irreversible-action confirmations, and any output where fragment ambiguity could be misread. Resume caveman after.
|
||||
@@ -4,14 +4,14 @@ description: >
|
||||
Compress natural language memory files (CLAUDE.md, todos, preferences) into caveman format
|
||||
to save input tokens. Preserves all technical substance, code, URLs, and structure.
|
||||
Compressed version overwrites the original file. Human-readable backup saved as FILE.original.md.
|
||||
Trigger: /caveman-compress <filepath> or "compress memory file"
|
||||
Trigger: /caveman-compress FILEPATH or "compress memory file"
|
||||
---
|
||||
|
||||
# Caveman Compress
|
||||
|
||||
## Purpose
|
||||
|
||||
Compress natural language files (CLAUDE.md, todos, preferences) into caveman-speak to reduce input tokens. Compressed version overwrites original. Human-readable backup saved as `<filename>.original.md`.
|
||||
Compress natural language files (CLAUDE.md, todos, preferences) into caveman-speak to reduce input tokens. Compressed version overwrites original. Human-readable backup saved as `<filename>.original.md`, but NOT beside the source file — it lives in an out-of-tree data dir (`$XDG_DATA_HOME/caveman-compress/backups/<parent-dir-name>/`, or `%LOCALAPPDATA%\caveman-compress\backups\<parent-dir-name>\` on Windows) so skill auto-loaders don't re-ingest it as a live file.
|
||||
|
||||
## Trigger
|
||||
|
||||
@@ -19,19 +19,19 @@ Compress natural language files (CLAUDE.md, todos, preferences) into caveman-spe
|
||||
|
||||
## Process
|
||||
|
||||
1. This SKILL.md lives alongside `memory/` in the same directory. Find that directory.
|
||||
1. The compression scripts live in `scripts/` (adjacent to this SKILL.md). If the path is not immediately available, search for `scripts/__main__.py` next to this SKILL.md.
|
||||
|
||||
2. Run:
|
||||
```
|
||||
cd <directory_containing_this_SKILL.md> && python3 -m scripts <absolute_filepath>
|
||||
```
|
||||
2. From the directory containing this SKILL.md, run:
|
||||
|
||||
python3 -m scripts <absolute_filepath>
|
||||
|
||||
3. The CLI will:
|
||||
- detect file type (no tokens)
|
||||
- call Claude to compress
|
||||
- validate output (no tokens)
|
||||
- if errors: cherry-pick fix with Claude (targeted fixes only, no recompression)
|
||||
- retry up to 2 times
|
||||
- detect file type (no tokens)
|
||||
- call Claude to compress
|
||||
- validate output (no tokens)
|
||||
- if errors: cherry-pick fix with Claude (targeted fixes only, no recompression)
|
||||
- retry up to 2 times
|
||||
- if still failing after 2 retries: report error to user, leave original file untouched
|
||||
|
||||
4. Return result to user
|
||||
|
||||
@@ -103,9 +103,9 @@ Compressed:
|
||||
|
||||
## Boundaries
|
||||
|
||||
- ONLY compress natural language files (.md, .txt, extensionless)
|
||||
- ONLY compress natural language files (.md, .txt, .typ, .typst, .tex, extensionless)
|
||||
- NEVER modify: .py, .js, .ts, .json, .yaml, .yml, .toml, .env, .lock, .css, .html, .xml, .sql, .sh
|
||||
- If file has mixed content (prose + code), compress ONLY the prose sections
|
||||
- If unsure whether something is code or prose, leave it unchanged
|
||||
- Original file is backed up as FILE.original.md before overwriting
|
||||
- Original file is backed up as FILE.original.md before overwriting — in the out-of-tree backup data dir (see Purpose), not beside the source file
|
||||
- Never compress FILE.original.md (skip it)
|
||||
@@ -0,0 +1,9 @@
|
||||
"""Caveman compress scripts.
|
||||
|
||||
This package provides tools to compress natural language markdown files
|
||||
into caveman format to save input tokens.
|
||||
"""
|
||||
|
||||
__all__ = ["cli", "compress", "detect", "validate"]
|
||||
|
||||
__version__ = "1.0.0"
|
||||
@@ -11,7 +11,7 @@ except ImportError:
|
||||
|
||||
try:
|
||||
import tiktoken
|
||||
_enc = tiktoken.get_encoding("cl100k_base")
|
||||
_enc = tiktoken.get_encoding("o200k_base")
|
||||
except ImportError:
|
||||
_enc = None
|
||||
|
||||
@@ -23,12 +23,12 @@ def count_tokens(text):
|
||||
|
||||
|
||||
def benchmark_pair(orig_path: Path, comp_path: Path):
|
||||
orig_text = orig_path.read_text()
|
||||
comp_text = comp_path.read_text()
|
||||
orig_text = orig_path.read_text(encoding="utf-8", errors="ignore")
|
||||
comp_text = comp_path.read_text(encoding="utf-8", errors="ignore")
|
||||
|
||||
orig_tokens = count_tokens(orig_text)
|
||||
comp_tokens = count_tokens(comp_text)
|
||||
saved = 100 * (orig_tokens - comp_tokens) / orig_tokens
|
||||
saved = 100 * (orig_tokens - comp_tokens) / orig_tokens if orig_tokens > 0 else 0.0
|
||||
result = validate(orig_path, comp_path)
|
||||
|
||||
return (comp_path.name, orig_tokens, comp_tokens, saved, result.is_valid)
|
||||
@@ -44,8 +44,8 @@ def print_table(rows):
|
||||
def main():
|
||||
# Direct file pair: python3 benchmark.py original.md compressed.md
|
||||
if len(sys.argv) == 3:
|
||||
orig = Path(sys.argv[1])
|
||||
comp = Path(sys.argv[2])
|
||||
orig = Path(sys.argv[1]).resolve()
|
||||
comp = Path(sys.argv[2]).resolve()
|
||||
if not orig.exists():
|
||||
print(f"❌ Not found: {orig}")
|
||||
sys.exit(1)
|
||||
@@ -56,7 +56,9 @@ def main():
|
||||
return
|
||||
|
||||
# Glob mode: repo_root/tests/caveman-compress/
|
||||
tests_dir = Path(__file__).parent.parent.parent / "tests" / "caveman-compress"
|
||||
# __file__ lives at <repo_root>/skills/caveman-compress/scripts/benchmark.py
|
||||
# Walk up four dirs: scripts → caveman-compress → skills → repo_root.
|
||||
tests_dir = Path(__file__).resolve().parents[3] / "tests" / "caveman-compress"
|
||||
if not tests_dir.exists():
|
||||
print(f"❌ Tests dir not found: {tests_dir}")
|
||||
sys.exit(1)
|
||||
@@ -1,15 +1,27 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Caveman Memory CLI
|
||||
Caveman Compress CLI
|
||||
|
||||
Usage:
|
||||
caveman <filepath>
|
||||
"""
|
||||
|
||||
import sys
|
||||
|
||||
# Force UTF-8 on stdout/stderr before any code can print. Windows consoles
|
||||
# default to cp1252 and crash on the ❌ glyphs in error/validation branches,
|
||||
# masking the real error and leaving the user with a half-compressed file.
|
||||
for _stream in (sys.stdout, sys.stderr):
|
||||
reconfigure = getattr(_stream, "reconfigure", None)
|
||||
if callable(reconfigure):
|
||||
try:
|
||||
reconfigure(encoding="utf-8", errors="replace")
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
from .compress import compress_file
|
||||
from .compress import backup_dir_for, compress_file
|
||||
from .detect import detect_file_type, should_compress
|
||||
|
||||
|
||||
@@ -33,6 +45,8 @@ def main():
|
||||
print(f"❌ Not a file: {filepath}")
|
||||
sys.exit(1)
|
||||
|
||||
filepath = filepath.resolve()
|
||||
|
||||
# Detect file type
|
||||
file_type = detect_file_type(filepath)
|
||||
|
||||
@@ -50,7 +64,7 @@ def main():
|
||||
|
||||
if success:
|
||||
print("\nCompression completed successfully")
|
||||
backup_path = filepath.with_name(filepath.stem + ".original.md")
|
||||
backup_path = backup_dir_for(filepath) / (filepath.stem + ".original.md")
|
||||
print(f"Compressed: {filepath}")
|
||||
print(f"Original: {backup_path}")
|
||||
sys.exit(0)
|
||||
@@ -0,0 +1,414 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Caveman Memory Compression Orchestrator
|
||||
|
||||
Usage:
|
||||
python scripts/compress.py <filepath>
|
||||
"""
|
||||
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import stat
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
from typing import List
|
||||
|
||||
OUTER_FENCE_REGEX = re.compile(
|
||||
r"\A\s*(`{3,}|~{3,})[^\n]*\n(.*)\n\1\s*\Z", re.DOTALL
|
||||
)
|
||||
|
||||
# YAML frontmatter: starts at file start with --- on its own line, ends with --- on its own line.
|
||||
# Captures the entire block (including delimiters and trailing newline) and the body after.
|
||||
FRONTMATTER_REGEX = re.compile(
|
||||
r"\A(---\r?\n.*?\r?\n---\r?\n)(.*)", re.DOTALL
|
||||
)
|
||||
|
||||
|
||||
def split_frontmatter(text: str):
|
||||
"""Split YAML frontmatter from body. Returns (frontmatter, body).
|
||||
|
||||
Memory files (and many other markdown docs) start with a YAML frontmatter
|
||||
block delimited by `---` lines. The compression LLM has a habit of stripping
|
||||
or rewriting these despite preserve-structure rules in the prompt — so we
|
||||
surgically remove the frontmatter before compression and prepend it back
|
||||
verbatim to the output. Files without frontmatter pass through unchanged.
|
||||
"""
|
||||
m = FRONTMATTER_REGEX.match(text)
|
||||
if m:
|
||||
return m.group(1), m.group(2)
|
||||
return "", text
|
||||
|
||||
# Filenames and paths that almost certainly hold secrets or PII. Compressing
|
||||
# them ships raw bytes to the Anthropic API — a third-party data boundary that
|
||||
# developers on sensitive codebases cannot cross. detect.py already skips .env
|
||||
# by extension, but credentials.md / secrets.txt / ~/.aws/credentials would
|
||||
# slip through the natural-language filter. This is a hard refuse before read.
|
||||
SENSITIVE_BASENAME_REGEX = re.compile(
|
||||
r"(?ix)^("
|
||||
r"\.env(\..+)?"
|
||||
r"|\.netrc"
|
||||
r"|credentials(\..+)?"
|
||||
r"|secrets?(\..+)?"
|
||||
r"|passwords?(\..+)?"
|
||||
r"|id_(rsa|dsa|ecdsa|ed25519)(\.pub)?"
|
||||
r"|authorized_keys"
|
||||
r"|known_hosts"
|
||||
r"|.*\.(pem|key|p12|pfx|crt|cer|jks|keystore|asc|gpg)"
|
||||
r")$"
|
||||
)
|
||||
|
||||
SENSITIVE_PATH_COMPONENTS = frozenset({".ssh", ".aws", ".gnupg", ".kube", ".docker"})
|
||||
|
||||
SENSITIVE_NAME_TOKENS = (
|
||||
"secret", "credential", "password", "passwd",
|
||||
"apikey", "accesskey", "token", "privatekey",
|
||||
)
|
||||
|
||||
|
||||
def backup_dir_for(filepath: Path) -> Path:
|
||||
"""Resolve the out-of-tree backup directory for a given source file.
|
||||
|
||||
Backups must live OUTSIDE the source directory so skill auto-loaders
|
||||
(Claude Code rules/, opencode instructions/, etc.) stop re-ingesting the
|
||||
`.original.md` copies as live files. Base dir is platform-aware:
|
||||
- Windows: %LOCALAPPDATA%\\caveman-compress\\backups
|
||||
- else: $XDG_DATA_HOME/caveman-compress/backups if set,
|
||||
else ~/.local/share/caveman-compress/backups
|
||||
|
||||
The source file's parent-dir name is mirrored under the base to reduce
|
||||
cross-project collisions (e.g. two `task.md` files in different repos).
|
||||
"""
|
||||
if os.name == "nt" or sys.platform == "win32":
|
||||
local_appdata = os.environ.get("LOCALAPPDATA")
|
||||
base = Path(local_appdata) if local_appdata else Path.home() / "AppData" / "Local"
|
||||
base = base / "caveman-compress" / "backups"
|
||||
else:
|
||||
xdg = os.environ.get("XDG_DATA_HOME")
|
||||
base = Path(xdg) if xdg else Path.home() / ".local" / "share"
|
||||
base = base / "caveman-compress" / "backups"
|
||||
return base / filepath.parent.name
|
||||
|
||||
|
||||
def is_sensitive_path(filepath: Path) -> bool:
|
||||
"""Heuristic denylist for files that must never be shipped to a third-party API."""
|
||||
name = filepath.name
|
||||
if SENSITIVE_BASENAME_REGEX.match(name):
|
||||
return True
|
||||
lowered_parts = {p.lower() for p in filepath.parts}
|
||||
if lowered_parts & SENSITIVE_PATH_COMPONENTS:
|
||||
return True
|
||||
# Normalize separators so "api-key" and "api_key" both match "apikey".
|
||||
lower = re.sub(r"[_\-\s.]", "", name.lower())
|
||||
return any(tok in lower for tok in SENSITIVE_NAME_TOKENS)
|
||||
|
||||
|
||||
def strip_llm_wrapper(text: str) -> str:
|
||||
"""Strip outer ```markdown ... ``` fence when it wraps the entire output."""
|
||||
m = OUTER_FENCE_REGEX.match(text)
|
||||
if m:
|
||||
return m.group(2)
|
||||
return text
|
||||
|
||||
|
||||
def write_text_atomic(path: Path, text: str) -> None:
|
||||
"""Write ``text`` to ``path`` atomically as UTF-8.
|
||||
|
||||
Path.write_text() truncates the destination before encoding the string —
|
||||
a UnicodeEncodeError (or any other failure) partway through leaves a
|
||||
0-byte file, destroying whatever was there before (issue #655). Encode
|
||||
first, write the bytes to a sibling temp file, fsync, then os.replace()
|
||||
so the destination only ever moves from one complete, valid file to
|
||||
another. Preserves the original file's permission bits across the swap.
|
||||
"""
|
||||
data = text.encode("utf-8")
|
||||
fd, tmp_name = tempfile.mkstemp(
|
||||
dir=str(path.parent), prefix=path.name + ".", suffix=".tmp"
|
||||
)
|
||||
tmp_path = Path(tmp_name)
|
||||
try:
|
||||
with os.fdopen(fd, "wb") as f:
|
||||
f.write(data)
|
||||
f.flush()
|
||||
os.fsync(f.fileno())
|
||||
if path.exists():
|
||||
os.chmod(tmp_path, stat.S_IMODE(path.stat().st_mode))
|
||||
os.replace(tmp_path, path)
|
||||
except Exception:
|
||||
try:
|
||||
tmp_path.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
raise
|
||||
|
||||
|
||||
def first_nonblank_line(text: str) -> str:
|
||||
"""Return the first non-blank line, stripped — used to detect a prose
|
||||
preamble smuggled in ahead of the real content (issue #588)."""
|
||||
for line in text.splitlines():
|
||||
if line.strip():
|
||||
return line.strip()
|
||||
return ""
|
||||
|
||||
|
||||
def _write_target(filepath: Path, text: str, backup_path: Path) -> None:
|
||||
"""Write to the target file, surfacing the backup location if the write
|
||||
itself fails. write_text_atomic already leaves the target untouched on
|
||||
failure, but the caller still needs to know where the pre-compression
|
||||
original lives instead of being left to guess (issue #652)."""
|
||||
try:
|
||||
write_text_atomic(filepath, text)
|
||||
except Exception:
|
||||
print(f"❌ Write to {filepath} failed. Original preserved at backup: {backup_path}")
|
||||
raise
|
||||
|
||||
|
||||
from .detect import should_compress
|
||||
from .validate import validate
|
||||
|
||||
MAX_RETRIES = 2
|
||||
|
||||
|
||||
# ---------- Claude Calls ----------
|
||||
|
||||
|
||||
def call_claude(prompt: str) -> str:
|
||||
"""Send a prompt to Claude.
|
||||
|
||||
Prefers the Anthropic SDK when ANTHROPIC_API_KEY is set; otherwise falls
|
||||
back to the ``claude --print`` CLI (which handles desktop auth).
|
||||
|
||||
On Windows the CLI subprocess decoding defaults to the system codepage
|
||||
(cp1251 / cp1252) and crashes on UTF-8 output — see issue #152. Pinning
|
||||
``encoding="utf-8"`` with ``errors="replace"`` matches the CLI's actual
|
||||
native I/O and prevents the UnicodeDecodeError before validation can
|
||||
report. Windows users with non-ASCII content can also set
|
||||
``ANTHROPIC_API_KEY`` to route through the SDK and skip the subprocess.
|
||||
"""
|
||||
api_key = os.environ.get("ANTHROPIC_API_KEY")
|
||||
if api_key:
|
||||
try:
|
||||
import anthropic
|
||||
|
||||
client = anthropic.Anthropic(api_key=api_key)
|
||||
msg = client.messages.create(
|
||||
model=os.environ.get("CAVEMAN_MODEL", "claude-sonnet-4-5"),
|
||||
max_tokens=8192,
|
||||
messages=[{"role": "user", "content": prompt}],
|
||||
)
|
||||
return strip_llm_wrapper(msg.content[0].text.strip())
|
||||
except ImportError:
|
||||
pass # anthropic not installed, fall back to CLI
|
||||
# Fallback: use claude CLI (handles desktop auth).
|
||||
# Resolve binary via shutil.which so Windows .cmd/.bat shims (e.g.
|
||||
# %APPDATA%\npm\claude.CMD) work without shell=True. On POSIX,
|
||||
# shutil.which returns the same absolute path as the implicit lookup,
|
||||
# so this is a no-op there. Falls back to bare "claude" if not found
|
||||
# on PATH so subprocess raises a clear FileNotFoundError.
|
||||
claude_bin = shutil.which("claude") or "claude"
|
||||
try:
|
||||
result = subprocess.run(
|
||||
[claude_bin, "--print"],
|
||||
input=prompt,
|
||||
text=True,
|
||||
capture_output=True,
|
||||
check=True,
|
||||
encoding="utf-8",
|
||||
errors="replace",
|
||||
)
|
||||
return strip_llm_wrapper(result.stdout.strip())
|
||||
except subprocess.CalledProcessError as e:
|
||||
raise RuntimeError(f"Claude call failed:\n{e.stderr}")
|
||||
|
||||
|
||||
def build_compress_prompt(original: str) -> str:
|
||||
return f"""
|
||||
Compress this markdown into caveman format.
|
||||
|
||||
STRICT RULES:
|
||||
- Do NOT modify anything inside ``` code blocks
|
||||
- Do NOT modify anything inside inline backticks
|
||||
- Preserve ALL URLs exactly
|
||||
- Preserve ALL headings exactly
|
||||
- Preserve file paths and commands
|
||||
- Return ONLY the compressed markdown body — do NOT wrap the entire output in a ```markdown fence or any other fence. Inner code blocks from the original stay as-is; do not add a new outer fence around the whole file.
|
||||
|
||||
Only compress natural language.
|
||||
|
||||
TEXT:
|
||||
{original}
|
||||
"""
|
||||
|
||||
|
||||
def build_fix_prompt(original: str, compressed: str, errors: List[str]) -> str:
|
||||
errors_str = "\n".join(f"- {e}" for e in errors)
|
||||
return f"""You are fixing a caveman-compressed markdown file. Specific validation errors were found.
|
||||
|
||||
CRITICAL RULES:
|
||||
- DO NOT recompress or rephrase the file
|
||||
- ONLY fix the listed errors — leave everything else exactly as-is
|
||||
- The ORIGINAL is provided as reference only (to restore missing content)
|
||||
- Preserve caveman style in all untouched sections
|
||||
|
||||
ERRORS TO FIX:
|
||||
{errors_str}
|
||||
|
||||
HOW TO FIX:
|
||||
- Missing URL: find it in ORIGINAL, restore it exactly where it belongs in COMPRESSED
|
||||
- Code block mismatch: find the exact code block in ORIGINAL, restore it in COMPRESSED
|
||||
- Heading mismatch: restore the exact heading text from ORIGINAL into COMPRESSED
|
||||
- Do not touch any section not mentioned in the errors
|
||||
|
||||
ORIGINAL (reference only):
|
||||
{original}
|
||||
|
||||
COMPRESSED (fix this):
|
||||
{compressed}
|
||||
|
||||
Return ONLY the fixed compressed file. No explanation.
|
||||
"""
|
||||
|
||||
|
||||
# ---------- Core Logic ----------
|
||||
|
||||
|
||||
def compress_file(filepath: Path) -> bool:
|
||||
# Resolve and validate path
|
||||
filepath = filepath.resolve()
|
||||
MAX_FILE_SIZE = 500_000 # 500KB
|
||||
if not filepath.exists():
|
||||
raise FileNotFoundError(f"File not found: {filepath}")
|
||||
if filepath.stat().st_size > MAX_FILE_SIZE:
|
||||
raise ValueError(f"File too large to compress safely (max 500KB): {filepath}")
|
||||
|
||||
# Refuse files that look like they contain secrets or PII. Compressing ships
|
||||
# the raw bytes to the Anthropic API — a third-party boundary — so we fail
|
||||
# loudly rather than silently exfiltrate credentials or keys. Override is
|
||||
# intentional: the user must rename the file if the heuristic is wrong.
|
||||
if is_sensitive_path(filepath):
|
||||
raise ValueError(
|
||||
f"Refusing to compress {filepath}: filename looks sensitive "
|
||||
"(credentials, keys, secrets, or known private paths). "
|
||||
"Compression sends file contents to the Anthropic API. "
|
||||
"Rename the file if this is a false positive."
|
||||
)
|
||||
|
||||
print(f"Processing: {filepath}")
|
||||
|
||||
if not should_compress(filepath):
|
||||
print("Skipping (not natural language)")
|
||||
return False
|
||||
|
||||
original_text = filepath.read_text(encoding="utf-8", errors="ignore")
|
||||
# Store backup outside the source directory so skill auto-loaders don't
|
||||
# re-ingest the `.original.md` copy as a live file. Mirror the source's
|
||||
# parent-dir name + stem under a platform-aware base to reduce collisions.
|
||||
backup_dir = backup_dir_for(filepath)
|
||||
backup_dir.mkdir(parents=True, exist_ok=True)
|
||||
backup_path = backup_dir / (filepath.stem + ".original.md")
|
||||
|
||||
if not original_text.strip():
|
||||
print("❌ Refusing to compress: file is empty or whitespace-only.")
|
||||
return False
|
||||
|
||||
# Check if backup already exists to prevent accidental overwriting
|
||||
if backup_path.exists():
|
||||
print(f"⚠️ Backup file already exists: {backup_path}")
|
||||
print("The original backup may contain important content.")
|
||||
print("Aborting to prevent data loss. Please remove or rename the backup file if you want to proceed.")
|
||||
return False
|
||||
|
||||
# Split YAML frontmatter off before compression. Claude tends to strip or
|
||||
# rewrite frontmatter despite preserve-structure rules; we keep it verbatim
|
||||
# by removing it from the input and re-prepending it to the output.
|
||||
frontmatter, body = split_frontmatter(original_text)
|
||||
if frontmatter:
|
||||
print(f"Detected YAML frontmatter ({len(frontmatter)} chars) — preserving verbatim")
|
||||
|
||||
if not body.strip():
|
||||
print("❌ Refusing to compress: body is empty after frontmatter removal.")
|
||||
return False
|
||||
|
||||
# Step 1: Compress (body only, frontmatter excluded)
|
||||
print("Compressing with Claude...")
|
||||
compressed_body = call_claude(build_compress_prompt(body))
|
||||
|
||||
if compressed_body is None or not compressed_body.strip():
|
||||
print("❌ Compression aborted: Claude returned an empty response.")
|
||||
print(" Original file is untouched (no backup created).")
|
||||
return False
|
||||
|
||||
# Compare the BODY (not the whole file) — frontmatter is preserved verbatim
|
||||
# and would never change, so identity must be judged on the compressible part.
|
||||
if compressed_body.strip() == body.strip():
|
||||
print("❌ Compression aborted: output is identical to input.")
|
||||
print(" Likely causes: Claude refused, returned the prompt verbatim, or the file is")
|
||||
print(" already in caveman form. Original file is untouched (no backup created).")
|
||||
return False
|
||||
|
||||
# Reassemble: frontmatter (verbatim) + compressed body
|
||||
compressed = frontmatter + compressed_body
|
||||
|
||||
# Save original as backup, then verify the backup readback before
|
||||
# touching the input file. If the filesystem dropped bytes (encoding,
|
||||
# antivirus, disk full), unlink the bad backup and abort instead of
|
||||
# leaving the user with a corrupt backup + compressed primary.
|
||||
write_text_atomic(backup_path, original_text)
|
||||
backup_readback = backup_path.read_text(encoding="utf-8", errors="ignore")
|
||||
if backup_readback != original_text:
|
||||
print(f"❌ Backup write verification failed: {backup_path}")
|
||||
print(" In-memory original differs from on-disk backup. Aborting before touching the input file.")
|
||||
try:
|
||||
backup_path.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
_write_target(filepath, compressed, backup_path)
|
||||
|
||||
# Step 2: Validate + Retry
|
||||
for attempt in range(MAX_RETRIES):
|
||||
print(f"\nValidation attempt {attempt + 1}")
|
||||
|
||||
result = validate(backup_path, filepath)
|
||||
|
||||
if result.is_valid:
|
||||
print("Validation passed")
|
||||
break
|
||||
|
||||
print("❌ Validation failed:")
|
||||
for err in result.errors:
|
||||
print(f" - {err}")
|
||||
|
||||
if attempt == MAX_RETRIES - 1:
|
||||
# Restore original on failure
|
||||
_write_target(filepath, original_text, backup_path)
|
||||
backup_path.unlink(missing_ok=True)
|
||||
print("❌ Failed after retries — original restored")
|
||||
return False
|
||||
|
||||
print("Fixing with Claude...")
|
||||
compressed = call_claude(
|
||||
build_fix_prompt(original_text, compressed, result.errors)
|
||||
)
|
||||
|
||||
if compressed is None or not compressed.strip():
|
||||
print("❌ Fix attempt aborted: Claude returned an empty response.")
|
||||
print(" Skipping this attempt.")
|
||||
continue
|
||||
|
||||
# Guard against a prose preamble smuggled in ahead of the real fixed
|
||||
# content (issue #588). Only enforced when the original starts with a
|
||||
# structural anchor (frontmatter `---` or a heading) — plain-prose
|
||||
# first lines get legitimately rewritten by compression, and requiring
|
||||
# them verbatim would reject every valid fix.
|
||||
anchor = first_nonblank_line(original_text)
|
||||
if anchor.startswith(("---", "#")) and first_nonblank_line(compressed) != anchor:
|
||||
print("❌ Fix attempt aborted: output does not start with the original's first line.")
|
||||
print(" Possible preamble leak. Skipping this attempt.")
|
||||
continue
|
||||
|
||||
_write_target(filepath, compressed, backup_path)
|
||||
|
||||
return True
|
||||
@@ -6,7 +6,7 @@ import re
|
||||
from pathlib import Path
|
||||
|
||||
# Extensions that are natural language and compressible
|
||||
COMPRESSIBLE_EXTENSIONS = {".md", ".txt", ".markdown", ".rst"}
|
||||
COMPRESSIBLE_EXTENSIONS = {".md", ".txt", ".markdown", ".rst", ".typ", ".typst", ".tex"}
|
||||
|
||||
# Extensions that are code/config and should be skipped
|
||||
SKIP_EXTENSIONS = {
|
||||
@@ -17,6 +17,16 @@ SKIP_EXTENSIONS = {
|
||||
".dockerfile", ".makefile", ".csv", ".ini", ".cfg",
|
||||
}
|
||||
|
||||
# Well-known build/config files that carry no (or a misleading) extension —
|
||||
# `Dockerfile` has no suffix so `.dockerfile` above never matches it, and
|
||||
# `CMakeLists.txt` would ride the compressible `.txt` rule. Checked by
|
||||
# basename before any extension rule.
|
||||
KNOWN_CODE_FILENAMES = {
|
||||
"dockerfile", "makefile", "gnumakefile", "jenkinsfile", "vagrantfile",
|
||||
"rakefile", "gemfile", "justfile", "procfile", "brewfile",
|
||||
"cmakelists.txt",
|
||||
}
|
||||
|
||||
# Patterns that indicate a line is code
|
||||
CODE_PATTERNS = [
|
||||
re.compile(r"^\s*(import |from .+ import |require\(|const |let |var )"),
|
||||
@@ -67,6 +77,10 @@ def detect_file_type(filepath: Path) -> str:
|
||||
"""
|
||||
ext = filepath.suffix.lower()
|
||||
|
||||
# Known code filenames win over any extension rule
|
||||
if filepath.name.lower() in KNOWN_CODE_FILENAMES:
|
||||
return "code"
|
||||
|
||||
# Extension-based classification
|
||||
if ext in COMPRESSIBLE_EXTENSIONS:
|
||||
return "natural_language"
|
||||
@@ -76,12 +90,16 @@ def detect_file_type(filepath: Path) -> str:
|
||||
# Extensionless files (like CLAUDE.md, TODO) — check content
|
||||
if not ext:
|
||||
try:
|
||||
text = filepath.read_text(errors="ignore")
|
||||
text = filepath.read_text(encoding="utf-8", errors="ignore")
|
||||
except (OSError, PermissionError):
|
||||
return "unknown"
|
||||
|
||||
lines = text.splitlines()[:50]
|
||||
|
||||
# Shebang means executable script, never prose
|
||||
if text.startswith("#!"):
|
||||
return "code"
|
||||
|
||||
if _is_json_content(text[:10000]):
|
||||
return "config"
|
||||
if _is_yaml_content(lines):
|
||||
@@ -115,7 +133,7 @@ if __name__ == "__main__":
|
||||
sys.exit(1)
|
||||
|
||||
for path_str in sys.argv[1:]:
|
||||
p = Path(path_str)
|
||||
p = Path(path_str).resolve()
|
||||
file_type = detect_file_type(p)
|
||||
compress = should_compress(p)
|
||||
print(f" {p.name:30s} type={file_type:20s} compress={compress}")
|
||||
@@ -1,14 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
import re
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
|
||||
URL_REGEX = re.compile(r"https?://[^\s)]+")
|
||||
CODE_BLOCK_REGEX = re.compile(r"```.*?```", re.DOTALL)
|
||||
FENCE_OPEN_REGEX = re.compile(r"^(\s{0,3})(`{3,}|~{3,})(.*)$")
|
||||
HEADING_REGEX = re.compile(r"^(#{1,6})\s+(.*)", re.MULTILINE)
|
||||
BULLET_REGEX = re.compile(r"^\s*[-*+]\s+", re.MULTILINE)
|
||||
|
||||
# crude but effective path detection
|
||||
PATH_REGEX = re.compile(r"(\./|\../|/|[A-Za-z]:\\)[\w\-/\\\.]+")
|
||||
# Requires either a path prefix (./ ../ / or drive letter) or a slash/backslash within the match
|
||||
PATH_REGEX = re.compile(r"(?:\./|\.\./|/|[A-Za-z]:\\)[\w\-/\\\.]+|[\w\-\.]+[/\\][\w\-/\\\.]+")
|
||||
|
||||
|
||||
class ValidationResult:
|
||||
@@ -26,7 +28,7 @@ class ValidationResult:
|
||||
|
||||
|
||||
def read_file(path: Path) -> str:
|
||||
return path.read_text(errors="ignore")
|
||||
return path.read_text(encoding="utf-8")
|
||||
|
||||
|
||||
# ---------- Extractors ----------
|
||||
@@ -37,7 +39,47 @@ def extract_headings(text):
|
||||
|
||||
|
||||
def extract_code_blocks(text):
|
||||
return CODE_BLOCK_REGEX.findall(text)
|
||||
"""Line-based fenced code block extractor.
|
||||
|
||||
Handles ``` and ~~~ fences with variable length (CommonMark: closing
|
||||
fence must use same char and be at least as long as opening). Supports
|
||||
nested fences (e.g. an outer 4-backtick block wrapping inner 3-backtick
|
||||
content).
|
||||
"""
|
||||
blocks = []
|
||||
lines = text.split("\n")
|
||||
i = 0
|
||||
n = len(lines)
|
||||
while i < n:
|
||||
m = FENCE_OPEN_REGEX.match(lines[i])
|
||||
if not m:
|
||||
i += 1
|
||||
continue
|
||||
fence_char = m.group(2)[0]
|
||||
fence_len = len(m.group(2))
|
||||
open_line = lines[i]
|
||||
block_lines = [open_line]
|
||||
i += 1
|
||||
closed = False
|
||||
while i < n:
|
||||
close_m = FENCE_OPEN_REGEX.match(lines[i])
|
||||
if (
|
||||
close_m
|
||||
and close_m.group(2)[0] == fence_char
|
||||
and len(close_m.group(2)) >= fence_len
|
||||
and close_m.group(3).strip() == ""
|
||||
):
|
||||
block_lines.append(lines[i])
|
||||
closed = True
|
||||
i += 1
|
||||
break
|
||||
block_lines.append(lines[i])
|
||||
i += 1
|
||||
if closed:
|
||||
blocks.append("\n".join(block_lines))
|
||||
# Unclosed fences are silently skipped — they indicate malformed markdown
|
||||
# and including them would cause false-positive validation failures.
|
||||
return blocks
|
||||
|
||||
|
||||
def extract_urls(text):
|
||||
@@ -52,6 +94,20 @@ def count_bullets(text):
|
||||
return len(BULLET_REGEX.findall(text))
|
||||
|
||||
|
||||
def extract_inline_codes(text):
|
||||
"""Backtick-delimited inline spans, with fenced code blocks stripped first.
|
||||
|
||||
Previously used a column-0-anchored regex to strip fences, which misses
|
||||
fences indented 1-3 spaces (valid CommonMark). Reuse extract_code_blocks
|
||||
(FENCE_OPEN_REGEX-based, indentation-aware) instead so an indented fence's
|
||||
body backticks don't leak into inline-code pairing.
|
||||
"""
|
||||
text_without_fences = text
|
||||
for block in extract_code_blocks(text):
|
||||
text_without_fences = text_without_fences.replace(block, "", 1)
|
||||
return re.findall(r"`([^`]+)`", text_without_fences)
|
||||
|
||||
|
||||
# ---------- Validators ----------
|
||||
|
||||
|
||||
@@ -103,6 +159,22 @@ def validate_bullets(orig, comp, result):
|
||||
result.add_warning(f"Bullet count changed too much: {b1} -> {b2}")
|
||||
|
||||
|
||||
def validate_inline_codes(orig, comp, result):
|
||||
c1 = Counter(extract_inline_codes(orig))
|
||||
c2 = Counter(extract_inline_codes(comp))
|
||||
|
||||
if c1 != c2:
|
||||
lost = set(c1.keys()) - set(c2.keys())
|
||||
added = set(c2.keys()) - set(c1.keys())
|
||||
for code, count in c1.items():
|
||||
if code in c2 and c2[code] < count:
|
||||
lost.add(f"{code} (lost {count - c2[code]} of {count} occurrences)")
|
||||
if lost:
|
||||
result.add_error(f"Inline code lost: {lost}")
|
||||
if added:
|
||||
result.add_warning(f"Inline code added: {added}")
|
||||
|
||||
|
||||
# ---------- Main ----------
|
||||
|
||||
|
||||
@@ -117,6 +189,7 @@ def validate(original_path: Path, compressed_path: Path) -> ValidationResult:
|
||||
validate_urls(orig, comp, result)
|
||||
validate_paths(orig, comp, result)
|
||||
validate_bullets(orig, comp, result)
|
||||
validate_inline_codes(orig, comp, result)
|
||||
|
||||
return result
|
||||
|
||||
@@ -130,8 +203,8 @@ if __name__ == "__main__":
|
||||
print("Usage: python validate.py <original> <compressed>")
|
||||
sys.exit(1)
|
||||
|
||||
orig = Path(sys.argv[1])
|
||||
comp = Path(sys.argv[2])
|
||||
orig = Path(sys.argv[1]).resolve()
|
||||
comp = Path(sys.argv[2]).resolve()
|
||||
|
||||
res = validate(orig, comp)
|
||||
|
||||
@@ -0,0 +1,12 @@
|
||||
---
|
||||
name: caveman-stats
|
||||
description: >
|
||||
Show real token usage and estimated savings for the current session.
|
||||
Reads directly from the Claude Code session log — no AI estimation.
|
||||
Triggers on /caveman-stats. Output is injected by the mode-tracker hook;
|
||||
the model itself does not compute the numbers.
|
||||
---
|
||||
|
||||
This skill is delivered by `hooks/caveman-stats.js` (read by `hooks/caveman-mode-tracker.js` on `/caveman-stats`). The model does not need to do anything when this skill fires — the hook returns `decision: "block"` with the formatted stats as the reason. The user sees the numbers immediately.
|
||||
|
||||
Output also includes `Est. rule overhead` and `Est. net` lines wherever a savings estimate exists with a known turn count. Rule overhead is the estimated per-turn INPUT-token cost of the injected caveman rules (default 1,250 tokens/turn, override with `CAVEMAN_RULE_OVERHEAD_TOKENS`) times the turn count. Net is savings minus that overhead — when negative, the output says so plainly and suggests turning caveman off for that workload, rather than hiding the net-negative regime behind a gross-savings number (see `docs/HONEST-NUMBERS.md`).
|
||||
@@ -1,120 +1,88 @@
|
||||
---
|
||||
name: caveman
|
||||
description: >
|
||||
Ultra-compressed communication mode. Cuts token usage ~75% by speaking like caveman
|
||||
while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra.
|
||||
Ultra-compressed communication mode. Cuts output tokens 65% (measured) by speaking like caveman
|
||||
while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra,
|
||||
wenyan-lite, wenyan-full, wenyan-ultra.
|
||||
Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens",
|
||||
"be brief", or invokes /caveman. Also auto-triggers when token efficiency is requested.
|
||||
---
|
||||
|
||||
# Caveman Mode
|
||||
Respond terse like smart caveman. All technical substance stay. Only fluff die.
|
||||
|
||||
## Core Rule
|
||||
## Persistence
|
||||
|
||||
Respond like smart caveman. Cut articles, filler, pleasantries. Keep all technical substance.
|
||||
ACTIVE EVERY RESPONSE. No revert after many turns. No filler drift. Still active if unsure. Off only: "stop caveman" / "normal mode".
|
||||
|
||||
Default intensity: **full**. Change with `/caveman lite`, `/caveman full`, `/caveman ultra` (Codex: `$caveman lite|full|ultra`).
|
||||
Default: **full**. Switch: `/caveman lite|full|ultra|wenyan-lite|wenyan-full|wenyan-ultra|off`.
|
||||
|
||||
## Grammar
|
||||
## Rules
|
||||
|
||||
- Drop articles (a, an, the)
|
||||
- Drop filler (just, really, basically, actually, simply)
|
||||
- Drop pleasantries (sure, certainly, of course, happy to)
|
||||
- Short synonyms (big not extensive, fix not "implement a solution for")
|
||||
- No hedging (skip "it might be worth considering")
|
||||
- Fragments fine. No need full sentence
|
||||
- Technical terms stay exact. "Polymorphism" stays "polymorphism"
|
||||
- Code blocks unchanged. Caveman speak around code, not in code
|
||||
- Error messages quoted exact. Caveman only for explanation
|
||||
Drop: articles (a/an/the), filler (just/really/basically/actually/simply), pleasantries (sure/certainly/of course/happy to), hedging. Fragments OK. Short synonyms (big not extensive, fix not "implement a solution for"). No tool-call narration, no decorative tables/emoji, no dumping long raw error logs unless asked — quote shortest decisive line. Standard well-known tech acronyms OK (DB/API/HTTP); never invent new abbreviations (cfg/impl/req/res/fn) — tokenizer split them same as full word: zero token saved, reader still decode. Full word cheaper AND clearer. No causal arrows (→) either — own token, save nothing. Technical terms exact. Code blocks unchanged. Errors quoted exact.
|
||||
|
||||
## Pattern
|
||||
Never drop not/never/no/only/except — flip meaning worse than any token saved. Numbers, units exact.
|
||||
|
||||
```
|
||||
[thing] [action] [reason]. [next step].
|
||||
```
|
||||
Tool calls: fire direct. No preamble, plan, or progress note before or between calls. After result: next call direct or final answer — never announce next call. Text before call only to clarify, warn security/irreversible, or resolve ambiguity.
|
||||
|
||||
Not:
|
||||
> Sure! I'd be happy to help you with that. The issue you're experiencing is likely caused by...
|
||||
Preserve user's dominant language exactly — reply in the language user writes, never switch regardless of example text or multilingual context elsewhere. Compress the style, not the language. Every emitted line in that language — openings, pre-tool status lines, all — not just final reply. ALWAYS keep technical terms, code, API names, CLI commands, commit-type keywords (feat/fix/...), and exact error strings verbatim — unless user explicitly ask for translation.
|
||||
|
||||
Yes:
|
||||
> Bug in auth middleware. Token expiry check use `<` not `<=`. Fix:
|
||||
'Drop articles' = article languages only. Where small markers carry case/role (particles, postpositions), keep them — grammar, not filler; compress politeness/filler instead.
|
||||
|
||||
## Examples
|
||||
No self-reference. Never name or announce the style. No "caveman mode on", "me caveman think", no third-person caveman tags. Output caveman-only — never normal answer plus "Caveman:" recap. Exception: user explicitly ask what the mode is.
|
||||
|
||||
**User:** Why is my React component re-rendering?
|
||||
Pattern: `[thing] [action] [reason]. [next step].`
|
||||
|
||||
**Normal (69 tokens):** "The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object."
|
||||
Not: "Sure! I'd be happy to help you with that. The issue you're experiencing is likely caused by..."
|
||||
Yes: "Bug in auth middleware. Token expiry check use `<` not `<=`. Fix:"
|
||||
|
||||
**Caveman (19 tokens):** "New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`."
|
||||
## Intensity
|
||||
|
||||
---
|
||||
| Level | What change |
|
||||
|-------|------------|
|
||||
| **lite** | No filler/hedging. Keep articles + full sentences. Professional but tight |
|
||||
| **full** | Drop articles, fragments OK, short synonyms. Classic caveman. No tool-call narration, no decorative tables/emoji, no long raw error-log dumps unless asked. Standard acronyms OK; no invented abbreviations |
|
||||
| **ultra** | Strip conjunctions when cause-then-effect stay unambiguous. One word when one word enough. State each fact once. NO prose abbreviations (cfg/impl/req/res/fn/auth), NO arrows (X → Y) — measured zero token saving under tokenizer, cost decode clarity. Code symbols, function names, API names, error strings: never touch |
|
||||
| **wenyan-lite** | Semi-classical. Drop filler/hedging but keep grammar structure, classical register |
|
||||
| **wenyan-full** | Maximum classical terseness. Fully 文言文. 80-90% character reduction — chars, not tokens. Classical sentence patterns, verbs precede objects, subjects often omitted, classical particles (之/乃/為/其) |
|
||||
| **wenyan-ultra** | Extreme abbreviation while keeping classical Chinese feel. Maximum compression, ultra terse |
|
||||
|
||||
**User:** How do I set up a PostgreSQL connection pool?
|
||||
Example — "Why React component re-render?"
|
||||
- lite: "Your component re-renders because you create a new object reference each render. Wrap it in `useMemo`."
|
||||
- full: "New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`."
|
||||
- ultra: "Inline obj prop, new ref, re-render. `useMemo`."
|
||||
- wenyan-lite: "組件頻重繪,以每繪新生對象參照故。以 useMemo 包之。"
|
||||
- wenyan-full: "每繪新生對象參照,故重繪;以 useMemo 包之則免。"
|
||||
- wenyan-ultra: "新參照則重繪。useMemo 包之。"
|
||||
|
||||
**Caveman:**
|
||||
```
|
||||
Use `pg` pool:
|
||||
```
|
||||
```js
|
||||
const pool = new Pool({
|
||||
max: 20,
|
||||
idleTimeoutMillis: 30000,
|
||||
connectionTimeoutMillis: 2000,
|
||||
})
|
||||
```
|
||||
```
|
||||
max = concurrent connections. Keep under DB limit. idleTimeout kill stale conn.
|
||||
```
|
||||
Example — "Explain database connection pooling."
|
||||
- lite: "Connection pooling reuses open connections instead of creating new ones per request. Avoids repeated handshake overhead."
|
||||
- full: "Pool reuse open DB connections. No new connection per request. Skip handshake overhead."
|
||||
- ultra: "Pool reuse open DB connections. No per-request handshake."
|
||||
- wenyan-full: "池蓄已開之連,不逐請而新開,省握手之費。"
|
||||
- wenyan-ultra: "池蓄連,免逐請新開,省握手。"
|
||||
|
||||
## Intensity Levels
|
||||
Classical chars = wenyan modes only. Never swap a word to a classical char to shrink at non-wenyan levels.
|
||||
|
||||
### Lite — trim the fat
|
||||
## Auto-Clarity
|
||||
|
||||
Professional tone, just no fluff. Grammar stays intact.
|
||||
Drop caveman when:
|
||||
- Security warnings
|
||||
- Irreversible action confirmations
|
||||
- Multi-step sequences where fragment order or omitted conjunctions risk misread
|
||||
- Compression itself creates technical ambiguity (e.g., `"migrate table drop column backup first"` — order unclear without articles/conjunctions)
|
||||
- User asks to clarify or repeats question
|
||||
|
||||
- Drop filler and pleasantries (same list as full)
|
||||
- Drop hedging
|
||||
- Keep articles, keep full sentences
|
||||
- Prefer short synonyms where natural
|
||||
Resume caveman after clear part done.
|
||||
|
||||
### Full (default)
|
||||
Example shows FORMAT only — write warning in session language, not example's.
|
||||
|
||||
Classic caveman. Rules from Grammar section above apply.
|
||||
|
||||
### Ultra — maximum grunt
|
||||
|
||||
Telegraphic. Every word earn its place or die.
|
||||
|
||||
- All full rules, plus:
|
||||
- Abbreviate common terms (DB, auth, config, req, res, fn, impl)
|
||||
- Strip conjunctions where possible
|
||||
- One word answer when one word enough
|
||||
- Arrow notation for causality (X → Y)
|
||||
|
||||
## Intensity Examples
|
||||
|
||||
**User:** Why is my React component re-rendering?
|
||||
|
||||
**Lite:** "Your component re-renders because you create a new object reference each render. Inline object props fail shallow comparison every time. Wrap it in `useMemo`."
|
||||
|
||||
**Full:** "New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`."
|
||||
|
||||
**Ultra:** "Inline obj prop → new ref → re-render. `useMemo`."
|
||||
|
||||
---
|
||||
|
||||
**User:** Explain database connection pooling.
|
||||
|
||||
**Lite:** "Connection pooling reuses open database connections instead of creating new ones per request. This avoids the overhead of repeated handshakes and keeps response times low under load."
|
||||
|
||||
**Full:** "Pool reuse open DB connections. No new connection per request. Skip repeated handshake overhead. Response time stay low under load."
|
||||
|
||||
**Ultra:** "Pool = reuse DB conn. Skip handshake overhead → fast under load."
|
||||
Example — destructive op:
|
||||
> **Warning:** This will permanently delete all rows in the `users` table and cannot be undone.
|
||||
> ```sql
|
||||
> DROP TABLE users;
|
||||
> ```
|
||||
> Caveman resume. Verify backup exist first.
|
||||
|
||||
## Boundaries
|
||||
|
||||
- Code: write normal. Caveman English only
|
||||
- Git commits: normal
|
||||
- PR descriptions: normal
|
||||
- User say "stop caveman" or "normal mode": revert immediately
|
||||
- Intensity level persist until changed or session end
|
||||
Persisted outside chat: write normal prose — code, comments, commits, docs, issue/PR/MR text, memory files, third-party messages (/caveman-compress exempt). "stop caveman" or "normal mode": revert. Level persist until changed or session end.
|
||||
@@ -0,0 +1,6 @@
|
||||
interface:
|
||||
display_name: "Caveman"
|
||||
short_description: "Talk like caveman. Cut filler. Keep technical accuracy."
|
||||
icon_small: "./assets/caveman-small.svg"
|
||||
icon_large: "./assets/caveman.svg"
|
||||
default_prompt: "Use $caveman to answer briefly, cut filler, and preserve exact technical substance."
|
||||
@@ -0,0 +1,7 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 64 64" role="img" aria-label="Caveman">
|
||||
<circle cx="32" cy="32" r="28" fill="#6B7280"/>
|
||||
<path d="M18 40c4-10 24-10 28 0" fill="none" stroke="#F9FAFB" stroke-linecap="round" stroke-width="4"/>
|
||||
<circle cx="24" cy="27" r="4" fill="#F9FAFB"/>
|
||||
<circle cx="40" cy="27" r="4" fill="#F9FAFB"/>
|
||||
<path d="M21 18c3-5 8-8 15-8 5 0 10 2 14 7" fill="none" stroke="#D1D5DB" stroke-linecap="round" stroke-width="4"/>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 471 B |
@@ -0,0 +1,7 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 128 128" role="img" aria-label="Caveman">
|
||||
<rect width="128" height="128" rx="24" fill="#6B7280"/>
|
||||
<path d="M28 79c8-22 48-22 56 0" fill="none" stroke="#F9FAFB" stroke-linecap="round" stroke-width="8"/>
|
||||
<circle cx="48" cy="52" r="8" fill="#F9FAFB"/>
|
||||
<circle cx="80" cy="52" r="8" fill="#F9FAFB"/>
|
||||
<path d="M43 35c6-10 17-15 31-15 12 0 24 5 31 14" fill="none" stroke="#D1D5DB" stroke-linecap="round" stroke-width="8"/>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 487 B |
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"version": 1,
|
||||
"skills": {
|
||||
"cavecrew": {
|
||||
"source": "JuliusBrussee/caveman",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/cavecrew/SKILL.md",
|
||||
"computedHash": "06d45a7308d8603365313decf400020106b985cb5ce500ee169bfba9c71dd147"
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,61 @@
|
||||
# cavecrew
|
||||
|
||||
Decision guide. When to delegate to caveman subagents instead of doing the work inline.
|
||||
|
||||
## What it does
|
||||
|
||||
Tells the main thread when to spawn a caveman-style subagent versus the vanilla equivalent. The win: subagent tool-results inject back into main context verbatim, and caveman output is roughly 1/3 the size of vanilla prose. Across 20 delegations in one session, that is the difference between context exhaustion and finishing the task.
|
||||
|
||||
Three subagents:
|
||||
|
||||
| Subagent | Job | Use when |
|
||||
|----------|-----|----------|
|
||||
| `cavecrew-investigator` | Locate code (read-only) | "Where is X defined / what calls Y / list uses of Z" |
|
||||
| `cavecrew-builder` | Surgical edit, 1-2 files | Scope is obvious, ≤2 files. Refuses 3+ file scope. |
|
||||
| `cavecrew-reviewer` | Diff/file review | One-line findings with severity emoji |
|
||||
|
||||
Use vanilla `Explore` or `Code Reviewer` when you want prose, architecture commentary, or rationale. Use main thread directly for one-line answers and 3+ file refactors.
|
||||
|
||||
This skill is a decision guide, not a slash command. It activates when the conversation mentions delegation.
|
||||
|
||||
## How to invoke
|
||||
|
||||
Triggers on phrases like "delegate to subagent", "use cavecrew", "spawn investigator", "save context", "compressed agent output".
|
||||
|
||||
## Example chaining
|
||||
|
||||
Locate → fix → verify (most common):
|
||||
|
||||
1. `cavecrew-investigator` returns site list (`path:line — symbol — note`)
|
||||
2. Main thread picks 1-2 sites, hands paths to `cavecrew-builder`
|
||||
3. `cavecrew-reviewer` audits the resulting diff
|
||||
|
||||
Parallel scout: spawn 2-3 `cavecrew-investigator` calls in one message with different angles (defs, callers, tests). Aggregate in main.
|
||||
|
||||
## Model overrides
|
||||
|
||||
By default, `cavecrew-reviewer` and `cavecrew-investigator` pin `model: haiku` in their frontmatter; `cavecrew-builder` has no `model:` line (uses the API session default). Set env vars in your shell before launching Claude Code to override per-agent:
|
||||
|
||||
| Env var | Agent |
|
||||
|---|---|
|
||||
| `CAVECREW_REVIEWER_MODEL` | `cavecrew-reviewer` |
|
||||
| `CAVECREW_BUILDER_MODEL` | `cavecrew-builder` |
|
||||
| `CAVECREW_INVESTIGATOR_MODEL` | `cavecrew-investigator` |
|
||||
|
||||
Example — run reviewer on sonnet, keep others on default:
|
||||
|
||||
```sh
|
||||
export CAVECREW_REVIEWER_MODEL=sonnet
|
||||
```
|
||||
|
||||
Use the same model name strings you'd use in any Claude Code agent frontmatter (e.g. `haiku`, `sonnet`, `opus`).
|
||||
|
||||
Overrides patch only the `model:` line in the installed agent's frontmatter; the prompt body is untouched and keeps receiving upstream updates. Plugin installs only — standalone hook installs have no local agent files to patch. Unset or blank = no change. The patch persists in the installed file until the plugin is updated or reinstalled.
|
||||
|
||||
## See also
|
||||
|
||||
- [`SKILL.md`](./SKILL.md) — full decision matrix and output contracts
|
||||
- [`agents/cavecrew-investigator.md`](../../agents/cavecrew-investigator.md)
|
||||
- [`agents/cavecrew-builder.md`](../../agents/cavecrew-builder.md)
|
||||
- [`agents/cavecrew-reviewer.md`](../../agents/cavecrew-reviewer.md)
|
||||
- [Caveman README](../../README.md) — repo overview
|
||||
@@ -0,0 +1,82 @@
|
||||
---
|
||||
name: cavecrew
|
||||
description: >
|
||||
Decision guide for delegating to caveman-style subagents. Tells the main
|
||||
thread WHEN to spawn `cavecrew-investigator` (locate code), `cavecrew-builder`
|
||||
(1-2 file edit), or `cavecrew-reviewer` (diff review) instead of doing the
|
||||
work inline or using vanilla `Explore`. Subagent output is caveman-compressed
|
||||
so the tool-result injected back into main context is ~60% smaller — main
|
||||
context lasts longer across long sessions.
|
||||
Trigger: "delegate to subagent", "use cavecrew", "spawn investigator/builder/reviewer",
|
||||
"save context", "compressed agent output".
|
||||
---
|
||||
|
||||
Cavecrew = three subagent presets that emit caveman output. Same job as Anthropic defaults (`Explore`, edit-style agents, reviewer); difference is the tool-result they return is compressed, so main context shrinks per delegation.
|
||||
|
||||
## When to use cavecrew vs alternatives
|
||||
|
||||
| Task | Use |
|
||||
|---|---|
|
||||
| "Where is X defined / what calls Y / list uses of Z" | `cavecrew-investigator` |
|
||||
| Same but you also want suggestions/architecture commentary | `Explore` (vanilla) |
|
||||
| Surgical edit, ≤2 files, scope obvious | `cavecrew-builder` |
|
||||
| New feature / 3+ files / cross-cutting refactor | Main thread or `feature-dev:code-architect` |
|
||||
| Review diff, branch, or file for bugs | `cavecrew-reviewer` |
|
||||
| Deep code review with rationale + alternatives | `Code Reviewer` (vanilla) |
|
||||
| One-line answer you already know | Main thread, no subagent |
|
||||
|
||||
Rule of thumb: **if you'd want the subagent's output in 1/3 the tokens, pick cavecrew. If you'd want prose, pick vanilla.**
|
||||
|
||||
## Why this exists (the real win)
|
||||
|
||||
Subagent tool results get injected into main context verbatim. A vanilla `Explore` that returns 2k tokens of prose costs 2k tokens of main-context budget every time. The same finding from `cavecrew-investigator` returns ~700 tokens. Across 20 delegations in one session that's the difference between context exhaustion and finishing the task.
|
||||
|
||||
## Output contracts
|
||||
|
||||
What main thread can rely on per agent:
|
||||
|
||||
**`cavecrew-investigator`**
|
||||
```
|
||||
<Header>:
|
||||
- path:line — `symbol` — short note
|
||||
totals: <counts>.
|
||||
```
|
||||
Or `No match.` Always file-path-first, line-number-attached, backticked symbols. Safe to grep with `path:\d+`.
|
||||
|
||||
**`cavecrew-builder`**
|
||||
```
|
||||
<path:line-range> — <change ≤10 words>.
|
||||
verified: <re-read OK | mismatch @ path:line>.
|
||||
```
|
||||
Or one of: `too-big.` / `needs-confirm.` / `ambiguous.` / `regressed.` (terminal first token).
|
||||
|
||||
**`cavecrew-reviewer`**
|
||||
```
|
||||
path:line: <emoji> <severity>: <problem>. <fix>.
|
||||
totals: N🔴 N🟡 N🔵 N❓
|
||||
```
|
||||
Or `No issues.` Findings sorted file → line ascending.
|
||||
|
||||
## Chaining patterns
|
||||
|
||||
**Locate → fix → verify** (most common):
|
||||
1. `cavecrew-investigator` returns site list.
|
||||
2. Main thread picks 1-2 sites, hands paths to `cavecrew-builder`.
|
||||
3. `cavecrew-reviewer` audits the diff.
|
||||
|
||||
**Parallel scout** (when investigation is broad):
|
||||
Spawn 2-3 `cavecrew-investigator` calls in one message (different angles: defs vs callers vs tests). Aggregate in main thread.
|
||||
|
||||
**Single-shot edit** (when site is already known):
|
||||
Skip investigator. Hand exact path:line to `cavecrew-builder` directly.
|
||||
|
||||
## What NOT to do
|
||||
|
||||
- Don't use `cavecrew-builder` when you don't already know the file. Spawn investigator first or main thread will eat tokens passing context.
|
||||
- Don't chain `cavecrew-investigator → cavecrew-builder` for a 5-file refactor. Builder will return `too-big.` and you'll have wasted a turn.
|
||||
- Don't ask `cavecrew-reviewer` for "general feedback" — it returns findings only, no architecture opinions. Use `Code Reviewer` for that.
|
||||
- Don't expect prose. Cavecrew output is structured, sometimes terse to the point of cryptic. If a human will read it directly, paraphrase.
|
||||
|
||||
## Auto-clarity (inherited)
|
||||
|
||||
Subagents drop caveman → normal English for security warnings, irreversible-action confirmations, and any output where fragment ambiguity could be misread. Resume caveman after.
|
||||
@@ -0,0 +1,44 @@
|
||||
# caveman-commit
|
||||
|
||||
Terse Conventional Commits. Why over what.
|
||||
|
||||
## What it does
|
||||
|
||||
Generates commit messages in Conventional Commits format. Subject ≤50 chars, hard cap 72. Imperative mood. Body only when the *why* is non-obvious or there are breaking changes. No AI attribution, no "this commit does X", no emoji unless the project uses them. Body always required for breaking changes, security fixes, data migrations, and reverts — future debuggers need the context.
|
||||
|
||||
Outputs only the message. Does not stage, commit, or amend.
|
||||
|
||||
## How to invoke
|
||||
|
||||
```
|
||||
/caveman-commit
|
||||
```
|
||||
|
||||
Also triggers on phrases like "write a commit", "commit message", "generate commit".
|
||||
|
||||
## Example output
|
||||
|
||||
Diff: new endpoint for user profile.
|
||||
|
||||
```
|
||||
feat(api): add GET /users/:id/profile
|
||||
|
||||
Mobile client needs profile data without the full user payload
|
||||
to reduce LTE bandwidth on cold-launch screens.
|
||||
|
||||
Closes #128
|
||||
```
|
||||
|
||||
Diff: breaking API rename.
|
||||
|
||||
```
|
||||
feat(api)!: rename /v1/orders to /v1/checkout
|
||||
|
||||
BREAKING CHANGE: clients on /v1/orders must migrate to /v1/checkout
|
||||
before 2026-06-01. Old route returns 410 after that date.
|
||||
```
|
||||
|
||||
## See also
|
||||
|
||||
- [`SKILL.md`](./SKILL.md) — full LLM-facing instructions
|
||||
- [Caveman README](../../README.md) — repo overview
|
||||
@@ -0,0 +1,65 @@
|
||||
---
|
||||
name: caveman-commit
|
||||
description: >
|
||||
Ultra-compressed commit message generator. Cuts noise from commit messages while preserving
|
||||
intent and reasoning. Conventional Commits format. Subject ≤50 chars, body only when "why"
|
||||
isn't obvious. Use when user says "write a commit", "commit message", "generate commit",
|
||||
"/commit", or invokes /caveman-commit. Auto-triggers when staging changes.
|
||||
---
|
||||
|
||||
Write commit messages terse and exact. Conventional Commits format. No fluff. Why over what.
|
||||
|
||||
## Rules
|
||||
|
||||
**Subject line:**
|
||||
- `<type>(<scope>): <imperative summary>` — `<scope>` optional
|
||||
- Types: `feat`, `fix`, `refactor`, `perf`, `docs`, `test`, `chore`, `build`, `ci`, `style`, `revert`
|
||||
- Imperative mood: "add", "fix", "remove" — not "added", "adds", "adding"
|
||||
- ≤50 chars when possible, hard cap 72
|
||||
- No trailing period
|
||||
- Match project convention for capitalization after the colon
|
||||
|
||||
**Body (only if needed):**
|
||||
- Skip entirely when subject is self-explanatory
|
||||
- Add body only for: non-obvious *why*, breaking changes, migration notes, linked issues
|
||||
- Wrap at 72 chars
|
||||
- Bullets `-` not `*`
|
||||
- Reference issues/PRs at end: `Closes #42`, `Refs #17`
|
||||
|
||||
**What NEVER goes in:**
|
||||
- "This commit does X", "I", "we", "now", "currently" — the diff says what
|
||||
- "As requested by..." — use Co-authored-by trailer
|
||||
- "Generated with Claude Code" or any AI attribution — unless the user's own rule requires an `Assisted-by`/AI-attribution trailer, then add it as a trailer
|
||||
- Emoji (unless project convention requires)
|
||||
- Restating the file name when scope already says it
|
||||
|
||||
## Examples
|
||||
|
||||
Diff: new endpoint for user profile with body explaining the why
|
||||
- ❌ "feat: add a new endpoint to get user profile information from the database"
|
||||
- ✅
|
||||
```
|
||||
feat(api): add GET /users/:id/profile
|
||||
|
||||
Mobile client needs profile data without the full user payload
|
||||
to reduce LTE bandwidth on cold-launch screens.
|
||||
|
||||
Closes #128
|
||||
```
|
||||
|
||||
Diff: breaking API change
|
||||
- ✅
|
||||
```
|
||||
feat(api)!: rename /v1/orders to /v1/checkout
|
||||
|
||||
BREAKING CHANGE: clients on /v1/orders must migrate to /v1/checkout
|
||||
before 2026-06-01. Old route returns 410 after that date.
|
||||
```
|
||||
|
||||
## Auto-Clarity
|
||||
|
||||
Always include body for: breaking changes, security fixes, data migrations, anything reverting a prior commit. Never compress these into subject-only — future debuggers need the context.
|
||||
|
||||
## Boundaries
|
||||
|
||||
Only generates the commit message. Does not run `git commit`, does not stage files, does not amend. Output the message as a code block ready to paste. "stop caveman-commit" or "normal mode": revert to verbose commit style.
|
||||
@@ -25,7 +25,7 @@ CLAUDE.md ← compressed (Claude reads this — fewer tokens every sess
|
||||
CLAUDE.original.md ← human-readable backup (you edit this)
|
||||
```
|
||||
|
||||
Original never lost. You can read and edit `.original.md`. Run skill again to re-compress after edits.
|
||||
Original never lost. Backup lives in a data dir, not next to your file — `$XDG_DATA_HOME/caveman-compress/backups/<parent-dir-name>/` (macOS/Linux) or `%LOCALAPPDATA%\caveman-compress\backups\<parent-dir-name>\` (Windows) — so skill auto-loaders don't re-read it as a live file. You can read and edit `.original.md` there. Run skill again to re-compress after edits.
|
||||
|
||||
## Benchmarks
|
||||
|
||||
@@ -35,10 +35,10 @@ Real results on real project files:
|
||||
|------|----------:|----------:|------:|
|
||||
| `claude-md-preferences.md` | 706 | 285 | **59.6%** |
|
||||
| `project-notes.md` | 1145 | 535 | **53.3%** |
|
||||
| `claude-md-project.md` | 1122 | 687 | **38.8%** |
|
||||
| `claude-md-project.md` | 1122 | 636 | **43.3%** |
|
||||
| `todo-list.md` | 627 | 388 | **38.1%** |
|
||||
| `mixed-with-code.md` | 888 | 574 | **35.4%** |
|
||||
| **Average** | **898** | **494** | **45%** |
|
||||
| `mixed-with-code.md` | 888 | 560 | **36.9%** |
|
||||
| **Average** | **898** | **481** | **46%** |
|
||||
|
||||
All validations passed ✅ — headings, code blocks, URLs, file paths preserved exactly.
|
||||
|
||||
@@ -55,7 +55,7 @@ All validations passed ✅ — headings, code blocks, URLs, file paths preserved
|
||||
</td>
|
||||
<td width="50%">
|
||||
|
||||
### 🪨 Caveman (285 tokens)
|
||||
### <img src="../../docs/assets/dancing-rock.svg" width="20" height="20" alt="rock"/> Caveman (285 tokens)
|
||||
|
||||
> "Prefer TypeScript strict mode always. No `any` unless unavoidable — comment why if used. Proper types catch bugs early."
|
||||
|
||||
@@ -65,16 +65,18 @@ All validations passed ✅ — headings, code blocks, URLs, file paths preserved
|
||||
|
||||
**Same instructions. 60% fewer tokens. Every. Single. Session.**
|
||||
|
||||
## Security
|
||||
|
||||
`caveman-compress` is flagged as Snyk High Risk due to subprocess and file I/O patterns detected by static analysis. This is a false positive — see [SECURITY.md](./SECURITY.md) for a full explanation of what the skill does and does not do.
|
||||
|
||||
## Install
|
||||
|
||||
```bash
|
||||
cp -r ~/.claude/skills/caveman-compress <path-to-skill>
|
||||
```
|
||||
Compress is built in with the `caveman` plugin. Install `caveman` once, then use `/caveman-compress`.
|
||||
|
||||
Or if you have the caveman repo:
|
||||
If you need local files, the compress skill lives at:
|
||||
|
||||
```bash
|
||||
cp -r skills/caveman-compress ~/.claude/skills/caveman-compress
|
||||
caveman-compress/
|
||||
```
|
||||
|
||||
**Requires:** Python 3.10+
|
||||
@@ -96,7 +98,7 @@ Examples:
|
||||
|
||||
| Type | Compress? |
|
||||
|------|-----------|
|
||||
| `.md`, `.txt`, `.rst` | ✅ Yes |
|
||||
| `.md`, `.txt`, `.rst`, `.typ`, `.typst`, `.tex` | ✅ Yes |
|
||||
| Extensionless natural language | ✅ Yes |
|
||||
| `.py`, `.js`, `.ts`, `.json`, `.yaml` | ❌ Skip (code/config) |
|
||||
| `*.original.md` | ❌ Skip (backup files) |
|
||||
@@ -142,15 +144,15 @@ Caveman compress natural language. It never touch:
|
||||
|
||||
`CLAUDE.md` loads on **every session start**. A 1000-token project memory file costs tokens every single time you open a project. Over 100 sessions that's 100,000 tokens of overhead — just for context you already wrote.
|
||||
|
||||
Caveman cut that by ~45% on average. Same instructions. Same accuracy. Less waste.
|
||||
Caveman cut that by ~46% on average. Same instructions. Same accuracy. Less waste.
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────┐
|
||||
│ TOKEN SAVINGS PER FILE ████████ 45% │
|
||||
│ SESSIONS THAT BENEFIT ████████ 100% │
|
||||
│ INFORMATION PRESERVED ████████ 100% │
|
||||
│ SETUP TIME █ 1x │
|
||||
└──────────────────────────────────────────┘
|
||||
┌────────────────────────────────────────────┐
|
||||
│ TOKEN SAVINGS PER FILE █████ 46% │
|
||||
│ SESSIONS THAT BENEFIT ██████████ 100% │
|
||||
│ INFORMATION PRESERVED ██████████ 100% │
|
||||
│ SETUP TIME █ 1x │
|
||||
└────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
## Part of Caveman
|
||||
@@ -158,4 +160,4 @@ Caveman cut that by ~45% on average. Same instructions. Same accuracy. Less wast
|
||||
This skill is part of the [caveman](https://github.com/JuliusBrussee/caveman) toolkit — making Claude use fewer tokens without losing accuracy.
|
||||
|
||||
- **caveman** — make Claude *speak* like caveman (cuts response tokens ~65%)
|
||||
- **caveman-compress** — make Claude *read* less (cuts context tokens ~45%)
|
||||
- **caveman-compress** — make Claude *read* less (cuts context tokens ~46%)
|
||||
@@ -0,0 +1,31 @@
|
||||
# Security
|
||||
|
||||
## Snyk High Risk Rating
|
||||
|
||||
`caveman-compress` receives a Snyk High Risk rating due to static analysis heuristics. This document explains what the skill does and does not do.
|
||||
|
||||
### What triggers the rating
|
||||
|
||||
1. **subprocess usage**: The skill calls the `claude` CLI via `subprocess.run()` as a fallback when `ANTHROPIC_API_KEY` is not set. The subprocess call uses a fixed argument list — no shell interpolation occurs. User file content is passed via stdin, not as a shell argument.
|
||||
|
||||
2. **File read/write**: The skill reads the file the user explicitly points it at, compresses it, and writes the result back to the same path. A `.original.md` backup is saved to an out-of-tree data dir (`$XDG_DATA_HOME/caveman-compress/backups/<parent-dir-name>/`, or `%LOCALAPPDATA%\caveman-compress\backups\<parent-dir-name>\` on Windows). Beyond the target file and that backup location, no files are read or written.
|
||||
|
||||
### What the skill does NOT do
|
||||
|
||||
- Does not execute user file content as code
|
||||
- Does not make network requests except to Anthropic's API (via SDK or CLI)
|
||||
- Does not access files outside the path the user provides
|
||||
- Does not use shell=True or string interpolation in subprocess calls
|
||||
- Does not collect or transmit any data beyond the file being compressed
|
||||
|
||||
### Auth behavior
|
||||
|
||||
If `ANTHROPIC_API_KEY` is set, the skill uses the Anthropic Python SDK directly (no subprocess). If not set, it falls back to the `claude` CLI, which uses the user's existing Claude desktop authentication.
|
||||
|
||||
### File size limit
|
||||
|
||||
Files larger than 500KB are rejected before any API call is made.
|
||||
|
||||
### Reporting a vulnerability
|
||||
|
||||
If you believe you've found a genuine security issue, please open a GitHub issue with the label `security`.
|
||||
@@ -0,0 +1,111 @@
|
||||
---
|
||||
name: caveman-compress
|
||||
description: >
|
||||
Compress natural language memory files (CLAUDE.md, todos, preferences) into caveman format
|
||||
to save input tokens. Preserves all technical substance, code, URLs, and structure.
|
||||
Compressed version overwrites the original file. Human-readable backup saved as FILE.original.md.
|
||||
Trigger: /caveman-compress FILEPATH or "compress memory file"
|
||||
---
|
||||
|
||||
# Caveman Compress
|
||||
|
||||
## Purpose
|
||||
|
||||
Compress natural language files (CLAUDE.md, todos, preferences) into caveman-speak to reduce input tokens. Compressed version overwrites original. Human-readable backup saved as `<filename>.original.md`, but NOT beside the source file — it lives in an out-of-tree data dir (`$XDG_DATA_HOME/caveman-compress/backups/<parent-dir-name>/`, or `%LOCALAPPDATA%\caveman-compress\backups\<parent-dir-name>\` on Windows) so skill auto-loaders don't re-ingest it as a live file.
|
||||
|
||||
## Trigger
|
||||
|
||||
`/caveman-compress <filepath>` or when user asks to compress a memory file.
|
||||
|
||||
## Process
|
||||
|
||||
1. The compression scripts live in `scripts/` (adjacent to this SKILL.md). If the path is not immediately available, search for `scripts/__main__.py` next to this SKILL.md.
|
||||
|
||||
2. From the directory containing this SKILL.md, run:
|
||||
|
||||
python3 -m scripts <absolute_filepath>
|
||||
|
||||
3. The CLI will:
|
||||
- detect file type (no tokens)
|
||||
- call Claude to compress
|
||||
- validate output (no tokens)
|
||||
- if errors: cherry-pick fix with Claude (targeted fixes only, no recompression)
|
||||
- retry up to 2 times
|
||||
- if still failing after 2 retries: report error to user, leave original file untouched
|
||||
|
||||
4. Return result to user
|
||||
|
||||
## Compression Rules
|
||||
|
||||
### Remove
|
||||
- Articles: a, an, the
|
||||
- Filler: just, really, basically, actually, simply, essentially, generally
|
||||
- Pleasantries: "sure", "certainly", "of course", "happy to", "I'd recommend"
|
||||
- Hedging: "it might be worth", "you could consider", "it would be good to"
|
||||
- Redundant phrasing: "in order to" → "to", "make sure to" → "ensure", "the reason is because" → "because"
|
||||
- Connective fluff: "however", "furthermore", "additionally", "in addition"
|
||||
|
||||
### Preserve EXACTLY (never modify)
|
||||
- Code blocks (fenced ``` and indented)
|
||||
- Inline code (`backtick content`)
|
||||
- URLs and links (full URLs, markdown links)
|
||||
- File paths (`/src/components/...`, `./config.yaml`)
|
||||
- Commands (`npm install`, `git commit`, `docker build`)
|
||||
- Technical terms (library names, API names, protocols, algorithms)
|
||||
- Proper nouns (project names, people, companies)
|
||||
- Dates, version numbers, numeric values
|
||||
- Environment variables (`$HOME`, `NODE_ENV`)
|
||||
|
||||
### Preserve Structure
|
||||
- All markdown headings (keep exact heading text, compress body below)
|
||||
- Bullet point hierarchy (keep nesting level)
|
||||
- Numbered lists (keep numbering)
|
||||
- Tables (compress cell text, keep structure)
|
||||
- Frontmatter/YAML headers in markdown files
|
||||
|
||||
### Compress
|
||||
- Use short synonyms: "big" not "extensive", "fix" not "implement a solution for", "use" not "utilize"
|
||||
- Fragments OK: "Run tests before commit" not "You should always run tests before committing"
|
||||
- Drop "you should", "make sure to", "remember to" — just state the action
|
||||
- Merge redundant bullets that say the same thing differently
|
||||
- Keep one example where multiple examples show the same pattern
|
||||
|
||||
CRITICAL RULE:
|
||||
Anything inside ``` ... ``` must be copied EXACTLY.
|
||||
Do not:
|
||||
- remove comments
|
||||
- remove spacing
|
||||
- reorder lines
|
||||
- shorten commands
|
||||
- simplify anything
|
||||
|
||||
Inline code (`...`) must be preserved EXACTLY.
|
||||
Do not modify anything inside backticks.
|
||||
|
||||
If file contains code blocks:
|
||||
- Treat code blocks as read-only regions
|
||||
- Only compress text outside them
|
||||
- Do not merge sections around code
|
||||
|
||||
## Pattern
|
||||
|
||||
Original:
|
||||
> You should always make sure to run the test suite before pushing any changes to the main branch. This is important because it helps catch bugs early and prevents broken builds from being deployed to production.
|
||||
|
||||
Compressed:
|
||||
> Run tests before push to main. Catch bugs early, prevent broken prod deploys.
|
||||
|
||||
Original:
|
||||
> The application uses a microservices architecture with the following components. The API gateway handles all incoming requests and routes them to the appropriate service. The authentication service is responsible for managing user sessions and JWT tokens.
|
||||
|
||||
Compressed:
|
||||
> Microservices architecture. API gateway route all requests to services. Auth service manage user sessions + JWT tokens.
|
||||
|
||||
## Boundaries
|
||||
|
||||
- ONLY compress natural language files (.md, .txt, .typ, .typst, .tex, extensionless)
|
||||
- NEVER modify: .py, .js, .ts, .json, .yaml, .yml, .toml, .env, .lock, .css, .html, .xml, .sql, .sh
|
||||
- If file has mixed content (prose + code), compress ONLY the prose sections
|
||||
- If unsure whether something is code or prose, leave it unchanged
|
||||
- Original file is backed up as FILE.original.md before overwriting — in the out-of-tree backup data dir (see Purpose), not beside the source file
|
||||
- Never compress FILE.original.md (skip it)
|
||||
@@ -0,0 +1,9 @@
|
||||
"""Caveman compress scripts.
|
||||
|
||||
This package provides tools to compress natural language markdown files
|
||||
into caveman format to save input tokens.
|
||||
"""
|
||||
|
||||
__all__ = ["cli", "compress", "detect", "validate"]
|
||||
|
||||
__version__ = "1.0.0"
|
||||
@@ -0,0 +1,3 @@
|
||||
from .cli import main
|
||||
|
||||
main()
|
||||
@@ -0,0 +1,80 @@
|
||||
#!/usr/bin/env python3
|
||||
from pathlib import Path
|
||||
import sys
|
||||
|
||||
# Support both direct execution and module import
|
||||
try:
|
||||
from .validate import validate
|
||||
except ImportError:
|
||||
sys.path.insert(0, str(Path(__file__).parent))
|
||||
from validate import validate
|
||||
|
||||
try:
|
||||
import tiktoken
|
||||
_enc = tiktoken.get_encoding("o200k_base")
|
||||
except ImportError:
|
||||
_enc = None
|
||||
|
||||
|
||||
def count_tokens(text):
|
||||
if _enc is None:
|
||||
return len(text.split()) # fallback: word count
|
||||
return len(_enc.encode(text))
|
||||
|
||||
|
||||
def benchmark_pair(orig_path: Path, comp_path: Path):
|
||||
orig_text = orig_path.read_text(encoding="utf-8", errors="ignore")
|
||||
comp_text = comp_path.read_text(encoding="utf-8", errors="ignore")
|
||||
|
||||
orig_tokens = count_tokens(orig_text)
|
||||
comp_tokens = count_tokens(comp_text)
|
||||
saved = 100 * (orig_tokens - comp_tokens) / orig_tokens if orig_tokens > 0 else 0.0
|
||||
result = validate(orig_path, comp_path)
|
||||
|
||||
return (comp_path.name, orig_tokens, comp_tokens, saved, result.is_valid)
|
||||
|
||||
|
||||
def print_table(rows):
|
||||
print("\n| File | Original | Compressed | Saved % | Valid |")
|
||||
print("|------|----------|------------|---------|-------|")
|
||||
for r in rows:
|
||||
print(f"| {r[0]} | {r[1]} | {r[2]} | {r[3]:.1f}% | {'✅' if r[4] else '❌'} |")
|
||||
|
||||
|
||||
def main():
|
||||
# Direct file pair: python3 benchmark.py original.md compressed.md
|
||||
if len(sys.argv) == 3:
|
||||
orig = Path(sys.argv[1]).resolve()
|
||||
comp = Path(sys.argv[2]).resolve()
|
||||
if not orig.exists():
|
||||
print(f"❌ Not found: {orig}")
|
||||
sys.exit(1)
|
||||
if not comp.exists():
|
||||
print(f"❌ Not found: {comp}")
|
||||
sys.exit(1)
|
||||
print_table([benchmark_pair(orig, comp)])
|
||||
return
|
||||
|
||||
# Glob mode: repo_root/tests/caveman-compress/
|
||||
# __file__ lives at <repo_root>/skills/caveman-compress/scripts/benchmark.py
|
||||
# Walk up four dirs: scripts → caveman-compress → skills → repo_root.
|
||||
tests_dir = Path(__file__).resolve().parents[3] / "tests" / "caveman-compress"
|
||||
if not tests_dir.exists():
|
||||
print(f"❌ Tests dir not found: {tests_dir}")
|
||||
sys.exit(1)
|
||||
|
||||
rows = []
|
||||
for orig in sorted(tests_dir.glob("*.original.md")):
|
||||
comp = orig.with_name(orig.stem.removesuffix(".original") + ".md")
|
||||
if comp.exists():
|
||||
rows.append(benchmark_pair(orig, comp))
|
||||
|
||||
if not rows:
|
||||
print("No compressed file pairs found.")
|
||||
return
|
||||
|
||||
print_table(rows)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,85 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Caveman Compress CLI
|
||||
|
||||
Usage:
|
||||
caveman <filepath>
|
||||
"""
|
||||
|
||||
import sys
|
||||
|
||||
# Force UTF-8 on stdout/stderr before any code can print. Windows consoles
|
||||
# default to cp1252 and crash on the ❌ glyphs in error/validation branches,
|
||||
# masking the real error and leaving the user with a half-compressed file.
|
||||
for _stream in (sys.stdout, sys.stderr):
|
||||
reconfigure = getattr(_stream, "reconfigure", None)
|
||||
if callable(reconfigure):
|
||||
try:
|
||||
reconfigure(encoding="utf-8", errors="replace")
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
from .compress import backup_dir_for, compress_file
|
||||
from .detect import detect_file_type, should_compress
|
||||
|
||||
|
||||
def print_usage():
|
||||
print("Usage: caveman <filepath>")
|
||||
|
||||
|
||||
def main():
|
||||
if len(sys.argv) != 2:
|
||||
print_usage()
|
||||
sys.exit(1)
|
||||
|
||||
filepath = Path(sys.argv[1])
|
||||
|
||||
# Check file exists
|
||||
if not filepath.exists():
|
||||
print(f"❌ File not found: {filepath}")
|
||||
sys.exit(1)
|
||||
|
||||
if not filepath.is_file():
|
||||
print(f"❌ Not a file: {filepath}")
|
||||
sys.exit(1)
|
||||
|
||||
filepath = filepath.resolve()
|
||||
|
||||
# Detect file type
|
||||
file_type = detect_file_type(filepath)
|
||||
|
||||
print(f"Detected: {file_type}")
|
||||
|
||||
# Check if compressible
|
||||
if not should_compress(filepath):
|
||||
print("Skipping: file is not natural language (code/config)")
|
||||
sys.exit(0)
|
||||
|
||||
print("Starting caveman compression...\n")
|
||||
|
||||
try:
|
||||
success = compress_file(filepath)
|
||||
|
||||
if success:
|
||||
print("\nCompression completed successfully")
|
||||
backup_path = backup_dir_for(filepath) / (filepath.stem + ".original.md")
|
||||
print(f"Compressed: {filepath}")
|
||||
print(f"Original: {backup_path}")
|
||||
sys.exit(0)
|
||||
else:
|
||||
print("\n❌ Compression failed after retries")
|
||||
sys.exit(2)
|
||||
|
||||
except KeyboardInterrupt:
|
||||
print("\nInterrupted by user")
|
||||
sys.exit(130)
|
||||
|
||||
except Exception as e:
|
||||
print(f"\n❌ Error: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,414 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Caveman Memory Compression Orchestrator
|
||||
|
||||
Usage:
|
||||
python scripts/compress.py <filepath>
|
||||
"""
|
||||
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import stat
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
from typing import List
|
||||
|
||||
OUTER_FENCE_REGEX = re.compile(
|
||||
r"\A\s*(`{3,}|~{3,})[^\n]*\n(.*)\n\1\s*\Z", re.DOTALL
|
||||
)
|
||||
|
||||
# YAML frontmatter: starts at file start with --- on its own line, ends with --- on its own line.
|
||||
# Captures the entire block (including delimiters and trailing newline) and the body after.
|
||||
FRONTMATTER_REGEX = re.compile(
|
||||
r"\A(---\r?\n.*?\r?\n---\r?\n)(.*)", re.DOTALL
|
||||
)
|
||||
|
||||
|
||||
def split_frontmatter(text: str):
|
||||
"""Split YAML frontmatter from body. Returns (frontmatter, body).
|
||||
|
||||
Memory files (and many other markdown docs) start with a YAML frontmatter
|
||||
block delimited by `---` lines. The compression LLM has a habit of stripping
|
||||
or rewriting these despite preserve-structure rules in the prompt — so we
|
||||
surgically remove the frontmatter before compression and prepend it back
|
||||
verbatim to the output. Files without frontmatter pass through unchanged.
|
||||
"""
|
||||
m = FRONTMATTER_REGEX.match(text)
|
||||
if m:
|
||||
return m.group(1), m.group(2)
|
||||
return "", text
|
||||
|
||||
# Filenames and paths that almost certainly hold secrets or PII. Compressing
|
||||
# them ships raw bytes to the Anthropic API — a third-party data boundary that
|
||||
# developers on sensitive codebases cannot cross. detect.py already skips .env
|
||||
# by extension, but credentials.md / secrets.txt / ~/.aws/credentials would
|
||||
# slip through the natural-language filter. This is a hard refuse before read.
|
||||
SENSITIVE_BASENAME_REGEX = re.compile(
|
||||
r"(?ix)^("
|
||||
r"\.env(\..+)?"
|
||||
r"|\.netrc"
|
||||
r"|credentials(\..+)?"
|
||||
r"|secrets?(\..+)?"
|
||||
r"|passwords?(\..+)?"
|
||||
r"|id_(rsa|dsa|ecdsa|ed25519)(\.pub)?"
|
||||
r"|authorized_keys"
|
||||
r"|known_hosts"
|
||||
r"|.*\.(pem|key|p12|pfx|crt|cer|jks|keystore|asc|gpg)"
|
||||
r")$"
|
||||
)
|
||||
|
||||
SENSITIVE_PATH_COMPONENTS = frozenset({".ssh", ".aws", ".gnupg", ".kube", ".docker"})
|
||||
|
||||
SENSITIVE_NAME_TOKENS = (
|
||||
"secret", "credential", "password", "passwd",
|
||||
"apikey", "accesskey", "token", "privatekey",
|
||||
)
|
||||
|
||||
|
||||
def backup_dir_for(filepath: Path) -> Path:
|
||||
"""Resolve the out-of-tree backup directory for a given source file.
|
||||
|
||||
Backups must live OUTSIDE the source directory so skill auto-loaders
|
||||
(Claude Code rules/, opencode instructions/, etc.) stop re-ingesting the
|
||||
`.original.md` copies as live files. Base dir is platform-aware:
|
||||
- Windows: %LOCALAPPDATA%\\caveman-compress\\backups
|
||||
- else: $XDG_DATA_HOME/caveman-compress/backups if set,
|
||||
else ~/.local/share/caveman-compress/backups
|
||||
|
||||
The source file's parent-dir name is mirrored under the base to reduce
|
||||
cross-project collisions (e.g. two `task.md` files in different repos).
|
||||
"""
|
||||
if os.name == "nt" or sys.platform == "win32":
|
||||
local_appdata = os.environ.get("LOCALAPPDATA")
|
||||
base = Path(local_appdata) if local_appdata else Path.home() / "AppData" / "Local"
|
||||
base = base / "caveman-compress" / "backups"
|
||||
else:
|
||||
xdg = os.environ.get("XDG_DATA_HOME")
|
||||
base = Path(xdg) if xdg else Path.home() / ".local" / "share"
|
||||
base = base / "caveman-compress" / "backups"
|
||||
return base / filepath.parent.name
|
||||
|
||||
|
||||
def is_sensitive_path(filepath: Path) -> bool:
|
||||
"""Heuristic denylist for files that must never be shipped to a third-party API."""
|
||||
name = filepath.name
|
||||
if SENSITIVE_BASENAME_REGEX.match(name):
|
||||
return True
|
||||
lowered_parts = {p.lower() for p in filepath.parts}
|
||||
if lowered_parts & SENSITIVE_PATH_COMPONENTS:
|
||||
return True
|
||||
# Normalize separators so "api-key" and "api_key" both match "apikey".
|
||||
lower = re.sub(r"[_\-\s.]", "", name.lower())
|
||||
return any(tok in lower for tok in SENSITIVE_NAME_TOKENS)
|
||||
|
||||
|
||||
def strip_llm_wrapper(text: str) -> str:
|
||||
"""Strip outer ```markdown ... ``` fence when it wraps the entire output."""
|
||||
m = OUTER_FENCE_REGEX.match(text)
|
||||
if m:
|
||||
return m.group(2)
|
||||
return text
|
||||
|
||||
|
||||
def write_text_atomic(path: Path, text: str) -> None:
|
||||
"""Write ``text`` to ``path`` atomically as UTF-8.
|
||||
|
||||
Path.write_text() truncates the destination before encoding the string —
|
||||
a UnicodeEncodeError (or any other failure) partway through leaves a
|
||||
0-byte file, destroying whatever was there before (issue #655). Encode
|
||||
first, write the bytes to a sibling temp file, fsync, then os.replace()
|
||||
so the destination only ever moves from one complete, valid file to
|
||||
another. Preserves the original file's permission bits across the swap.
|
||||
"""
|
||||
data = text.encode("utf-8")
|
||||
fd, tmp_name = tempfile.mkstemp(
|
||||
dir=str(path.parent), prefix=path.name + ".", suffix=".tmp"
|
||||
)
|
||||
tmp_path = Path(tmp_name)
|
||||
try:
|
||||
with os.fdopen(fd, "wb") as f:
|
||||
f.write(data)
|
||||
f.flush()
|
||||
os.fsync(f.fileno())
|
||||
if path.exists():
|
||||
os.chmod(tmp_path, stat.S_IMODE(path.stat().st_mode))
|
||||
os.replace(tmp_path, path)
|
||||
except Exception:
|
||||
try:
|
||||
tmp_path.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
raise
|
||||
|
||||
|
||||
def first_nonblank_line(text: str) -> str:
|
||||
"""Return the first non-blank line, stripped — used to detect a prose
|
||||
preamble smuggled in ahead of the real content (issue #588)."""
|
||||
for line in text.splitlines():
|
||||
if line.strip():
|
||||
return line.strip()
|
||||
return ""
|
||||
|
||||
|
||||
def _write_target(filepath: Path, text: str, backup_path: Path) -> None:
|
||||
"""Write to the target file, surfacing the backup location if the write
|
||||
itself fails. write_text_atomic already leaves the target untouched on
|
||||
failure, but the caller still needs to know where the pre-compression
|
||||
original lives instead of being left to guess (issue #652)."""
|
||||
try:
|
||||
write_text_atomic(filepath, text)
|
||||
except Exception:
|
||||
print(f"❌ Write to {filepath} failed. Original preserved at backup: {backup_path}")
|
||||
raise
|
||||
|
||||
|
||||
from .detect import should_compress
|
||||
from .validate import validate
|
||||
|
||||
MAX_RETRIES = 2
|
||||
|
||||
|
||||
# ---------- Claude Calls ----------
|
||||
|
||||
|
||||
def call_claude(prompt: str) -> str:
|
||||
"""Send a prompt to Claude.
|
||||
|
||||
Prefers the Anthropic SDK when ANTHROPIC_API_KEY is set; otherwise falls
|
||||
back to the ``claude --print`` CLI (which handles desktop auth).
|
||||
|
||||
On Windows the CLI subprocess decoding defaults to the system codepage
|
||||
(cp1251 / cp1252) and crashes on UTF-8 output — see issue #152. Pinning
|
||||
``encoding="utf-8"`` with ``errors="replace"`` matches the CLI's actual
|
||||
native I/O and prevents the UnicodeDecodeError before validation can
|
||||
report. Windows users with non-ASCII content can also set
|
||||
``ANTHROPIC_API_KEY`` to route through the SDK and skip the subprocess.
|
||||
"""
|
||||
api_key = os.environ.get("ANTHROPIC_API_KEY")
|
||||
if api_key:
|
||||
try:
|
||||
import anthropic
|
||||
|
||||
client = anthropic.Anthropic(api_key=api_key)
|
||||
msg = client.messages.create(
|
||||
model=os.environ.get("CAVEMAN_MODEL", "claude-sonnet-4-5"),
|
||||
max_tokens=8192,
|
||||
messages=[{"role": "user", "content": prompt}],
|
||||
)
|
||||
return strip_llm_wrapper(msg.content[0].text.strip())
|
||||
except ImportError:
|
||||
pass # anthropic not installed, fall back to CLI
|
||||
# Fallback: use claude CLI (handles desktop auth).
|
||||
# Resolve binary via shutil.which so Windows .cmd/.bat shims (e.g.
|
||||
# %APPDATA%\npm\claude.CMD) work without shell=True. On POSIX,
|
||||
# shutil.which returns the same absolute path as the implicit lookup,
|
||||
# so this is a no-op there. Falls back to bare "claude" if not found
|
||||
# on PATH so subprocess raises a clear FileNotFoundError.
|
||||
claude_bin = shutil.which("claude") or "claude"
|
||||
try:
|
||||
result = subprocess.run(
|
||||
[claude_bin, "--print"],
|
||||
input=prompt,
|
||||
text=True,
|
||||
capture_output=True,
|
||||
check=True,
|
||||
encoding="utf-8",
|
||||
errors="replace",
|
||||
)
|
||||
return strip_llm_wrapper(result.stdout.strip())
|
||||
except subprocess.CalledProcessError as e:
|
||||
raise RuntimeError(f"Claude call failed:\n{e.stderr}")
|
||||
|
||||
|
||||
def build_compress_prompt(original: str) -> str:
|
||||
return f"""
|
||||
Compress this markdown into caveman format.
|
||||
|
||||
STRICT RULES:
|
||||
- Do NOT modify anything inside ``` code blocks
|
||||
- Do NOT modify anything inside inline backticks
|
||||
- Preserve ALL URLs exactly
|
||||
- Preserve ALL headings exactly
|
||||
- Preserve file paths and commands
|
||||
- Return ONLY the compressed markdown body — do NOT wrap the entire output in a ```markdown fence or any other fence. Inner code blocks from the original stay as-is; do not add a new outer fence around the whole file.
|
||||
|
||||
Only compress natural language.
|
||||
|
||||
TEXT:
|
||||
{original}
|
||||
"""
|
||||
|
||||
|
||||
def build_fix_prompt(original: str, compressed: str, errors: List[str]) -> str:
|
||||
errors_str = "\n".join(f"- {e}" for e in errors)
|
||||
return f"""You are fixing a caveman-compressed markdown file. Specific validation errors were found.
|
||||
|
||||
CRITICAL RULES:
|
||||
- DO NOT recompress or rephrase the file
|
||||
- ONLY fix the listed errors — leave everything else exactly as-is
|
||||
- The ORIGINAL is provided as reference only (to restore missing content)
|
||||
- Preserve caveman style in all untouched sections
|
||||
|
||||
ERRORS TO FIX:
|
||||
{errors_str}
|
||||
|
||||
HOW TO FIX:
|
||||
- Missing URL: find it in ORIGINAL, restore it exactly where it belongs in COMPRESSED
|
||||
- Code block mismatch: find the exact code block in ORIGINAL, restore it in COMPRESSED
|
||||
- Heading mismatch: restore the exact heading text from ORIGINAL into COMPRESSED
|
||||
- Do not touch any section not mentioned in the errors
|
||||
|
||||
ORIGINAL (reference only):
|
||||
{original}
|
||||
|
||||
COMPRESSED (fix this):
|
||||
{compressed}
|
||||
|
||||
Return ONLY the fixed compressed file. No explanation.
|
||||
"""
|
||||
|
||||
|
||||
# ---------- Core Logic ----------
|
||||
|
||||
|
||||
def compress_file(filepath: Path) -> bool:
|
||||
# Resolve and validate path
|
||||
filepath = filepath.resolve()
|
||||
MAX_FILE_SIZE = 500_000 # 500KB
|
||||
if not filepath.exists():
|
||||
raise FileNotFoundError(f"File not found: {filepath}")
|
||||
if filepath.stat().st_size > MAX_FILE_SIZE:
|
||||
raise ValueError(f"File too large to compress safely (max 500KB): {filepath}")
|
||||
|
||||
# Refuse files that look like they contain secrets or PII. Compressing ships
|
||||
# the raw bytes to the Anthropic API — a third-party boundary — so we fail
|
||||
# loudly rather than silently exfiltrate credentials or keys. Override is
|
||||
# intentional: the user must rename the file if the heuristic is wrong.
|
||||
if is_sensitive_path(filepath):
|
||||
raise ValueError(
|
||||
f"Refusing to compress {filepath}: filename looks sensitive "
|
||||
"(credentials, keys, secrets, or known private paths). "
|
||||
"Compression sends file contents to the Anthropic API. "
|
||||
"Rename the file if this is a false positive."
|
||||
)
|
||||
|
||||
print(f"Processing: {filepath}")
|
||||
|
||||
if not should_compress(filepath):
|
||||
print("Skipping (not natural language)")
|
||||
return False
|
||||
|
||||
original_text = filepath.read_text(encoding="utf-8", errors="ignore")
|
||||
# Store backup outside the source directory so skill auto-loaders don't
|
||||
# re-ingest the `.original.md` copy as a live file. Mirror the source's
|
||||
# parent-dir name + stem under a platform-aware base to reduce collisions.
|
||||
backup_dir = backup_dir_for(filepath)
|
||||
backup_dir.mkdir(parents=True, exist_ok=True)
|
||||
backup_path = backup_dir / (filepath.stem + ".original.md")
|
||||
|
||||
if not original_text.strip():
|
||||
print("❌ Refusing to compress: file is empty or whitespace-only.")
|
||||
return False
|
||||
|
||||
# Check if backup already exists to prevent accidental overwriting
|
||||
if backup_path.exists():
|
||||
print(f"⚠️ Backup file already exists: {backup_path}")
|
||||
print("The original backup may contain important content.")
|
||||
print("Aborting to prevent data loss. Please remove or rename the backup file if you want to proceed.")
|
||||
return False
|
||||
|
||||
# Split YAML frontmatter off before compression. Claude tends to strip or
|
||||
# rewrite frontmatter despite preserve-structure rules; we keep it verbatim
|
||||
# by removing it from the input and re-prepending it to the output.
|
||||
frontmatter, body = split_frontmatter(original_text)
|
||||
if frontmatter:
|
||||
print(f"Detected YAML frontmatter ({len(frontmatter)} chars) — preserving verbatim")
|
||||
|
||||
if not body.strip():
|
||||
print("❌ Refusing to compress: body is empty after frontmatter removal.")
|
||||
return False
|
||||
|
||||
# Step 1: Compress (body only, frontmatter excluded)
|
||||
print("Compressing with Claude...")
|
||||
compressed_body = call_claude(build_compress_prompt(body))
|
||||
|
||||
if compressed_body is None or not compressed_body.strip():
|
||||
print("❌ Compression aborted: Claude returned an empty response.")
|
||||
print(" Original file is untouched (no backup created).")
|
||||
return False
|
||||
|
||||
# Compare the BODY (not the whole file) — frontmatter is preserved verbatim
|
||||
# and would never change, so identity must be judged on the compressible part.
|
||||
if compressed_body.strip() == body.strip():
|
||||
print("❌ Compression aborted: output is identical to input.")
|
||||
print(" Likely causes: Claude refused, returned the prompt verbatim, or the file is")
|
||||
print(" already in caveman form. Original file is untouched (no backup created).")
|
||||
return False
|
||||
|
||||
# Reassemble: frontmatter (verbatim) + compressed body
|
||||
compressed = frontmatter + compressed_body
|
||||
|
||||
# Save original as backup, then verify the backup readback before
|
||||
# touching the input file. If the filesystem dropped bytes (encoding,
|
||||
# antivirus, disk full), unlink the bad backup and abort instead of
|
||||
# leaving the user with a corrupt backup + compressed primary.
|
||||
write_text_atomic(backup_path, original_text)
|
||||
backup_readback = backup_path.read_text(encoding="utf-8", errors="ignore")
|
||||
if backup_readback != original_text:
|
||||
print(f"❌ Backup write verification failed: {backup_path}")
|
||||
print(" In-memory original differs from on-disk backup. Aborting before touching the input file.")
|
||||
try:
|
||||
backup_path.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
_write_target(filepath, compressed, backup_path)
|
||||
|
||||
# Step 2: Validate + Retry
|
||||
for attempt in range(MAX_RETRIES):
|
||||
print(f"\nValidation attempt {attempt + 1}")
|
||||
|
||||
result = validate(backup_path, filepath)
|
||||
|
||||
if result.is_valid:
|
||||
print("Validation passed")
|
||||
break
|
||||
|
||||
print("❌ Validation failed:")
|
||||
for err in result.errors:
|
||||
print(f" - {err}")
|
||||
|
||||
if attempt == MAX_RETRIES - 1:
|
||||
# Restore original on failure
|
||||
_write_target(filepath, original_text, backup_path)
|
||||
backup_path.unlink(missing_ok=True)
|
||||
print("❌ Failed after retries — original restored")
|
||||
return False
|
||||
|
||||
print("Fixing with Claude...")
|
||||
compressed = call_claude(
|
||||
build_fix_prompt(original_text, compressed, result.errors)
|
||||
)
|
||||
|
||||
if compressed is None or not compressed.strip():
|
||||
print("❌ Fix attempt aborted: Claude returned an empty response.")
|
||||
print(" Skipping this attempt.")
|
||||
continue
|
||||
|
||||
# Guard against a prose preamble smuggled in ahead of the real fixed
|
||||
# content (issue #588). Only enforced when the original starts with a
|
||||
# structural anchor (frontmatter `---` or a heading) — plain-prose
|
||||
# first lines get legitimately rewritten by compression, and requiring
|
||||
# them verbatim would reject every valid fix.
|
||||
anchor = first_nonblank_line(original_text)
|
||||
if anchor.startswith(("---", "#")) and first_nonblank_line(compressed) != anchor:
|
||||
print("❌ Fix attempt aborted: output does not start with the original's first line.")
|
||||
print(" Possible preamble leak. Skipping this attempt.")
|
||||
continue
|
||||
|
||||
_write_target(filepath, compressed, backup_path)
|
||||
|
||||
return True
|
||||
@@ -0,0 +1,139 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Detect whether a file is natural language (compressible) or code/config (skip)."""
|
||||
|
||||
import json
|
||||
import re
|
||||
from pathlib import Path
|
||||
|
||||
# Extensions that are natural language and compressible
|
||||
COMPRESSIBLE_EXTENSIONS = {".md", ".txt", ".markdown", ".rst", ".typ", ".typst", ".tex"}
|
||||
|
||||
# Extensions that are code/config and should be skipped
|
||||
SKIP_EXTENSIONS = {
|
||||
".py", ".js", ".ts", ".tsx", ".jsx", ".json", ".yaml", ".yml",
|
||||
".toml", ".env", ".lock", ".css", ".scss", ".html", ".xml",
|
||||
".sql", ".sh", ".bash", ".zsh", ".go", ".rs", ".java", ".c",
|
||||
".cpp", ".h", ".hpp", ".rb", ".php", ".swift", ".kt", ".lua",
|
||||
".dockerfile", ".makefile", ".csv", ".ini", ".cfg",
|
||||
}
|
||||
|
||||
# Well-known build/config files that carry no (or a misleading) extension —
|
||||
# `Dockerfile` has no suffix so `.dockerfile` above never matches it, and
|
||||
# `CMakeLists.txt` would ride the compressible `.txt` rule. Checked by
|
||||
# basename before any extension rule.
|
||||
KNOWN_CODE_FILENAMES = {
|
||||
"dockerfile", "makefile", "gnumakefile", "jenkinsfile", "vagrantfile",
|
||||
"rakefile", "gemfile", "justfile", "procfile", "brewfile",
|
||||
"cmakelists.txt",
|
||||
}
|
||||
|
||||
# Patterns that indicate a line is code
|
||||
CODE_PATTERNS = [
|
||||
re.compile(r"^\s*(import |from .+ import |require\(|const |let |var )"),
|
||||
re.compile(r"^\s*(def |class |function |async function |export )"),
|
||||
re.compile(r"^\s*(if\s*\(|for\s*\(|while\s*\(|switch\s*\(|try\s*\{)"),
|
||||
re.compile(r"^\s*[\}\]\);]+\s*$"), # closing braces/brackets
|
||||
re.compile(r"^\s*@\w+"), # decorators/annotations
|
||||
re.compile(r'^\s*"[^"]+"\s*:\s*'), # JSON-like key-value
|
||||
re.compile(r"^\s*\w+\s*=\s*[{\[\(\"']"), # assignment with literal
|
||||
]
|
||||
|
||||
|
||||
def _is_code_line(line: str) -> bool:
|
||||
"""Check if a line looks like code."""
|
||||
return any(p.match(line) for p in CODE_PATTERNS)
|
||||
|
||||
|
||||
def _is_json_content(text: str) -> bool:
|
||||
"""Check if content is valid JSON."""
|
||||
try:
|
||||
json.loads(text)
|
||||
return True
|
||||
except (json.JSONDecodeError, ValueError):
|
||||
return False
|
||||
|
||||
|
||||
def _is_yaml_content(lines: list[str]) -> bool:
|
||||
"""Heuristic: check if content looks like YAML."""
|
||||
yaml_indicators = 0
|
||||
for line in lines[:30]:
|
||||
stripped = line.strip()
|
||||
if stripped.startswith("---"):
|
||||
yaml_indicators += 1
|
||||
elif re.match(r"^\w[\w\s]*:\s", stripped):
|
||||
yaml_indicators += 1
|
||||
elif stripped.startswith("- ") and ":" in stripped:
|
||||
yaml_indicators += 1
|
||||
# If most non-empty lines look like YAML
|
||||
non_empty = sum(1 for l in lines[:30] if l.strip())
|
||||
return non_empty > 0 and yaml_indicators / non_empty > 0.6
|
||||
|
||||
|
||||
def detect_file_type(filepath: Path) -> str:
|
||||
"""Classify a file as 'natural_language', 'code', 'config', or 'unknown'.
|
||||
|
||||
Returns:
|
||||
One of: 'natural_language', 'code', 'config', 'unknown'
|
||||
"""
|
||||
ext = filepath.suffix.lower()
|
||||
|
||||
# Known code filenames win over any extension rule
|
||||
if filepath.name.lower() in KNOWN_CODE_FILENAMES:
|
||||
return "code"
|
||||
|
||||
# Extension-based classification
|
||||
if ext in COMPRESSIBLE_EXTENSIONS:
|
||||
return "natural_language"
|
||||
if ext in SKIP_EXTENSIONS:
|
||||
return "code" if ext not in {".json", ".yaml", ".yml", ".toml", ".ini", ".cfg", ".env"} else "config"
|
||||
|
||||
# Extensionless files (like CLAUDE.md, TODO) — check content
|
||||
if not ext:
|
||||
try:
|
||||
text = filepath.read_text(encoding="utf-8", errors="ignore")
|
||||
except (OSError, PermissionError):
|
||||
return "unknown"
|
||||
|
||||
lines = text.splitlines()[:50]
|
||||
|
||||
# Shebang means executable script, never prose
|
||||
if text.startswith("#!"):
|
||||
return "code"
|
||||
|
||||
if _is_json_content(text[:10000]):
|
||||
return "config"
|
||||
if _is_yaml_content(lines):
|
||||
return "config"
|
||||
|
||||
code_lines = sum(1 for l in lines if l.strip() and _is_code_line(l))
|
||||
non_empty = sum(1 for l in lines if l.strip())
|
||||
if non_empty > 0 and code_lines / non_empty > 0.4:
|
||||
return "code"
|
||||
|
||||
return "natural_language"
|
||||
|
||||
return "unknown"
|
||||
|
||||
|
||||
def should_compress(filepath: Path) -> bool:
|
||||
"""Return True if the file is natural language and should be compressed."""
|
||||
if not filepath.is_file():
|
||||
return False
|
||||
# Skip backup files
|
||||
if filepath.name.endswith(".original.md"):
|
||||
return False
|
||||
return detect_file_type(filepath) == "natural_language"
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
import sys
|
||||
|
||||
if len(sys.argv) < 2:
|
||||
print("Usage: python detect.py <file1> [file2] ...")
|
||||
sys.exit(1)
|
||||
|
||||
for path_str in sys.argv[1:]:
|
||||
p = Path(path_str).resolve()
|
||||
file_type = detect_file_type(p)
|
||||
compress = should_compress(p)
|
||||
print(f" {p.name:30s} type={file_type:20s} compress={compress}")
|
||||
@@ -0,0 +1,221 @@
|
||||
#!/usr/bin/env python3
|
||||
import re
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
|
||||
URL_REGEX = re.compile(r"https?://[^\s)]+")
|
||||
FENCE_OPEN_REGEX = re.compile(r"^(\s{0,3})(`{3,}|~{3,})(.*)$")
|
||||
HEADING_REGEX = re.compile(r"^(#{1,6})\s+(.*)", re.MULTILINE)
|
||||
BULLET_REGEX = re.compile(r"^\s*[-*+]\s+", re.MULTILINE)
|
||||
|
||||
# crude but effective path detection
|
||||
# Requires either a path prefix (./ ../ / or drive letter) or a slash/backslash within the match
|
||||
PATH_REGEX = re.compile(r"(?:\./|\.\./|/|[A-Za-z]:\\)[\w\-/\\\.]+|[\w\-\.]+[/\\][\w\-/\\\.]+")
|
||||
|
||||
|
||||
class ValidationResult:
|
||||
def __init__(self):
|
||||
self.is_valid = True
|
||||
self.errors = []
|
||||
self.warnings = []
|
||||
|
||||
def add_error(self, msg):
|
||||
self.is_valid = False
|
||||
self.errors.append(msg)
|
||||
|
||||
def add_warning(self, msg):
|
||||
self.warnings.append(msg)
|
||||
|
||||
|
||||
def read_file(path: Path) -> str:
|
||||
return path.read_text(encoding="utf-8")
|
||||
|
||||
|
||||
# ---------- Extractors ----------
|
||||
|
||||
|
||||
def extract_headings(text):
|
||||
return [(level, title.strip()) for level, title in HEADING_REGEX.findall(text)]
|
||||
|
||||
|
||||
def extract_code_blocks(text):
|
||||
"""Line-based fenced code block extractor.
|
||||
|
||||
Handles ``` and ~~~ fences with variable length (CommonMark: closing
|
||||
fence must use same char and be at least as long as opening). Supports
|
||||
nested fences (e.g. an outer 4-backtick block wrapping inner 3-backtick
|
||||
content).
|
||||
"""
|
||||
blocks = []
|
||||
lines = text.split("\n")
|
||||
i = 0
|
||||
n = len(lines)
|
||||
while i < n:
|
||||
m = FENCE_OPEN_REGEX.match(lines[i])
|
||||
if not m:
|
||||
i += 1
|
||||
continue
|
||||
fence_char = m.group(2)[0]
|
||||
fence_len = len(m.group(2))
|
||||
open_line = lines[i]
|
||||
block_lines = [open_line]
|
||||
i += 1
|
||||
closed = False
|
||||
while i < n:
|
||||
close_m = FENCE_OPEN_REGEX.match(lines[i])
|
||||
if (
|
||||
close_m
|
||||
and close_m.group(2)[0] == fence_char
|
||||
and len(close_m.group(2)) >= fence_len
|
||||
and close_m.group(3).strip() == ""
|
||||
):
|
||||
block_lines.append(lines[i])
|
||||
closed = True
|
||||
i += 1
|
||||
break
|
||||
block_lines.append(lines[i])
|
||||
i += 1
|
||||
if closed:
|
||||
blocks.append("\n".join(block_lines))
|
||||
# Unclosed fences are silently skipped — they indicate malformed markdown
|
||||
# and including them would cause false-positive validation failures.
|
||||
return blocks
|
||||
|
||||
|
||||
def extract_urls(text):
|
||||
return set(URL_REGEX.findall(text))
|
||||
|
||||
|
||||
def extract_paths(text):
|
||||
return set(PATH_REGEX.findall(text))
|
||||
|
||||
|
||||
def count_bullets(text):
|
||||
return len(BULLET_REGEX.findall(text))
|
||||
|
||||
|
||||
def extract_inline_codes(text):
|
||||
"""Backtick-delimited inline spans, with fenced code blocks stripped first.
|
||||
|
||||
Previously used a column-0-anchored regex to strip fences, which misses
|
||||
fences indented 1-3 spaces (valid CommonMark). Reuse extract_code_blocks
|
||||
(FENCE_OPEN_REGEX-based, indentation-aware) instead so an indented fence's
|
||||
body backticks don't leak into inline-code pairing.
|
||||
"""
|
||||
text_without_fences = text
|
||||
for block in extract_code_blocks(text):
|
||||
text_without_fences = text_without_fences.replace(block, "", 1)
|
||||
return re.findall(r"`([^`]+)`", text_without_fences)
|
||||
|
||||
|
||||
# ---------- Validators ----------
|
||||
|
||||
|
||||
def validate_headings(orig, comp, result):
|
||||
h1 = extract_headings(orig)
|
||||
h2 = extract_headings(comp)
|
||||
|
||||
if len(h1) != len(h2):
|
||||
result.add_error(f"Heading count mismatch: {len(h1)} vs {len(h2)}")
|
||||
|
||||
if h1 != h2:
|
||||
result.add_warning("Heading text/order changed")
|
||||
|
||||
|
||||
def validate_code_blocks(orig, comp, result):
|
||||
c1 = extract_code_blocks(orig)
|
||||
c2 = extract_code_blocks(comp)
|
||||
|
||||
if c1 != c2:
|
||||
result.add_error("Code blocks not preserved exactly")
|
||||
|
||||
|
||||
def validate_urls(orig, comp, result):
|
||||
u1 = extract_urls(orig)
|
||||
u2 = extract_urls(comp)
|
||||
|
||||
if u1 != u2:
|
||||
result.add_error(f"URL mismatch: lost={u1 - u2}, added={u2 - u1}")
|
||||
|
||||
|
||||
def validate_paths(orig, comp, result):
|
||||
p1 = extract_paths(orig)
|
||||
p2 = extract_paths(comp)
|
||||
|
||||
if p1 != p2:
|
||||
result.add_warning(f"Path mismatch: lost={p1 - p2}, added={p2 - p1}")
|
||||
|
||||
|
||||
def validate_bullets(orig, comp, result):
|
||||
b1 = count_bullets(orig)
|
||||
b2 = count_bullets(comp)
|
||||
|
||||
if b1 == 0:
|
||||
return
|
||||
|
||||
diff = abs(b1 - b2) / b1
|
||||
|
||||
if diff > 0.15:
|
||||
result.add_warning(f"Bullet count changed too much: {b1} -> {b2}")
|
||||
|
||||
|
||||
def validate_inline_codes(orig, comp, result):
|
||||
c1 = Counter(extract_inline_codes(orig))
|
||||
c2 = Counter(extract_inline_codes(comp))
|
||||
|
||||
if c1 != c2:
|
||||
lost = set(c1.keys()) - set(c2.keys())
|
||||
added = set(c2.keys()) - set(c1.keys())
|
||||
for code, count in c1.items():
|
||||
if code in c2 and c2[code] < count:
|
||||
lost.add(f"{code} (lost {count - c2[code]} of {count} occurrences)")
|
||||
if lost:
|
||||
result.add_error(f"Inline code lost: {lost}")
|
||||
if added:
|
||||
result.add_warning(f"Inline code added: {added}")
|
||||
|
||||
|
||||
# ---------- Main ----------
|
||||
|
||||
|
||||
def validate(original_path: Path, compressed_path: Path) -> ValidationResult:
|
||||
result = ValidationResult()
|
||||
|
||||
orig = read_file(original_path)
|
||||
comp = read_file(compressed_path)
|
||||
|
||||
validate_headings(orig, comp, result)
|
||||
validate_code_blocks(orig, comp, result)
|
||||
validate_urls(orig, comp, result)
|
||||
validate_paths(orig, comp, result)
|
||||
validate_bullets(orig, comp, result)
|
||||
validate_inline_codes(orig, comp, result)
|
||||
|
||||
return result
|
||||
|
||||
|
||||
# ---------- CLI ----------
|
||||
|
||||
if __name__ == "__main__":
|
||||
import sys
|
||||
|
||||
if len(sys.argv) != 3:
|
||||
print("Usage: python validate.py <original> <compressed>")
|
||||
sys.exit(1)
|
||||
|
||||
orig = Path(sys.argv[1]).resolve()
|
||||
comp = Path(sys.argv[2]).resolve()
|
||||
|
||||
res = validate(orig, comp)
|
||||
|
||||
print(f"\nValid: {res.is_valid}")
|
||||
|
||||
if res.errors:
|
||||
print("\nErrors:")
|
||||
for e in res.errors:
|
||||
print(f" - {e}")
|
||||
|
||||
if res.warnings:
|
||||
print("\nWarnings:")
|
||||
for w in res.warnings:
|
||||
print(f" - {w}")
|
||||
@@ -0,0 +1,38 @@
|
||||
# caveman-help
|
||||
|
||||
Quick-reference card. One shot, no mode change.
|
||||
|
||||
## What it does
|
||||
|
||||
Prints a cheat sheet of all caveman modes, sibling skills, deactivation triggers, and how to set the default mode via env var or config file. One-shot display — does not flip the active mode, write flag files, or persist anything. Use when you forget the slash commands.
|
||||
|
||||
## How to invoke
|
||||
|
||||
```
|
||||
/caveman-help
|
||||
```
|
||||
|
||||
Also triggers on "caveman help", "what caveman commands", "how do I use caveman".
|
||||
|
||||
## Example output
|
||||
|
||||
```
|
||||
Modes:
|
||||
/caveman full (default)
|
||||
/caveman lite lighter
|
||||
/caveman ultra extreme
|
||||
/caveman wenyan classical Chinese
|
||||
|
||||
Skills:
|
||||
/caveman-commit terse Conventional Commits
|
||||
/caveman-review one-line PR comments
|
||||
/caveman-stats session token savings
|
||||
|
||||
Deactivate:
|
||||
"stop caveman" or "normal mode"
|
||||
```
|
||||
|
||||
## See also
|
||||
|
||||
- [`SKILL.md`](./SKILL.md) — full reference card
|
||||
- [Caveman README](../../README.md) — repo overview
|
||||
@@ -0,0 +1,63 @@
|
||||
---
|
||||
name: caveman-help
|
||||
description: >
|
||||
Quick-reference card for all caveman modes, skills, and commands.
|
||||
One-shot display, not a persistent mode. Trigger: /caveman-help,
|
||||
"caveman help", "what caveman commands", "how do I use caveman".
|
||||
---
|
||||
|
||||
# Caveman Help
|
||||
|
||||
Display this reference card when invoked. One-shot — do NOT change mode, write flag files, or persist anything. Output in caveman style.
|
||||
|
||||
## Modes
|
||||
|
||||
| Mode | Trigger | What change |
|
||||
|------|---------|-------------|
|
||||
| **Lite** | `/caveman lite` | Drop filler. Keep sentence structure. |
|
||||
| **Full** | `/caveman` | Drop articles, filler, pleasantries, hedging. Fragments OK. Default. |
|
||||
| **Ultra** | `/caveman ultra` | Extreme compression. Bare fragments. Tables over prose. |
|
||||
| **Wenyan-Lite** | `/caveman wenyan-lite` | Classical Chinese style, light compression. |
|
||||
| **Wenyan-Full** | `/caveman wenyan` | Full 文言文. Maximum classical terseness. |
|
||||
| **Wenyan-Ultra** | `/caveman wenyan-ultra` | Extreme. Ancient scholar on a budget. |
|
||||
|
||||
Mode stick until changed or session end.
|
||||
|
||||
## Skills
|
||||
|
||||
| Skill | Trigger | What it do |
|
||||
|-------|---------|-----------|
|
||||
| **caveman-commit** | `/caveman-commit` | Terse commit messages. Conventional Commits. ≤50 char subject. |
|
||||
| **caveman-review** | `/caveman-review` | One-line PR comments: `L42: bug: user null. Add guard.` |
|
||||
| **caveman-compress** | `/caveman-compress <file>` | Compress .md files to caveman prose. Saves ~46% input tokens. |
|
||||
| **caveman-help** | `/caveman-help` | This card. |
|
||||
|
||||
## Deactivate
|
||||
|
||||
Say "stop caveman" or "normal mode". Resume anytime with `/caveman`.
|
||||
|
||||
## Language
|
||||
|
||||
Keep user's language by default. User write Portuguese → reply Portuguese caveman. Compress the style, not the language. Technical terms, code, commands, commit types, and exact error strings stay verbatim unless user ask for translation.
|
||||
|
||||
## Configure Default Mode
|
||||
|
||||
Default mode = `full`. Change it:
|
||||
|
||||
**Environment variable** (highest priority):
|
||||
```bash
|
||||
export CAVEMAN_DEFAULT_MODE=ultra
|
||||
```
|
||||
|
||||
**Config file** (`~/.config/caveman/config.json` macOS/Linux, `%APPDATA%\caveman\config.json` Windows):
|
||||
```json
|
||||
{ "defaultMode": "lite" }
|
||||
```
|
||||
|
||||
Set `"off"` to disable auto-activation on session start. User can still activate manually with `/caveman`.
|
||||
|
||||
Resolution: env var > config file > `full`.
|
||||
|
||||
## More
|
||||
|
||||
Full docs: https://github.com/JuliusBrussee/caveman
|
||||
@@ -0,0 +1,33 @@
|
||||
# caveman-review
|
||||
|
||||
One-line PR comments. Location, problem, fix. No throat-clearing.
|
||||
|
||||
## What it does
|
||||
|
||||
Generates code review comments in `L<line>: <severity> <problem>. <fix>.` format. One line per finding. Severity emoji: 🔴 bug, 🟡 risk, 🔵 nit, ❓ question. Drops "I noticed that...", hedging, and restating what the diff already shows. Keeps exact line numbers, backticked symbols, and concrete fixes.
|
||||
|
||||
Auto-clarity: drops terse mode for CVE-class security findings, architectural disagreements, and onboarding contexts where the author needs the *why*. Resumes terse for the rest.
|
||||
|
||||
Output only — does not approve, request changes, or run linters.
|
||||
|
||||
## How to invoke
|
||||
|
||||
```
|
||||
/caveman-review
|
||||
```
|
||||
|
||||
Also triggers on "review this PR", "code review", "review the diff".
|
||||
|
||||
## Example output
|
||||
|
||||
```
|
||||
L42: 🔴 bug: user can be null after .find(). Add guard before .email.
|
||||
L88-140: 🔵 nit: 50-line fn does 4 things. Extract validate/normalize/persist.
|
||||
L23: 🟡 risk: no retry on 429. Wrap in withBackoff(3).
|
||||
L107: ❓ q: why drop the cache here? Reads on next request will miss.
|
||||
```
|
||||
|
||||
## See also
|
||||
|
||||
- [`SKILL.md`](./SKILL.md) — full LLM-facing instructions
|
||||
- [Caveman README](../../README.md) — repo overview
|
||||
@@ -0,0 +1,55 @@
|
||||
---
|
||||
name: caveman-review
|
||||
description: >
|
||||
Ultra-compressed code review comments. Cuts noise from PR feedback while preserving
|
||||
the actionable signal. Each comment is one line: location, problem, fix. Use when user
|
||||
says "review this PR", "code review", "review the diff", "/review", or invokes
|
||||
/caveman-review. Auto-triggers when reviewing pull requests.
|
||||
---
|
||||
|
||||
Write code review comments terse and actionable. One line per finding. Location, problem, fix. No throat-clearing.
|
||||
|
||||
## Rules
|
||||
|
||||
**Format:** `L<line>: <problem>. <fix>.` — or `<file>:L<line>: ...` when reviewing multi-file diffs.
|
||||
|
||||
**Severity prefix (optional, when mixed):**
|
||||
- `🔴 bug:` — broken behavior, will cause incident
|
||||
- `🟡 risk:` — works but fragile (race, missing null check, swallowed error)
|
||||
- `🔵 nit:` — style, naming, micro-optim. Author can ignore
|
||||
- `❓ q:` — genuine question, not a suggestion
|
||||
|
||||
**Drop:**
|
||||
- "I noticed that...", "It seems like...", "You might want to consider..."
|
||||
- "This is just a suggestion but..." — use `nit:` instead
|
||||
- "Great work!", "Looks good overall but..." — say it once at the top, not per comment
|
||||
- Restating what the line does — the reviewer can read the diff
|
||||
- Hedging ("perhaps", "maybe", "I think") — if unsure use `q:`
|
||||
|
||||
**Keep:**
|
||||
- Exact line numbers
|
||||
- Exact symbol/function/variable names in backticks
|
||||
- Concrete fix, not "consider refactoring this"
|
||||
- The *why* if the fix isn't obvious from the problem statement
|
||||
|
||||
## Examples
|
||||
|
||||
❌ "I noticed that on line 42 you're not checking if the user object is null before accessing the email property. This could potentially cause a crash if the user is not found in the database. You might want to add a null check here."
|
||||
|
||||
✅ `L42: 🔴 bug: user can be null after .find(). Add guard before .email.`
|
||||
|
||||
❌ "It looks like this function is doing a lot of things and might benefit from being broken up into smaller functions for readability."
|
||||
|
||||
✅ `L88-140: 🔵 nit: 50-line fn does 4 things. Extract validate/normalize/persist.`
|
||||
|
||||
❌ "Have you considered what happens if the API returns a 429? I think we should probably handle that case."
|
||||
|
||||
✅ `L23: 🟡 risk: no retry on 429. Wrap in withBackoff(3).`
|
||||
|
||||
## Auto-Clarity
|
||||
|
||||
Drop terse mode for: security findings (CVE-class bugs need full explanation + reference), architectural disagreements (need rationale, not just a one-liner), and onboarding contexts where the author is new and needs the "why". In those cases write a normal paragraph, then resume terse for the rest.
|
||||
|
||||
## Boundaries
|
||||
|
||||
Reviews only — does not write the code fix, does not approve/request-changes, does not run linters. Output the comment(s) ready to paste into the PR. "stop caveman-review" or "normal mode": revert to verbose review style.
|
||||
@@ -0,0 +1,36 @@
|
||||
# caveman-stats
|
||||
|
||||
Real session token receipts. No AI estimation.
|
||||
|
||||
## What it does
|
||||
|
||||
Reads the current Claude Code session log directly and reports actual input/output token usage plus estimated savings versus a non-caveman baseline. Numbers come from the JSONL session log on disk — the model itself does not compute or estimate them. Output is injected by the `caveman-mode-tracker` hook, which intercepts `/caveman-stats` and returns the formatted stats as a blocked-decision reason.
|
||||
|
||||
Output also includes an `Est. rule overhead` and `Est. net` line whenever the savings figure above them is unambiguous (a single benchmarked mode with a known turn count — no guessing across mixed or unattributed spans). Overhead estimates the per-turn INPUT-token cost of the rules the skill injects every turn — default 1,250 tokens/turn, override with `CAVEMAN_RULE_OVERHEAD_TOKENS` if you've measured your own setup. Net is savings minus that overhead. On short, terse replies this can go negative — caveman's OUTPUT savings don't clear its INPUT cost — and the line says so directly instead of hiding it behind a gross-savings number. Background: `docs/HONEST-NUMBERS.md`.
|
||||
|
||||
Each run also writes a lifetime-savings suffix file used by the statusline badge (`⛏ 12.4k`). That badge stays a gross-savings figure on purpose — it is a glanceable summary, not a full accounting; run `/caveman-stats` for the net picture.
|
||||
|
||||
## How to invoke
|
||||
|
||||
```
|
||||
/caveman-stats
|
||||
```
|
||||
|
||||
## Example output
|
||||
|
||||
```
|
||||
Session: 47 turns
|
||||
Input: 12,304 tokens
|
||||
Output: 3,891 tokens (caveman)
|
||||
Baseline: 11,247 tokens (estimated without caveman)
|
||||
Saved: 7,356 tokens (~65%)
|
||||
Est. rule overhead: 58,750 (input, ~1,250/turn over 47 turns)
|
||||
Est. net: -51,394 (caveman cost more than it saved for this workload — consider turning it off)
|
||||
```
|
||||
|
||||
(Numbers above are illustrative — see `docs/HONEST-NUMBERS.md` for why short, terse-reply sessions tend to land net-negative even at a healthy output-savings percentage.)
|
||||
|
||||
## See also
|
||||
|
||||
- [`SKILL.md`](./SKILL.md) — hook contract and mechanics
|
||||
- [Caveman README](../../README.md) — repo overview
|
||||