mirror of
https://github.com/runbear-io/beardrive.git
synced 2026-08-25 08:08:08 +02:00
docs.beardrive.ai shipped no robots.txt, so nothing on that host named the sitemap — a crawler arriving at the subdomain had to guess the URL or be handed it in Search Console. And every entry was a bare <loc>: no freshness signal at all, on the site whose whole value is being current. Each URL now carries the commit date of the markdown behind it, read from one `git log` for the whole tree. The build refuses to guess: in a shallow clone (the default for CI checkouts) git can only attribute every file to the single commit it has, so lastmod is omitted entirely rather than claiming the site changed wholesale on every deploy — Google discounts a sitemap that does that, which would cost more than the absent dates. Hosts that want the dates need full history; README says so. Declaring @astrojs/sitemap explicitly replaces the copy Starlight adds for itself rather than duplicating it — that's the supported way to reach these options, and Starlight's own version only configures i18n, which this single-language site doesn't use. Adds `npm run check:sitemap <origin>`: robots.txt -> index -> every advertised URL returns 200, the checks Search Console runs, against a local preview or against production. Run against production today it fails on the missing Sitemap: line, and reports the deployed site is several commits behind the repo — /concepts/permissions/ and /reference/migration/ are live 404s.
188 lines
7.5 KiB
Markdown
188 lines
7.5 KiB
Markdown
# docs.beardrive.ai
|
|
|
|
The public product documentation: CLI, sync model, self-hosting. Astro +
|
|
[Starlight](https://starlight.astro.build), static output, Pagefind search.
|
|
|
|
```sh
|
|
npm install
|
|
npm run dev # http://localhost:4321
|
|
npm run build # -> dist/
|
|
npm run preview
|
|
|
|
npm run check:sitemap http://localhost:4321 # after npm run preview
|
|
npm run check:sitemap https://docs.beardrive.ai # or against production
|
|
```
|
|
|
|
## The sitemap
|
|
|
|
`/sitemap-index.xml` -> `/sitemap-0.xml`, generated by `@astrojs/sitemap`.
|
|
Starlight would add that integration itself; `astro.config.mjs` declares it
|
|
explicitly instead, which replaces Starlight's default rather than duplicating
|
|
it and is what makes the `lastmod` option reachable.
|
|
|
|
Each URL carries a `<lastmod>` taken from the **commit date of the markdown
|
|
behind it** — see "Checkout depth" below for the one way that goes wrong. The
|
|
redirect stubs are correctly absent: Astro emits them as `noindex` meta-refresh
|
|
pages, and `@astrojs/sitemap` leaves them out.
|
|
|
|
`public/robots.txt` exists to carry the `Sitemap:` line: robots.txt is the one
|
|
file a crawler fetches without being told, and the landing page's robots.txt is
|
|
on a different host so it cannot point here.
|
|
|
|
**Verify that file survives the deploy.** This host is behind Cloudflare, which
|
|
was serving a managed robots.txt of its own (the content-signals policy block)
|
|
back when the origin had none — a body with no `Sitemap:` and, in fact, no
|
|
directives at all. Whether an origin robots.txt now replaces that or gets merged
|
|
with it is Cloudflare's call, not this repo's, so after deploying check the
|
|
served file rather than the built one:
|
|
|
|
```sh
|
|
curl -s https://docs.beardrive.ai/robots.txt | grep -i sitemap
|
|
```
|
|
|
|
`check:sitemap` is the same set of checks Search Console runs — follow
|
|
robots.txt to the index, parse it, confirm every advertised URL returns 200 —
|
|
and takes any origin, so it works against a local preview and against
|
|
production. It is worth running against production after a deploy: a sitemap
|
|
that has gone stale or started advertising 404s looks completely fine until
|
|
something crawls it, which is weeks later.
|
|
|
|
## Why this is a standalone site
|
|
|
|
Unlike the hub frontend (`internal/webapp/static`) and the cloud landing page
|
|
(`cloud/internal/landing/dist`), this is **not** embedded into a Go binary:
|
|
|
|
- Docs change far more often than the binary. Embedding them would mean a Go
|
|
rebuild and redeploy to fix a typo.
|
|
- A Pagefind search index has no business shipping inside every self-hoster's
|
|
install.
|
|
|
|
It lives in the OSS repo because that is what it documents. "Edit this page"
|
|
resolves to something an outside contributor can open a PR against.
|
|
|
|
## Design tokens
|
|
|
|
`scripts/tokens.mjs` reads the `@theme` block in
|
|
`internal/webapp/frontend/src/tw.css` — the source of truth for the BearDrive
|
|
palette — and emits `src/styles/tokens.gen.css`. That file is generated,
|
|
gitignored, and regenerated by `npm run dev` and `npm run build`, so the palette
|
|
cannot drift. `src/styles/custom.css` maps Starlight's `--sl-color-*` variables
|
|
onto it and invents no colors of its own.
|
|
|
|
(The cloud landing page can't do this — it sits in a different Go module and
|
|
keeps a *copy* of the tokens, policed by its own `check-tokens.mjs`.)
|
|
|
|
## Adding a page
|
|
|
|
Drop a `.md` file under `src/content/docs/<section>/` with `title` and
|
|
`description` frontmatter, then add its slug to the `sidebar` in
|
|
`astro.config.mjs`. Sidebar order is explicit, not alphabetical.
|
|
|
|
Write `description` for every page: it is the meta description, the search
|
|
result snippet, and what `llms.txt` shows.
|
|
|
|
## llms.txt
|
|
|
|
`starlight-llms-txt` generates `/llms.txt`, `/llms-small.txt`, and
|
|
`/llms-full.txt` at build time.
|
|
|
|
Convention puts `llms.txt` at the **root** domain, not a docs subdomain — so
|
|
`beardrive.ai/llms.txt` should redirect or proxy to `docs.beardrive.ai/llms.txt`.
|
|
That redirect lives in the cloud landing page and is the one cross-repo
|
|
coordination point this split introduces.
|
|
|
|
## Structure
|
|
|
|
The sidebar order in `astro.config.mjs` **is** the recommended path, and the
|
|
recommended path is agent-first:
|
|
|
|
- **Start here** — what it is, set up with your agent, first hour. No `brew
|
|
install` appears in this group.
|
|
- **Working with agents** — the workflows the product exists for.
|
|
- **Manual setup (optional)** — the CLI route: install, set up by hand, skills
|
|
and hooks in detail. Same destination, more steps; one click away, never on
|
|
the critical path.
|
|
- **Use cases** — job-shaped titles ("Share work across your team's agents"),
|
|
persona named in the first line and in the `description`. These pages ROUTE:
|
|
who it's for, what you get, the one setup difference, links out. The moment
|
|
one starts teaching a feature, it links to the guide that owns it instead.
|
|
- **Self-hosting**, **Reference**, **Concepts** — unchanged in intent.
|
|
|
|
Keep new onboarding content out of Manual. If a page teaches someone how to get
|
|
started, it belongs in Start here and should say what to ask an agent, not what
|
|
to type.
|
|
|
|
## Deploying
|
|
|
|
Static output in `dist/`. Any static host works; build command `npm run build`,
|
|
output directory `dist`, project root `web/docs`.
|
|
|
|
### Redirects
|
|
|
|
The docs were reorganized around the agent-first path, so three old URLs moved:
|
|
|
|
| Old | New |
|
|
|---|---|
|
|
| `/start/install` | `/manual/install/` |
|
|
| `/start/quickstart` | `/manual/setup-by-hand/` |
|
|
| `/guides/connect-an-agent` | `/start/setup/` |
|
|
|
|
`astro.config.mjs` declares these, which in a static build emits **meta-refresh
|
|
pages** — fine for humans, weak for search engines. Real 301s belong in the host.
|
|
|
|
**Firebase Hosting** (simplest static option on GCP — CDN, TLS, and custom
|
|
domains included):
|
|
|
|
```json
|
|
{
|
|
"hosting": {
|
|
"public": "dist",
|
|
"ignore": ["firebase.json", "**/.*"],
|
|
"redirects": [
|
|
{ "source": "/start/install", "destination": "/manual/install/", "type": 301 },
|
|
{ "source": "/start/quickstart", "destination": "/manual/setup-by-hand/", "type": 301 },
|
|
{ "source": "/guides/connect-an-agent", "destination": "/start/setup/", "type": 301 }
|
|
]
|
|
}
|
|
}
|
|
```
|
|
|
|
**Cloud Storage behind an external Application Load Balancer:** put the rules in
|
|
the URL map, which redirects before the bucket is ever reached.
|
|
|
|
```sh
|
|
gcloud compute url-maps edit docs-url-map # pathMatchers[].pathRules[]:
|
|
# - paths: ["/start/install"]
|
|
# urlRedirect:
|
|
# pathRedirect: "/manual/install/"
|
|
# redirectResponseCode: MOVED_PERMANENTLY_DEFAULT
|
|
# stripQuery: false
|
|
```
|
|
|
|
Whichever host wins, keep the Astro `redirects` block as well: it is the
|
|
portable fallback, and it keeps local `npm run preview` honest.
|
|
|
|
Note that the build reads a file **outside** `web/docs` (the token source), so
|
|
the host must check out the whole repository rather than just this
|
|
subdirectory.
|
|
|
|
### Checkout depth
|
|
|
|
Check out with **full history**, not a shallow clone. The sitemap's `<lastmod>`
|
|
for each page is the commit date of the markdown behind it, so a depth-1
|
|
checkout — the default for `actions/checkout` and for most build hosts — has
|
|
nothing to read the dates from.
|
|
|
|
```yaml
|
|
- uses: actions/checkout@v4
|
|
with:
|
|
fetch-depth: 0 # sitemap <lastmod> comes from commit dates
|
|
```
|
|
|
|
Getting this wrong degrades rather than breaks: `astro.config.mjs` detects the
|
|
shallow clone and emits **no** `lastmod` at all, because the alternative is
|
|
stamping all 25 pages with the one commit a shallow clone has, and a sitemap
|
|
that claims the whole site changed on every deploy is one Google learns to
|
|
ignore. So the symptom is a silently less useful sitemap — check for `<lastmod>`
|
|
in the deployed `sitemap-0.xml` after changing hosts.
|