9 Commits
Author SHA1 Message Date
only-cli 3ca462e353 fix: validate resolved IPs, not hostname strings, in the SSRF guard
PR #5's guard pattern-matched the URL's hostname against a regex, which
both under- and over-blocked: IPv4-mapped IPv6 loopback ([::ffff:127.0.0.1]),
0.0.0.0, and any DNS name that merely resolves to a private address all
sailed through, while a legitimate public hostname like 10.example.com was
wrongly rejected because it starts with "10.". It also only checked the
original URL, so a public URL that redirects to an internal address was
never re-validated.

This replaces the regex with net.isIP + dns.lookup: IP literals are checked
directly (including decoding IPv4-mapped/-compatible IPv6), and DNS names
are resolved first so every address they point to is validated before
connecting. The same check now reruns on every redirect hop for both the
impers and native-fetch transports. Resolving before connecting doesn't
pin the address for the actual connection (neither impers nor fetch expose
that here), so a name that re-resolves differently between this check and
the real connect remains a known, documented residual gap.
2026-08-19 14:43:43 -04:00
only-cli 01b607ac81 support youtube watch pages and transcripts
Watch pages need client JS to become interactive, but the title,
description, view count, and caption tracks already ship inline in
the initial HTML as ytInitialPlayerResponse, so this reads that
directly instead of waiting on the v0.3 headless fallback. Each
caption track becomes a numbered link, and oc do on it fetches the
timedtext transcript, collapsed into one block so it pages through
oc next/oc read like any other long document instead of costing one
block per caption line.

Adds youtubeToHTML and transcriptToHTML alongside feedToHTML in the
distiller, a youtube.com.json shortcut, and offline tests against a
fixture watch page.

Fixes #6
2026-08-19 13:27:52 -04:00
only-cli 42906ced60 spend turns, not tokens: do <n> reads text, near-budget pages print whole, timelines render readably
Three changes with one motive: a second command costs an agent more than
the lines it saves. 'oc do' on a heading or text block now prints the read
instead of refusing, since refusing spends a whole turn naming the command
that should have run. A page that would finish within about four times the
budget is printed whole rather than cut, because the cut moves tokens into
a second command instead of saving them. And social timelines stopped
rendering as one fused paragraph: linkedom splits text nodes around
apostrophes, so fragments are merged back by parent node and edge
whitespace, block elements now end lines, and repeated button labels are
trimmed sooner than links because a button label is never the content.
With that, x.com profiles and posts read without a login, so a six line
cli definition ships for the two x.com pages that work.
2026-08-19 07:59:29 -04:00
only-cli 0435b386ae Lead the render with the page's main content
The budget was being spent on whatever came first in the document, which on
most pages is menus. On the Reddit thread the benchmark uses, all 500 tokens
went to nav, sidebar, and moderator boxes, so an agent that opened the page to
read the discussion had to escalate to oc raw to see a single comment.

distill now finds the content (main, role=main, a single article, else the
densest run of prose) and emits it first, with the rest of the page after it
under a divider. Nothing is dropped, so do <n> still reaches every link.

Link labels that repeat down a page go with it: a thread stamps permalink,
save, and report onto every comment, which cost more than the comments. The
count and three examples are printed in place, and oc raw still has them.

Same thread, before and after: 660 blocks to 254, whole compact page 4,184
tokens to 2,712, and the first view now holds comments instead of a sidebar.
Pages small enough to print whole keep document order.
2026-08-18 17:20:55 -04:00
only-cli 6aaa8f1975 add oc find <query> so a long page answers a lookup in one command
find searches the distilled page the session already holds, prints one line
per match with the number to read it by, and costs no fetch. It matches the
query as a phrase, case insensitive, and falls back to matching the words
separately when the phrase is not there.

On the reddit thread from the benchmark: 'oc find w3m' is 115 tokens against
9,670 for oc raw, and it lands on the numbers to read.
2026-08-18 16:16:43 -04:00
only-cli 0f7a38d44a add oc next and oc read <n> so a long page costs a screenful, not a refetch
The compact view was all or nothing: an agent that needed more than the 500
token budget had only oc raw, ten to twenty times the price. Now open saves
the distilled page, next continues it where the view stopped, and read <n>
prints one region in full. Headings and text blocks long enough to be cut are
numbered so they can be addressed, and the marker prices what it left behind.

On one Reddit thread: open 475 tokens, next 455, read 88, raw 9,670.
2026-08-18 16:11:46 -04:00
only-cli 79ddeebc2e implement oc do so agents can follow a numbered link
The compact view hides link URLs because printing them is most of what
makes a page expensive, which left an agent re-fetching the same page as
--json or raw just to learn where [15] pointed. oc open now saves the
handles it numbered to a small JSON file per session under ~/.only-cli
(OC_HOME relocates it), and oc do <n> resolves one and renders the target
exactly as open would. Search engine tracking redirects are unwrapped so a
result link opens the destination instead of a script page.

Errors name the command that fixes them, since agents read them: an
out-of-range number reports the valid range, an input says to use fill, a
button says the page handles it itself.

Two hops on Hacker News cost about 3k characters this way, against roughly
23k for the re-fetch route.
2026-08-18 15:00:36 -04:00
only-cli 853bbca853 feed reading: offline tests, readme row, and a raw-mode title fix
The feed fixture covers Atom entries (escaped bodies, self-closed
categories, bylines) and an RSS item with a CDATA body. Raw markdown of
a feed exposed an old quirk: cleanDocument removed the head before
toMarkdown read the title, so any page whose body lacked a matching h1
lost its title in raw mode. cleanDocument now captures the title first
and returns both.
2026-08-18 10:40:08 -04:00
only-cli dcc0531ef3 only-cli v0.1: turn websites into a compact CLI for AI agents
Generic distillation engine (no per-site adapters): fetch via impers
impersonating Chrome with a firefox-fingerprint retry, distill to an
interaction tree, render under a hard token budget with numbered action
handles. Raw mode emits markdown via turndown or cleaned HTML. Per-site
CLI definitions for HN, Reddit, Bing, DuckDuckGo. Offline test suite,
agent skill, OIDC publish workflow.
2026-08-18 09:01:41 -04:00