Skip to content

Capturing sources

You capture a source by pointing Khiip at its URL:

Terminal window
khiipd capture https://example.com/some/article

Khiip routes the URL to the right extractor, which emits a Pydantic-typed payload, renders canonical Markdown into your vault, and preserves the raw Source-tier bytes. (Optionally, if you enable it, it also submits the URL to the Wayback Machine as a witness — see below.)

Sources today

SourceWhat it captures
XFull QRT chains, X-Article body (block-structured), embedded media, engagement metrics, community notes. Works anonymously via fxtwitter.
RedditPost + recursive comment tree (deep “load more” branches followed credential-free) + galleries (distinct full-res images) + crosspost + removed-status preservation. Credential-free by default (old.reddit HTML); an optional Reddit app adds rate headroom + gallery dimensions/captions — see Installation.
WikipediaStructured article via the MediaWiki action API (sections + page image + canonical URL) → REST summary (fallback); references + infobox best-effort.
Generic webArticle body via trafilatura (primary) → readability (fallback) → OG/JSON-LD enrichment.
YouTubeMetadata + transcripts via yt-dlp → oEmbed + transcript-api → Data API v3 (the optional API key widens the chain).
PDF (experimental)Text + structure via markitdown → pdfplumber (fallback). Not yet a first-class source.

First-class PDF (with a file-drop flow), Instagram, TikTok, Threads, and Bluesky are on the roadmap.

What lands per capture

  • Canonical Markdown with YAML frontmatter under ~/khiip-vault/captures/<source>/
  • A typed payload (TweetPayload, RedditPayload, WebPayload, WikiPayload, YouTubePayload, PDFPayload) — see Typed payloads
  • Raw Source-tier bytes preserved under your configured data_root, as insurance against upstream rot
  • A Wayback witness (opt-in; off by default) — archive.org’s anonymous Save-Page-Now is rate-limited and unreliable, so it’s off unless you set [archive] wayback_enabled = true. When on, it’s best-effort: the result lands in archive_urls and failures are quiet (no callout). Reliable archiving needs your own archive.org credentials (a BYO-credentials tier is planned).

Media

Media fetching walks a registry: HttpxFetcher (photos) → optionally YtDlpFetcher (video; opt-in via [media] download_videos = true) → GalleryDlFetcher (wide-coverage fallback). Video preservation is opt-in and off by default.

Refetching media

khiipd refetch <id> --media re-walks the registry on an existing capture, in place. By default it only fills gaps — an item already downloaded and still on disk is left untouched (no re-download, no churn). Add --force to re-attempt already-successful items.

Even under --force, the re-download is honest about what each fetcher can redo:

  • Photos (the HttpxFetcher class) are genuinely re-fetched. If the new bytes differ from the old, the previous copy is kept alongside the fresh one under a versioned name — <name>.khiip-v<timestamp><ext> — while the current bytes land back at the original filename. Identical bytes are a no-op (the file is left exactly as-is).
  • Video (yt-dlp) and gallery (gallery-dl) media stay skip-if-exists even under --force: if the file is already on disk, the underlying tool won’t overwrite it. (Their writers aren’t atomic, so Khiip keeps them pinned rather than risk clobbering a good binary mid-write.)

So --force refreshes photo media and versions the old copy; it does not re-pull an existing video or gallery file.

Hand-edited a captured note?

Every in-place refetch (--re-extract, --re-render, --media, --wayback) first checks whether you’ve edited the note since Khiip last wrote it. If you have, it refuses with a 409 rather than discard your edit — the CLI prints the conflict. Re-run with --force to overwrite the note (discarding your edit) on purpose.

--force never applies to the default network refetch: a network re-fetch writes a new capture instead of rewriting the existing note, so there’s nothing in place to override — khiipd refetch <id> --force (no dimension) returns a 400 saying so.

Partial success

If extraction succeeds but media or Wayback fails, the capture still lands — each sub-system reports its own status independently. See Failure handling (P-δ).