Files
context-extractor/CURSOR_PROMPT.md
T
ilia d26d1aaa4e
CI / automation-js (vitest + real Chromium) (pull_request) Successful in 1m15s
CI / Lint + tests (pull_request) Successful in 10m42s
Unify product version at 1.4.0, add ruff gate, drop stale GitHub workflow
- One version (1.4.0) across extension/manifest.json, automation/pyproject.toml,
  automation-js/package.json (+lockfile); README documents the bump-together policy
- Delete .github/workflows/ci.yml (behind the Gitea copy; CI runs on Gitea Actions)
- CURSOR_PROMPT.md: replace personal paths with <path-to-context-extractor>
- Add ruff (E4,E7,E9,F,I, line-length 120) to automation/ dev extras, Makefile lint,
  and the Gitea CI Lint + tests job as a hard gate
2026-07-26 15:47:37 -04:00

1.6 KiB

Cursor consumer prompt

Paste this into another Cursor chat when that project should use Context Extractor. Replace <path-to-context-extractor> with wherever you cloned this repo.


Use Context Extractor from <path-to-context-extractor> for browser capture / scrape debugging.

Setup (once):
  cd <path-to-context-extractor>/automation && python3 -m venv .venv && source .venv/bin/activate
  pip install -e ".[dev]"
  playwright install chromium
  # optional anti-detect: pip install -e ".[camoufox]"

Python scrape / CI loop:
  from playwright.sync_api import sync_playwright
  # or: from camoufox import Camoufox
  from context_extractor import ExtractorSession

  with sync_playwright() as p:
      page = p.chromium.launch(headless=True).new_page()
      session = ExtractorSession(page)          # attach BEFORE goto
      page.goto(URL, wait_until="domcontentloaded")
      prompt = session.build_ai_prompt(SELECTOR)  # None = whole body
      # also available: session.extract_markdown(), session.get_store(), page.content() for raw HTML

  CLI: context-extractor URL --selector CSS --engine chromium|camoufox --out prompt.md

Manual / interactive (Brave or Chrome):
  Load unpacked: <path-to-context-extractor>/extension
  Then use "Copy" on the broken page.

When debugging scrapers or flaky pages in THIS project:
  1. Capture with ExtractorSession / the extension.
  2. Paste the prompt into chat and ask for selector fixes, parse logic, or why the page failed.
  3. Prefer selectors from #main / data-* / stable IDs over brittle chrome-copy paths.
  Do not reinvent markdown extraction or console/network capture — reuse this package.