Files
context-extractor/CURSOR_PROMPT.md
T
ilia d26d1aaa4e
CI / automation-js (vitest + real Chromium) (pull_request) Successful in 1m15s
CI / Lint + tests (pull_request) Successful in 10m42s
Unify product version at 1.4.0, add ruff gate, drop stale GitHub workflow
- One version (1.4.0) across extension/manifest.json, automation/pyproject.toml,
  automation-js/package.json (+lockfile); README documents the bump-together policy
- Delete .github/workflows/ci.yml (behind the Gitea copy; CI runs on Gitea Actions)
- CURSOR_PROMPT.md: replace personal paths with <path-to-context-extractor>
- Add ruff (E4,E7,E9,F,I, line-length 120) to automation/ dev extras, Makefile lint,
  and the Gitea CI Lint + tests job as a hard gate
2026-07-26 15:47:37 -04:00

41 lines
1.6 KiB
Markdown

# Cursor consumer prompt
Paste this into another Cursor chat when that project should use Context Extractor.
Replace `<path-to-context-extractor>` with wherever you cloned this repo.
---
```text
Use Context Extractor from <path-to-context-extractor> for browser capture / scrape debugging.
Setup (once):
cd <path-to-context-extractor>/automation && python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
playwright install chromium
# optional anti-detect: pip install -e ".[camoufox]"
Python scrape / CI loop:
from playwright.sync_api import sync_playwright
# or: from camoufox import Camoufox
from context_extractor import ExtractorSession
with sync_playwright() as p:
page = p.chromium.launch(headless=True).new_page()
session = ExtractorSession(page) # attach BEFORE goto
page.goto(URL, wait_until="domcontentloaded")
prompt = session.build_ai_prompt(SELECTOR) # None = whole body
# also available: session.extract_markdown(), session.get_store(), page.content() for raw HTML
CLI: context-extractor URL --selector CSS --engine chromium|camoufox --out prompt.md
Manual / interactive (Brave or Chrome):
Load unpacked: <path-to-context-extractor>/extension
Then use "Copy" on the broken page.
When debugging scrapers or flaky pages in THIS project:
1. Capture with ExtractorSession / the extension.
2. Paste the prompt into chat and ask for selector fixes, parse logic, or why the page failed.
3. Prefer selectors from #main / data-* / stable IDs over brittle chrome-copy paths.
Do not reinvent markdown extraction or console/network capture — reuse this package.
```