PaperPod

Local, privacy-first tool that converts a single overhead video recording of documents (receipts, letters, multi-page stacks) into individual, properly named PDFs ready for a Paperless-ngx consume folder. All processing runs offline on your machine.

How it works

  1. Record one continuous overhead video, placing documents under the camera one at a time (optionally saying out loud what each one is).
  2. PaperPod finds "stable windows" where nothing is moving, picks the sharpest frame in each, detects the document outline, and produces a perspective-corrected crop. Detection scores candidate contours by shape (filled, convex, plausible aspect ratio) rather than just picking the largest one, and runs on a resolution-independent downscale so it works the same on FHD and 4K phone video.
  3. The crop is auto-rotated upright (0/90/180/270) — no need to place documents facing a particular way.
  4. Tesseract OCR (run against an illumination-normalized "flattened" version of the crop, similar to a scanner app's contrast enhancement) extracts a vendor/date/total for receipts, or a form code + organization name for recognized Canadian tax slips (T4, T4A, T5008, ...), and names the PDF.
  5. (Upcoming) Spoken descriptions (faster-whisper) as an alternate naming source; multi-page stacks grouped into single PDFs; a review step to rename/merge/split before export.

Status

Module Purpose Status
capture/ Frame sampling + audio extraction from video files Built
vision/ Motion detection, document contours, perspective crop, sharpness Built
events/ State machine: placed / page-flipped / cleared, pod grouping Planned
transcribe/ Local speech-to-text (faster-whisper) Planned
ocr/ OCR fallback naming (Tesseract): vendor, date, total Built
naming/ Final filename assembly + summary CSV Built
pdf/ PDF assembly (img2pdf) Built — single-page only, no consume-folder staging
review_cli/ Pre-export review (rename/merge/split) Planned

Known limitation: there is no page-flip / multi-page grouping yet (that's events/). Every detected document currently becomes its own single-page PDF, even if it was physically one page of a stack or one side of a double-sided document. Don't point export at the Paperless-ngx consume folder for multi-page documents until events/ exists — you'll get one PDF per page instead of one PDF per document.

Setup

Requires Python 3.11+ and ffmpeg (brew install ffmpeg).

python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Usage

# Generate a synthetic test video (no camera needed)
python sample_data/make_sample_video.py

# Inspect a video
python -m paperpod probe sample_data/videos/synthetic_sample.mp4

# Detect stable windows + document candidates
python -m paperpod detect sample_data/videos/synthetic_sample.mp4

# Detect, OCR-name, and export each document as its own PDF
python -m paperpod export sample_data/videos/receipts_sample.mp4

# Extract the audio track (16 kHz mono WAV, whisper-ready)
python -m paperpod extract-audio my_recording.mp4

detect writes to output/<video-name>/:

  • crops/window_NNN.png — perspective-corrected document candidates
  • frames/window_NNN_full.png — the full best frame per window (debugging)
  • report.json — video metadata, motion events, stable windows, detections
  • motion_scores.csv — per-sample motion scores, for tuning motion.threshold

export runs detect and then, for every detected document, additionally writes:

  • pdf/<name>.pdf — one single-page PDF per detected document, named YYYY-MM-DD_vendor.pdf from OCR-extracted date/vendor (or <form_code>_<org_name> for recognized tax slips), or UNSORTED_<timestamp>.pdf if OCR couldn't find anything usable
  • export_summary.csv — timestamp, pod_id, page_count, source_of_name (ocr/none), OCR vendor/form_code/date/total/confidence, detection confidence, rotation applied, final filename

Before naming, each crop is auto-rotated to be right-side-up (Tesseract's built-in orientation detection, falling back to a 4-way OCR-confidence sweep when a crop has too little text for that to work) — you don't need to worry about which way documents face when placing them.

Nothing is copied into a Paperless-ngx consume directory yet — review the pdf/ folder yourself before moving files anywhere.

Tuning

All thresholds live in config.yaml (motion sensitivity, stable window duration, contour size/shape filters, Canny thresholds, detection downscale width, auto-rotate/enhance toggles, speech matching window, output paths). Plot motion_scores.csv to pick a motion.threshold that separates your camera's noise floor from real hand movement.

Recording tips (from real-world testing): shoot in 4K, keep hands out of frame once a document is placed, and use a plain dark, matte (non-glossy) surface under documents — busy wood grain or shiny surfaces make contour detection and OCR meaningfully harder. Document detection and orientation correction are tuned against real phone-camera footage (receipts on a wood table and multi-page CRA tax slips), not just synthetic test videos.

Tests

python -m pytest

Tests are self-contained: they synthesize frames and tiny videos on the fly, no sample assets required.

Description
Local privacy-first tool — overhead video of documents → named PDFs for Paperless-ngx
Readme 150 KiB
Languages
Python 97.7%
Shell 2.3%