Discover MCPs & agents
Loading MCPs and agents…
Loading MCPs and agents…
A worked example of multi-agent agentic-IDE development — GitHub Copilot, OpenAI Codex, Claude Code and IBM Bob collaborating in one VS Code workspace via a shared docs/HANDOFF.md continuity doc — building a local-first, privacy-first invoice sorter for German/English tax preparation.
From the repo.

Note: This project and its documentation were developed under human direction and review. See CONTENT_PROVENANCE.md for details.
A local-first command-line and desktop tool that scans a folder of PDF and image invoices/receipts, extracts metadata, classifies each document into configurable categories, copies it into a category folder, and produces a Markdown summary for a tax advisor plus a JSONL audit log.
Built for a private user organizing invoices for a tax advisor. It runs entirely on your machine.
For the shortest setup and first dry run, see docs/QUICK_START.md.
For system design, architecture, data flow, and module responsibilities, see ARCHITECTURE.md. The editable Draw.io architecture diagrams contain separate static-structure and dynamic-flow pages.
┌─────────────────────────────────────────────────────┐
│ Input folder (PDFs / JPG / PNG / TIFF) │
└──────────────────────────┬──────────────────────────┘
│ scan
▼
┌─────────────────────────────────────────────────────┐
│ Orchestrator (orchestrator.py) │
│ for each file: │
│ 1. Extract text ──► Docling → light → manual │
│ 2. Extract metadata (vendor · date · amount) │
│ 3. Classify ────────► rule-based + confidence │
│ 4. Route ──────────► category folder or review │
│ 5. Copy / move / dry-run │
└──────┬──────────────────────────────────────────────┘
│ results + summary
▼
┌─────────────────────────────────────────────────────┐
│ Outputs │
│ • invoice_summary.md (Markdown report) │
│ • audit_log.jsonl (one JSON line per file) │
│ • performance_log.json (timing + token metrics) │
│ • Sorted_Invoices/<Category>/ (copied files) │
└─────────────────────────────────────────────────────┘
│ optional — local Ollama only, no upload
▼
┌─────────────────────────────────────────────────────┐
│ Local AI layer (127.0.0.1 only) │
│ • --ai-review post-sort Markdown review │
│ • Agent REST service advice / chat / exec report │
└─────────────────────────────────────────────────────┘
Both the CLI (invoice-sorter) and the PySide6 desktop GUI (invoice-sorter-gui)
call the same orchestrator.run(RunOptions) — the engine is UI-agnostic.
The AI layer is entirely optional; the deterministic pipeline never calls an LLM.
PDF, JPG, JPEG, PNG, TIFF.Sorted_Invoices/<Category>/ (copy, never move, by default).invoice_summary.md, audit_log.jsonl, and an anonymized
performance_log.json.Invoices contain sensitive personal and financial data, so:
Note on Docling: the optional Docling backend downloads layout/table/OCR models on first use (Hugging Face + ModelScope for RapidOCR). That is a one-time setup download — no invoice data leaves your machine. After a one-time warm-up the tool runs fully offline; enforce it with:
export HF_HUB_OFFLINE=1 export TRANSFORMERS_OFFLINE=1Verified on Apple Silicon (torch MPS): with these set, extraction runs with no network access.
Requires Python 3.12.
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e . # core (light deps)
Optional backends/extras:
pip install -e ".[light]" # pdfplumber/pypdf + pytesseract image OCR (fallback)
pip install -e ".[docling]" # Docling extraction (best tables/amounts; heavy)
pip install -e ".[gui]" # PySide6 desktop app
pip install -e ".[agent]" # in-app LangGraph agent (chat / advice / exec report)
pip install -e ".[docx]" # DOCX export (planned)
pip install -e ".[test]" # pytest
The extraction backend is auto-selected at runtime: Docling if installed, otherwise the light backend, otherwise files are flagged for manual review. Check which is active with:
python -c "from invoice_sorter.extraction_adapter import active_backend; print(active_backend())"
brew install tesseract
brew install tesseract-lang # adds German (deu); the base install is eng-only
invoice-sorter \
--input "/path/to/input/folder" \
--output "/path/to/output/folder" \
--config "config/categories.yaml" \
--backend auto \
--dry-run
Options:
| Option | Meaning |
|---|---|
--input | Input folder with PDFs and images (required) |
--output | Output folder for sorted invoices and reports (required) |
--config | Path to category configuration (default: bundled config/categories.yaml) |
--backend | Extraction backend: auto, docling, or light (default: auto) |
--dry-run | Analyze only; do not copy files |
--recursive / --no-recursive | Scan subfolders (default: on) |
--move | Move instead of copy (default: copy — safer) |
--ai-review | Append an optional local Ollama sorting review to invoice_summary.md |
--ai-model | Ollama model for --ai-review (default: $OLLAMA_MODEL or deepseek-r1:8b) |
--ai-base-url | Ollama URL for --ai-review (default: http://127.0.0.1:11434) |
--ai-prompt | Custom AI review prompt template file; {json_data} inserts the inspection payload |
--ai-temperature | Ollama sampling temperature from 0.0 to 2.0 (default: 0.2) |
--verbose | Print a per-file line (filenames; avoid when screen-sharing private data) |
--version | Print version and exit |
A native Apple macOS desktop workstation built with PySide6:
source .venv/bin/activate
pip install -e ".[gui]"
invoice-sorter-gui
You can also use ./quickstart.sh for one-click environment validation, backend detection, and launching.
QMenuBar) & Shortcuts: Full integration with Apple menu bar (File, Edit, Run, Tax Rules, AI Agent, View, Help) with standard shortcuts (Cmd+O, Cmd+Shift+O, Cmd+R, Cmd+P, Cmd+E, Cmd+Z, Cmd+K, Cmd+Shift+K, Cmd+D, Cmd+T, Cmd+,).Scanned Docs, Sorted / Classified, Needs Review, Total Gross Volume €), and pill status badges.⏹ Stop safely halts pipeline execution after the in-flight file finishes, writing partial reports and audit logs while preserving all processed rows.Category, Vendor, Invoice Date, Gross, Currency, Confidence, Status, Notes) for direct cell or dropdown editing, with field-level undo tracking (Cmd+Z) and CSV export.➕ Add Category automatically creates and updates a private config/categories.local.yaml (cloned from categories.yaml), keeping confidential vendor and category names strictly local and git-ignored.tax_knowledge.py): Built-in regulations base for tax years 2022–2026 (GWG 800€ net, Homeoffice 6€/day, Bewirtung 70%, 1-year digital AfA) with live Inspector tax hints and automatic prompt augmentation for AI endpoints.TaxRulesWindow: Dedicated German Tax Knowledge Hub window with year switcher and live legal updates.PreferencesWindow (Cmd+,): Modular preferences window organizing AI models, Ollama URLs, temperatures, backend selectors, and processing options.Cmd+1) and document inspector (Cmd+2) for a distraction-free workspace.The GUI automatically launches its local agent REST service on startup (127.0.0.1:8080). Document Advice, Document Chat, and Executive Report are available once a run completes.
Chat / Edit a document. Select a row and click Chat / Edit to open a
dialog that lets you (a) chat with the local agent about that one document
(it answers from the document's metadata only, incorporates active German tax rules, and can suggest a category from
your config) and (b) edit the category and metadata (vendor, dates,
invoice number, gross/VAT/net, currency) and Apply the changes back to the
results table. Category changes are tracked in the correction log (Undo / Export
Corrections). Requires the [agent] extra and a local Ollama server for chat;
editing works without them. The non-GUI edit logic lives in
corrections.py (apply_document_edits).
To correct several documents at once, select multiple table rows and click
Edit Category (Cmd+K). The chosen category is applied to every selected document;
edits and undo remain attached to the correct documents after table sorting.
Models are per-feature. The AI review model field drives the post-sort review, and a separate Agent models row lets you pick a different Ollama model for Advice, Exec Report, and Chat independently. Sensible per-use-case defaults are used (and each is overridable by an environment variable):
| Use case | Default model | Why this default | Environment override |
|---|---|---|---|
| AI review + general fallback | deepseek-r1:8b | Reasoning-focused model with a moderate local footprint for checking counts, confidence signals, and exceptions. | OLLAMA_MODEL |
| Document Advice | deepseek-r1:8b | The same reasoning behavior fits a focused decision about whether one document needs manual review. | OLLAMA_ADVICE_MODEL |
| Executive Report | qwen3-coder:30b | The largest installed default is reserved for the longer structured synthesis across the complete run summary. | OLLAMA_REPORT_MODEL |
| Chat | granite4:tiny-h | The smaller model reduces interactive turn latency while answering from one document's metadata. | OLLAMA_CHAT_MODEL |
These are practical defaults for the models installed on the target workstation, not claims that one model is universally best. Available memory, response time, language quality, and local evaluation results may justify different choices. Environment overrides are read when the application starts; the GUI fields can also override them for the current run.
Set them to models you have installed (ollama list). If a model is missing you
get a clear "pull it or choose another" message instead of an opaque error.
Reasoning-model <think> blocks and JSON-wrapped replies are cleaned
automatically.
Generate Exec PDF (Cmd+P) renders the Markdown report into a formatted PDF (headings,
tables, bold) via QTextDocument.setMarkdown — it no longer dumps raw Markdown.
--dry-run runs the entire analysis — scan, extract, classify, route — and
writes the report and audit log so you can review decisions, but it does not
create the Sorted_Invoices/ tree or copy any file. Re-run without --dry-run
to actually sort.
--ai-review calls a local Ollama server after deterministic sorting finishes
and appends a Local AI sorting review section to invoice_summary.md.
Classification remains rule-based; the AI review does not move files or change
categories.
invoice-sorter \
--input "/path/to/input/folder" \
--output "/path/to/output/folder" \
--backend auto \
--dry-run \
--ai-review \
--ai-model deepseek-r1:8b \
--ai-temperature 0.2
Privacy boundary: the AI review prompt is generated in application code and sends aggregate counts, confidence signals, manual-review reasons, and limited metadata to local Ollama. It never sends full extracted invoice text.
performance_log.json records per-document extraction/processing time under
anonymous IDs (doc_001, etc.). When Ollama is enabled, it also records model,
total/load/prompt-evaluation/inference durations, and prompt/output/total token
counts returned by Ollama. The Markdown report includes total extraction time and
a compact Ollama inference/token summary.
The default runtime template is
config/ai_review_prompt.txt. Copy it to the
git-ignored local override before editing:
cp config/ai_review_prompt.txt config/ai_review_prompt.local.txt
Keep {json_data} where the privacy-filtered inspection payload should be
inserted. Document IDs are pseudonymized, but the payload can contain extracted
metadata such as vendor, date, invoice number, amount, and currency for
low-confidence documents. It never contains full extracted invoice text.
If the placeholder is omitted, the application appends the JSON data after the
custom instructions. Run with:
invoice-sorter \
--input "/path/to/input/folder" \
--output "/path/to/output/folder" \
--dry-run \
--ai-review \
--ai-model deepseek-r1:8b \
--ai-temperature 0.2 \
--ai-prompt config/ai_review_prompt.local.txt
The GUI exposes the same setting through the AI review prompt field and file picker, plus an AI temperature control. Lower values are more deterministic; higher values allow more variation.
Categories live in config/categories.yaml. Each
category has keywords and optional vendors:
categories:
Internet:
keywords: [Internet, DSL, Glasfaser, Router, Mobilfunk]
vendors: [Telekom, Vodafone, 1&1, O2]
Folder names are derived automatically (umlauts transliterated, e.g.
Auto / Mobilität → Auto_Mobilitaet).
Keep private vendor names out of the repo: copy the file to a location
outside version control (or a git-ignored *.local.yaml) and pass it with
--config. The bundled config intentionally contains only generic examples.
scripts/suggest_local_config.py scans a real
folder, finds the files that land in manual review, extracts candidate vendor
tokens, auto-assigns well-known public vendors, and writes a git-ignored
config/categories.local.yaml. It prints only counts — your vendor names go
into the (git-ignored) file, never the console.
python scripts/suggest_local_config.py --input ./your_folder
# add --use-docling to extract with Docling instead of the light backend
Then open config/categories.local.yaml, move the # REVIEW vendor tokens
under the right categories, and run with --config config/categories.local.yaml.
Hybrid extraction (implemented). Docling's Markdown output (table cells,
#headers) classifies worse than plain text, but extracts amounts/VAT better. So the pipeline now uses two views: monetary metadata comes from Docling's rich text, missing non-monetary metadata can fall back to plain text, and classification runs on a plain-text view (the light backend's text when available, elsenormalize_for_classification()of the Markdown). You get Docling-quality amounts with light-quality sorting.
| Score | Meaning |
|---|---|
| 0.90 – 1.00 | Very likely correct |
| 0.70 – 0.89 | Probably correct |
| 0.50 – 0.69 | Needs review |
| below 0.50 | Unclear / manual review |
A file is routed to Unklar / Manuell prüfen when: text is too short / OCR is poor, no category keyword matches, several categories tie, no vendor is detected with low confidence, or confidence is below the configured threshold.
Unknown.Unknown.pdf_extraction_macos
project. Current Ollama features provide post-sort review, document advice,
document chat, and executive reports; they do not change classification. The
docling_preprocessor_factory repo can be wired into
extraction_adapter._extract_with_factory if preferred over plain Docling.Done already: CLI, Docling backend, hybrid extraction, backend selection, optional local Ollama review/advice/chat/reporting, PySide6 desktop GUI, rule-based classifier, Markdown report, JSONL audit log, dry-run, real-data tuning script.
Third-party packages and optional model artifacts retain their own terms. Before
shipping a bundled application, generate a resolved dependency/SBOM report and
review the exact PySide6/Qt, Docling, OCR, and model distribution configuration.
Built wheel and source archives include the project license, license policy, and
third-party notices. The source manifest explicitly excludes private
config/*.local.* files.
tax_preorganizer_public/
pyproject.toml # py3.12; extras: docling, light, gui, agent, docx, test
config/
categories.yaml # generic, committed
categories.local.yaml # git-ignored, your private vendors
ai_review_prompt.txt # default runtime Ollama review template
scripts/
suggest_local_config.py # build categories.local.yaml from a real folder
src/invoice_sorter/
cli.py # `invoice-sorter` entry point
gui.py # `invoice-sorter-gui` entry point (PySide6)
orchestrator.py # run(): scan -> per-file pipeline -> outputs
scanner.py # recursive file collection
extraction_adapter.py # backend selection: factory -> docling -> light
metadata_extraction.py # DE/EN amounts, dates, IBAN, invoice no., vendor
classifier.py # keyword/vendor scoring + confidence
routing.py # confident category vs. manual review
tax_knowledge.py # German tax regulations base & prompt context
corrections.py # field/category edit logic & undo tracking
file_operations.py # safe copy, collision-resolving names
audit_log.py # JSONL writer
performance_log.py # anonymized extraction/Ollama timing + tokens
report.py # Markdown report (RunSummary + build_report)
config.py / constants.py / models.py
tests/ # pytest suite (optional-feature tests skip if unavailable)
examples/sample_invoice_summary.md
The engine is UI-agnostic: both cli.py and gui.py call
orchestrator.run(RunOptions) and render the returned (results, summary).
pip install -e ".[test]"
python scripts/check_license_metadata.py
pytest
Replace {MCP_ENDPOINT_URL} with this MCP’s endpoint URL (from its repo or docs above). No API key — you connect directly.
Tool
OS
Config file: ~/.cursor/mcp.json
{
"mcpServers": {
"mcp-server": {
"url": "{MCP_ENDPOINT_URL}"
}
}
}Paste into mcpServers in the config file. Restart Cursor after saving.
If this MCP is also published on mcpchannel.ai, you can subscribe from Browse and use the gateway config there instead.