breakpilot-pwa

Archived

This repository has been archived on 2026-02-15. You can view files and clone it. You cannot open issues or pull requests or push a commit.

Files

T

History

BreakPilot Dev ee0c4b859c feat(klausur-service): Add Tesseract OCR, DSFA RAG, TrOCR, grid detection and vocab session store

New modules:
- tesseract_vocab_extractor.py: Bounding-box OCR with multi-PSM pipeline
- grid_detection_service.py: CV-based grid/table detection for worksheets
- vocab_session_store.py: PostgreSQL persistence for vocab sessions
- trocr_api.py: TrOCR handwriting recognition endpoint
- dsfa_rag_api.py + dsfa_corpus_ingestion.py: DSFA RAG corpus search

Changes:
- Dockerfile: Install tesseract-ocr + deu/eng language packs
- requirements.txt: Add PyMuPDF, pytesseract, Pillow
- main.py: Register new routers, init DB pools + Qdrant collections

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

2026-02-10 00:00:19 +01:00

backend

feat(klausur-service): Add Tesseract OCR, DSFA RAG, TrOCR, grid detection and vocab session store

2026-02-10 00:00:19 +01:00

docs

fix: Restore all files lost during destructive rebase

2026-02-09 09:51:32 +01:00

embedding-service

fix: Restore all files lost during destructive rebase

2026-02-09 09:51:32 +01:00

frontend

fix: Restore all files lost during destructive rebase

2026-02-09 09:51:32 +01:00

scripts

fix: Restore all files lost during destructive rebase

2026-02-09 09:51:32 +01:00

Dockerfile

feat(klausur-service): Add Tesseract OCR, DSFA RAG, TrOCR, grid detection and vocab session store

2026-02-10 00:00:19 +01:00