Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 17 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -72,10 +72,26 @@ jobs:
uv run python bin/kb/index --rebuild --json
KB_ROOT="$PWD" /tmp/brain-eval --json

ocr:
name: OCR (tesseract fixture)
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
- name: Install tesseract + poppler
run: |
sudo apt-get update
sudo apt-get install -y --no-install-recommends \
tesseract-ocr tesseract-ocr-eng tesseract-ocr-deu poppler-utils
- name: Go OCR tests (synthetic HELLO PNG)
run: go test ./internal/ocr -count=1

release:
name: Release (semver)
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
needs: test
needs: [test, ocr]
runs-on: ubuntu-latest
permissions:
contents: write
Expand Down
9 changes: 5 additions & 4 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ skills/ in-project agent skills (vendored, no external links)
bin/ self-describing tools bin/{subject}/{method}.go (shebang)
bin/brain/ search.go serve.go index.go add.go get.go stats.go eval.go watch.go
bin/chats/ sync.go import.go facts.go apply.go; libs in internal/chats
bin/mail/ sync.go import.go (index_mail → brain/index.go)
bin/mail/ sync.go import.go ocr.go (index_mail → brain/index.go)
bin/markdown/ import.go (H2 leaf split; Python bin/md/import fallback)
bin/postgres/ query.go (read-only YAML)
bin/git/ import.go (go-git history; Python shim execs it)
Expand Down Expand Up @@ -71,9 +71,9 @@ bin/brain/index.go --rebuild --with-facts --with-chats
- `sync` (Go) downloads messages + attachments; Gmail uses paginated list +
`body.attachmentId` (not partId) for attachments.
- `import` converts body + attachments to markdown. PDFs use poppler
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs fall back to
docling (isolated subprocess — its native onnx can segfault the parent).
Conversion never touches the brain DB (crash safety).
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs use
`pdftoppm` + tesseract `eng+deu` (`bin/mail/ocr.go`). Optional
`OCR_ENGINE=paddle`. Conversion never touches the brain DB (crash safety).
- `index_mail` is a deprecation shim for `bin/brain/index.go --rebuild`. Bulk
rebuild still deletes `var/kb.lbug` and creates FTS/HNSW last. Single-leaf
write is `bin/brain/add.go` (safe while indexes exist; do not DROP INDEX).
Expand Down Expand Up @@ -101,6 +101,7 @@ bin/git/import.go [REPO] [--json] [--limit N] # go-git history → commit le
bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence
bin/reasoner/bakeoff.go [--model ID] [--json] # D18 CPU tool-call bake-off
bin/postgres/query.go --profile onlyoffice -c 'SELECT 1'
bin/mail/ocr.go <image|pdf> # tesseract eng+deu (scans)
bin/md/tables # what the graph holds → YAML
bin/brain/deduce "question" # thinking wrapper
```
Expand Down
4 changes: 4 additions & 0 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,10 @@ ENV PYTHONUNBUFFERED=1 \

WORKDIR /app
RUN id -u 2dph 2>/dev/null || useradd --create-home --uid 1001 2dph
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
poppler-utils tesseract-ocr tesseract-ocr-eng tesseract-ocr-deu \
&& rm -rf /var/lib/apt/lists/*

COPY requirements.lock.txt /tmp/requirements.lock.txt
RUN python -m pip install --no-cache-dir -r /tmp/requirements.lock.txt \
Expand Down
20 changes: 13 additions & 7 deletions PLAN.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,9 @@ A brain that loves facts and deduction. Evidence-first knowledge graph + hybrid
RAG over the operational Brain/ops/eSlider stack. Built like Sherlock
Holmes: nothing is asserted unless it has proof.

Status: **in progress** — read path + MCP work; v1 goal is [epic #16](https://git.produktor.io/eSlider/2dph/issues/16)
(milestone [v1 detective brain](https://git.produktor.io/eSlider/2dph/milestone/12)).
Status: **v1 in** (epic [#16](https://git.produktor.io/eSlider/2dph/issues/16) closed).
v2 board: milestone [v2](https://git.produktor.io/eSlider/2dph/milestone/13) — OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6),
[#29](https://git.produktor.io/eSlider/2dph/issues/29) OQ1, [#30](https://git.produktor.io/eSlider/2dph/issues/30) OQ3.
Gap: [docs/roadmap.md](docs/roadmap.md).

## What
Expand Down Expand Up @@ -74,6 +75,7 @@ detective method: **a fact needs ≥2 independent sources or it is
reasoner/bakeoff.go CPU tool-call bake-off (D18; OpenAI tools)
chats/sync.go import.go facts.go apply.go
(libs in internal/chats; no chats index)
mail/ocr.go tesseract eng+deu (pdftoppm scans)
md/import (deprecated; bin/markdown/import.go)
brain/extract brain/audit brain/deduce (thinking wrapper)
web/search (deprecated shim → web/search.go)
Expand Down Expand Up @@ -115,10 +117,13 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
## Open questions (v2)

- OQ1: mutually-contradicting evidence — how to resolve (authority weighting,
temporal freshness, audit adjudication). **v2**; does not block epic #16.
- OQ2: OCR — poppler `pdftotext` fast-path exists; scans still docling.
[#6](https://git.produktor.io/eSlider/2dph/issues/6) (v2, does not block #16).
temporal freshness, audit adjudication). **v2**; [#29](https://git.produktor.io/eSlider/2dph/issues/29).
- OQ2: OCR — **in**. `pdftotext -layout` first; scans `pdftoppm` + tesseract
`eng+deu` (`bin/mail/ocr.go`, `internal/ocr`). No gocv, no gosseract CGO
(D21 Zig owns Ladybug CGO). Optional `OCR_ENGINE=paddle` / compose profile
`ocr-paddle`. Docling left the default path. [#6](https://git.produktor.io/eSlider/2dph/issues/6).
- OQ3: optional duckdb-md layer for `SELECT … FORMAT MARKDOWN` export/write-back.
[#30](https://git.produktor.io/eSlider/2dph/issues/30).
- OQ4: YAML-first storage for leafs — deferred: JSON is ~10x faster to
serialize and unambiguous; YAML only where humans edit files.

Expand All @@ -127,7 +132,8 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
1. `bin/mail/sync.go` (Go, 8 workers) — paginated Gmail/OnlyOffice download.
Gmail attachments key off `body.attachmentId`, not MIME `partId`.
2. `bin/mail/import.go --from-raw` — message.json → message.md; PDFs via
`pdftotext -layout` (~15ms) with docling subprocess fallback; ICS sidecars
`pdftotext -layout` (~15ms); textless/scanned PDFs `pdftoppm` + tesseract
`eng+deu`. ICS sidecars
Latin-1→UTF-8 normalized.
3. `bin/brain/index.go --rebuild` — fresh rebuild (repo corpus + mail) because ladybug
corrupts its WAL on bulk-insert into an already-indexed DB. Conversion and
Expand Down Expand Up @@ -177,4 +183,4 @@ Narrative: [docs/roadmap.md](docs/roadmap.md).
| 4 | [#15](https://git.produktor.io/eSlider/2dph/issues/15) | **in** — lever/loop documented (`search` → `get` → `audit`). |
| 5 | [#19](https://git.produktor.io/eSlider/2dph/issues/19) | **in** — CI recall SoT is `bin/brain/eval.go` via Zig. Python `bin/kb/eval` stays as an explicit fallback. |

Does **not** block epic close: [#6](https://git.produktor.io/eSlider/2dph/issues/6) OCR, OQ1, OQ3, OQ4.
Does **not** block epic close: OQ1 [#29](https://git.produktor.io/eSlider/2dph/issues/29), OQ3 [#30](https://git.produktor.io/eSlider/2dph/issues/30), OQ4. OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6) is **in**.
88 changes: 11 additions & 77 deletions bin/mail/import
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@
bin/mail/import --since 2026-01-01 only messages after a date
bin/mail/import --limit 50 cap messages per run
bin/mail/import --no-attachments body only, skip attachment conversion
bin/mail/import --ocr OCR scanned PDFs/images via docling
bin/mail/import --ocr OCR images (PDFs OCR when textless)
bin/mail/import --dry-run list messages without writing anything

Writes one directory per message: var/mail/{folder}/{message_id}/
Expand All @@ -16,10 +16,11 @@ Writes one directory per message: var/mail/{folder}/{message_id}/
attachments/*.md converted attachment content

Indexing is a separate step (`bin/brain/index.go --rebuild`): conversion can
crash in native docling and must not leave the brain DB mid-transaction.
crash and must not leave the brain DB mid-transaction.

Requires ONLYOFFICE_URL/USER/PASS in .env (or env). Idempotent: a message
already present (message.md exists) is skipped unless --force.
Requires ONLYOFFICE_URL/USER/PASS in .env (or env) except `--from-raw`.
Idempotent: a message already present (message.md exists) is skipped unless
--force.
"""
from __future__ import annotations

Expand All @@ -41,10 +42,10 @@ from mailconv import ( # noqa: E402
IMAGE_SUFFIXES,
LEGACY_OFFICE_SUFFIXES,
TEXT_SUFFIXES,
convert_pdf,
html_to_markdown,
is_convertible,
normalize_markdown,
subject_to_filename,
ocr_image,
zip_extract_safe,
)

Expand Down Expand Up @@ -147,77 +148,16 @@ def convert_file_to_md(path: Path, ocr: bool) -> str | None:
except Exception as e:
return f"\n<!-- conversion failed: {e} -->\n"
if suffix == ".pdf":
return _convert_pdf(path, ocr)
return convert_pdf(path, ocr)
if suffix in IMAGE_SUFFIXES and ocr:
return _convert_pdf(path, ocr)
return ocr_image(path) or "\n<!-- ocr unavailable -->\n"
if suffix in LEGACY_OFFICE_SUFFIXES:
return _convert_legacy(path)
if suffix in ARCHIVE_SUFFIXES:
return None # handled by caller (unpack + recurse)
return None


def _convert_pdf(path: Path, ocr: bool) -> str:
"""Convert one PDF to markdown.

Fast path: poppler's pdftotext (-layout) extracts exact text from
born-digital PDFs in ~15ms vs docling's 1-3s. Only textless PDFs (scanned
pages, layout-heavy) fall back to docling, which runs isolated in a
subprocess because its native onnx/RT-DETR has segfaulted the main process.
"""
text = _pdf_fast_text(path)
if ocr or text is None or not text.strip():
return _convert_pdf_docling(path, ocr)
return normalize_markdown(text)


def _pdf_fast_text(path: Path) -> str | None:
"""pdftotext -layout; None when poppler is unavailable (or the PDF has no text layer)."""
try:
proc = subprocess.run(
["pdftotext", "-layout", str(path), "-"],
capture_output=True, timeout=60)
except (OSError, subprocess.TimeoutExpired):
return None
if proc.returncode != 0:
return None
return proc.stdout.decode("utf-8", errors="replace")


def _convert_pdf_docling(path: Path, ocr: bool) -> str:
try:
proc = subprocess.run(
[sys.executable, os.path.abspath(__file__), "--pdf-worker", str(path),
"--ocr" if ocr else "--no-ocr"],
capture_output=True, text=True, timeout=600)
except subprocess.TimeoutExpired:
return "\n<!-- pdf conversion timed out -->\n"
if proc.returncode != 0:
tail = proc.stderr.strip().splitlines()[-3:]
return f"\n<!-- pdf conversion failed: {proc.returncode}: {' | '.join(tail)} -->\n"
return proc.stdout


def _pdf_worker(path: Path, ocr: bool) -> None:
"""docling worker entry: prints converted markdown on stdout, exits non-zero on error."""
try:
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.datamodel.pipeline_options import PdfPipelineOptions
opts = PdfPipelineOptions()
opts.do_ocr = bool(ocr)
opts.do_table_structure = True
conv = DocumentConverter(format_options={"pdf": PdfFormatOption(pipeline_options=opts)})
res = conv.convert(str(path))
sys.stdout.write(normalize_markdown(res.document.export_to_markdown()))
sys.exit(0)
except Exception as e:
# errors/stacktraces to stderr; the caller only reports a one-liner
print(f"pdf-worker: {e}", file=sys.stderr)
import traceback
traceback.print_exc(file=sys.stderr)
sys.exit(1)


def _convert_legacy(path: Path) -> str:
"""Legacy .doc/.xls/.ppt -> md via pandoc (installed) or a stub."""
try:
Expand Down Expand Up @@ -356,19 +296,12 @@ def main(argv: list[str]) -> int:
p.add_argument("--from-raw", default="",
help="convert Go-synced dirs (var/mail/<folder>/<id>/message.json) to markdown")
p.add_argument("--no-attachments", action="store_true", help="skip attachment download+convert")
p.add_argument("--ocr", action="store_true", help="OCR scanned PDFs/images via docling")
p.add_argument("--ocr", action="store_true", help="OCR images (PDFs OCR when textless)")
p.add_argument("--force", action="store_true", help="re-import even if message.md exists")
p.add_argument("--dry-run", action="store_true", help="list messages, write nothing")
p.add_argument("--json", action="store_true")
p.add_argument("--pdf-worker", default="", help=argparse.SUPPRESS)
p.add_argument("--no-ocr", action="store_true", help=argparse.SUPPRESS)
a = p.parse_args(argv)

if a.pdf_worker:
_pdf_worker(Path(a.pdf_worker), ocr=not a.no_ocr)
return 0

conf = load_env()
fid = folder_id(a.folder)
out_root = ROOT / "var" / "mail"
summary: list[dict] = []
Expand All @@ -394,6 +327,7 @@ def main(argv: list[str]) -> int:
target_dir=msg_dir.parent))
summary.append(entry)
else:
conf = load_env()
OOCLIENT = OOClient(conf)
if a.id:
messages = [{"id": i} for i in a.id]
Expand Down
48 changes: 48 additions & 0 deletions bin/mail/ocr.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
//usr/bin/env go run -tags=mail_ocr "$0" "$@"; exit
//go:build mail_ocr
//
// bin/mail/ocr.go - OCR an image or scanned PDF (tesseract eng+deu).
//
// ./bin/mail/ocr.go scan.png
// ./bin/mail/ocr.go scan.pdf
// OCR_ENGINE=paddle ./bin/mail/ocr.go scan.png
//
// PDFs try pdftotext -layout first; empty text layer uses pdftoppm + tesseract.
// No gocv. Tesseract CGO bindings are not used (D21 Zig owns Ladybug CGO).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main

import (
"fmt"
"os"
"strings"

"github.com/eSlider/2dph/internal/ocr"
)

func main() {
os.Exit(run(os.Args[1:]))
}

func run(args []string) int {
if len(args) != 1 || strings.HasPrefix(args[0], "-") {
fmt.Fprintln(os.Stderr, `usage: bin/mail/ocr.go <image|pdf>`)
return 2
}
path := args[0]
var (
text string
err error
)
if strings.HasSuffix(strings.ToLower(path), ".pdf") {
text, err = ocr.PDFFile(path)
} else {
text, err = ocr.ImageFile(path)
}
if err != nil {
fmt.Fprintf(os.Stderr, "mail/ocr: %v\n", err)
return 1
}
fmt.Println(text)
return 0
}
Loading
Loading