feat: add vision pipeline scaffolding with ONNX backend

Introduce the app/services/vision/ module with ABC interfaces, ONNX
Runtime backend, model registry, and per-task implementations:
- OpenCLIP ViT-B/32 embedder (image + text, 512-d)
- RapidOCR engine (PP-OCRv4 via ONNX, no PaddlePaddle)
- YOLOv8n object detector (raw ONNX, no ultralytics runtime)
- YuNet + SFace face processor (Apache 2.0, opencv_zoo, 128-d)
- DBSCAN face clustering helper

Add VisionSettings to config (mulita.yml + Pydantic), bootstrap_models.py
for first-boot weight downloads, models_data Docker volume, and ROCm
backend stub for future GPU acceleration.

No Celery tasks wired yet — models load but nothing invokes them.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-04-10 09:00:06 +02:00
parent dea04ceed9
commit 9282a5c734
16 changed files with 856 additions and 2 deletions

View File

@@ -22,3 +22,28 @@ performance:
cache_ttl: 3600
db_pool_size: 20
db_pool_recycle: 3600
# AI vision pipeline — embedding, OCR, object detection, face recognition.
# Runs on the dedicated `vision` Celery queue (PR4+). Set enabled: false
# to disable all vision processing.
vision:
enabled: true
backend: onnx # "onnx" (CPU) | "rocm" (future GPU)
models_dir: /data/models
embedder:
name: openclip_vitb32
batch_size: 8
ocr:
enabled: true
languages: [en]
min_confidence: 0.5
detector:
enabled: true
min_confidence: 0.35
max_detections: 50
faces:
enabled: true
min_face_size: 40
recognition_threshold: 0.4
cluster_eps: 0.35
worker_concurrency: 2