Drops face recognition, OCR, object detection, and semantic embeddings.
The sole remaining vision task is a CLIP-based binary classifier
(photography vs other); photos in "other" get needs_review=true so
screenshots, documents, memes and scans can be triaged from a new
filter pill in the UI.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- detect_objects, classify_content, recluster_faces now look up the
photo's user_id and set it on created Tag rows — fixes tags being
invisible to the owning user due to NULL user_id
- Initial admin setup creates source root at the mount root (/photos)
instead of a subdirectory, since the admin owns the entire library
- Revert to OpenCLIP ViT-B/32 (512-d) as default embedder — SigLIP
requires transformers version alignment not yet available in the
Docker image. SigLIP2 code remains for future enablement.
- Add transformers to requirements for future SigLIP support
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Centralize execution provider selection in providers.py with
auto-detection and graceful fallback. All ONNX sessions (embedder,
detector, face processor, recognizer) now use the configured providers.
- New VISION_EXECUTION_PROVIDERS env var: "auto" for GPU auto-detect,
or explicit "CUDAExecutionProvider,CPUExecutionProvider"
- Provider priority: CUDA > ROCm > OpenVINO > CPU (when set to "auto")
- docker-compose.yml includes commented-out NVIDIA GPU deploy section
- Supports onnxruntime-gpu as a drop-in replacement for onnxruntime
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace OpenCLIP ViT-B/32 (512-d, ~78% recall) with SigLIP2 ViT-B/16
(768-d, ~84% recall) as the default embedding model for significantly
better image-text retrieval quality.
- New SigLIP2Embedder class with 384px input and SigLIP normalization
- ONNX export pipeline for SigLIP2 visual + textual encoders
- Migration 0010: resize embeddings.vector from 512 to 768 dimensions
- Config-driven model selection: "siglip2_vitb16" (default) or
"openclip_vitb32" (legacy) — both models can coexist
- Content classifier follows the configured embedder family
- Existing embeddings cleared on migration; vision backfill regenerates
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add export_models.py for OpenCLIP ViT-B/32 and YOLOv8n ONNX export
- Fix ArgMax(13) ORT ARM64 incompatibility by passing eot_indices as a
separate ONNX input (computed outside the graph in embed.py)
- Use legacy TorchScript exporter (dynamo=False) for IR version 9 compat
- Upgrade onnxruntime to 1.18.1
- Rewrite bootstrap_models.py with clear separation of auto-downloadable
models (YuNet, SFace) vs manually-exported ones (OpenCLIP, YOLOv8n)
- Wire bootstrap into worker CMD (runs before Celery)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>