feat: split celery workers, fix asyncpg-in-fork, add pipeline progress UI
Three overlapping fixes so the ingestion pipeline actually runs and the
user can see what it's doing:
Pipeline recovery
- app/database.py: use NullPool when MULITA_CELERY_WORKER=1 so each
Celery task opens a fresh asyncpg connection on its own event loop.
Fixes "another operation in progress" and "Future attached to a
different loop" errors that were dropping ~every thumbnail +
extract_metadata task on the floor.
- app/tasks/thumbs.py: initialize photo=None before the try and rollback
on error so a transport failure in the initial SELECT doesn't raise
UnboundLocalError in the except block and leak rows stuck in 'pending'.
- app/services/vision/bootstrap_models.py: on missing model files,
invoke export_models automatically instead of just warning. First
boot of a fresh install now self-heals.
- app/services/vision/export_models.py: shutil.move instead of
Path.rename so the YOLO export survives the /app → /data/models
cross-volume hop.
- requirements.txt: add ultralytics so export works in a stock image.
Worker topology
- docker-compose.yml: replace the single worker with worker-light
(default/high/low queues, c=2, IO-bound) and worker-vision (vision
queue, c=5, OMP_NUM_THREADS=1 to avoid oversubscription on 6 cores).
Vision is pinned to ≤5 parallel inferences so ONNX doesn't each
spawn an all-cores intra-op pool.
- .env / .env.example: CELERYD_CONCURRENCY replaced with
CELERY_LIGHT_CONCURRENCY + CELERY_VISION_CONCURRENCY.
- Backfill queries in thumbs / scan / vision now ORDER BY taken_at
DESC NULLS LAST so newest photos finish first — the library fills
in top-down in the UI instead of arbitrary insertion order.
Settings visibility
- routers/library.py: new GET /maintenance/pipeline-stats returning
done/total per stage (thumbnails, exif, gps, phash, embeddings,
tags, ocr, faces, face clusters, duplicate groups). Worker-status
now also reports the `vision` queue depth, which was missing.
- services/api.ts: PipelineStats / PipelineStage / ScanStatus types
and the matching client call.
- components/dialogs/SettingsDialog.tsx:
- new Pipeline Progress card with one progress bar per stage
- inline scan banner (processed/total/current folder) inside the
Library section while a scan is running
- Tasks/min throughput computed by diffing worker processed counters
between polls
- Workers section calls out the vision queue and documents the
CELERY_LIGHT/VISION_CONCURRENCY + docker compose up -d scale path
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -15,8 +15,10 @@ SQLite (escape hatch via docker-compose.sqlite.yml): no Alembic. The
|
||||
historical inline ALTER TABLE block stays in place so existing dev
|
||||
installs keep upgrading.
|
||||
"""
|
||||
import os
|
||||
from sqlalchemy.ext.asyncio import AsyncSession, create_async_engine, async_sessionmaker
|
||||
from sqlalchemy.orm import declarative_base
|
||||
from sqlalchemy.pool import NullPool
|
||||
from sqlalchemy import text
|
||||
import logging
|
||||
from pathlib import Path
|
||||
@@ -28,6 +30,27 @@ logger = logging.getLogger(__name__)
|
||||
_is_sqlite = settings.database_url.startswith("sqlite")
|
||||
_is_postgres = settings.database_url.startswith("postgresql")
|
||||
|
||||
# When running inside a Celery worker we use NullPool rather than the
|
||||
# default connection pool. The reasons stack up:
|
||||
#
|
||||
# 1. Celery's prefork model forks the master *after* imports, so every
|
||||
# child inherits the same asyncpg Connection objects — they share
|
||||
# a socket, and two children using one concurrently raises
|
||||
# "another operation is in progress".
|
||||
#
|
||||
# 2. Task bodies run under `asyncio.run()`, which spins up a fresh
|
||||
# event loop per invocation. A pooled asyncpg Connection created
|
||||
# on loop A, returned to the pool, and checked out on loop B
|
||||
# raises "Future attached to a different loop".
|
||||
#
|
||||
# NullPool dodges both: every session checkout opens a brand-new
|
||||
# connection on the *current* loop and the connection is closed at
|
||||
# session end. Connection setup is cheap compared to task cost, so this
|
||||
# is the right default for the worker. The FastAPI backend keeps the
|
||||
# normal pool because it serves many short requests on a single long-
|
||||
# lived event loop, where pooling is a clear win.
|
||||
_is_celery_worker = os.environ.get("MULITA_CELERY_WORKER") == "1"
|
||||
|
||||
if _is_sqlite:
|
||||
db_path = Path(settings.database_url.replace("sqlite+aiosqlite:///", ""))
|
||||
db_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
@@ -39,6 +62,12 @@ if _is_sqlite:
|
||||
"timeout": 30,
|
||||
},
|
||||
)
|
||||
elif _is_celery_worker:
|
||||
engine = create_async_engine(
|
||||
settings.database_url,
|
||||
echo=False,
|
||||
poolclass=NullPool,
|
||||
)
|
||||
else:
|
||||
engine = create_async_engine(
|
||||
settings.database_url,
|
||||
|
||||
Reference in New Issue
Block a user