Files
mule-image/backend/app/services/duplicates.py
root 339e1be510 feat: hide-from-views flag on folders
Adds a per-folder "hide from views" toggle so noisy subtrees
(screenshots, WhatsApp dumps, work archives) can be excluded from
cross-cutting views without losing indexing. Photos under a hidden
folder are still scanned, thumbnailed, embedded, OCR'd, face-
extracted — they just stop appearing in All Photos, Rated, Colors,
Map, Tags, People, Search, Duplicates, and the sidebar counts.
Navigating directly into the folder still shows every photo.

Schema (migration 0007_folder_hidden):
- folders.is_hidden   user-set toggle, default false
- photos.is_hidden    denormalized effective flag (true iff any
                      ancestor folder is hidden), indexed so cross-
                      cutting queries stay on the existing planner
                      paths

The denorm is maintained by two paths:
- The scanner walks the ancestry chain on insert, with a per-scan
  memoized cache so each folder is resolved once per scan.
- POST /api/v1/folders/{id}/hide flips folders.is_hidden and runs a
  WITH RECURSIVE CTE to recompute every folder's effective state in
  one query, then bulk-updates photos WHERE IS DISTINCT FROM. Runs
  in ~10 ms on a 13k-photo library.

Filters added (cross-cutting queries):
- /library/stats — every sidebar badge via a shared `visible` filter
- /photos (list) — only when neither folder_id nor heap_id is set;
  folder browse and heap browse always show everything
- /photos/map
- /library/duplicates/groups
- /folders/tree photo_count subquery
- /tags count_subq (drives Tags + People sidebar counts)
- services/duplicates.regroup_duplicates (so hidden dupes never
  contaminate the Duplicates view)
- services/search.hybrid_search — both semantic (pgvector) and FTS
  legs join photos so rankings don't include hidden results

Intentionally NOT filtered:
- /photos?folder_id=X and /photos?heap_id=X (user-intentional browse)
- /library/maintenance/pipeline-stats (tracks real worker state)
- cleanup service (disk-level ops, not views)

Frontend:
- sourceFolders.setHidden(id, hidden) API client method
- FolderTreeNode.is_hidden carried through the tree into TreeItem
- LeftSidebar kebab menu: "Hide from views" / "Show in views" with a
  mutation that invalidates folders, photos, stats, and tags caches
- Hidden folder rows swap the Folder icon for EyeOff and render the
  label italic/muted so the state is visible at a glance

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 10:55:25 +02:00

230 lines
8.1 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""
Duplicate detection: group photos by perceptual-hash similarity.
Strategy
--------
Each photo carries a 16-char hex perceptual hash (`Photo.phash`) computed
by the thumbnail worker from the original-resolution decoded frame
(`app.tasks.thumbs._generate_thumbnails_async`). pHash is robust to
resize / re-encoding / mild edits, so two photos that are visually the
"same shot" land at small Hamming distance even when their bytes are
completely different.
This module turns those per-photo hashes into explicit *groups*. The
result is persisted in two columns:
* `Photo.duplicate_group_id` — shared by every member of a group
* `Photo.is_duplicate` — derived: True iff group_id IS NOT NULL
(kept as a column so the existing
PhotoThumbnail badge and /library/stats
count don't have to change).
The grouping is recomputed in batches by `regroup_duplicates`, NOT on
individual writes — that lets us use a single in-memory pass instead of
maintaining a per-row similarity index. Triggered automatically after
each scan and on demand from the Settings panel.
Complexity
----------
Pairwise O(N²) over photos with a non-null phash. At ~5 µs per Hamming
distance in CPython this is roughly:
1k photos → ~5 s
5k photos → ~125 s
10k photos → ~500 s
That's the wrong shape for libraries past ~5k. The drop-in replacement
is a BK-tree (e.g. `pybktree`) which gives O(log N) lookups for a fixed
Hamming threshold; swap it in here when someone trips the limit. The
public function signature stays the same.
Out of scope (deferred)
-----------------------
* Dismissing a group / "intentional duplicates" — would need a per-group
or per-pair flag plus a skip-set in this function so re-grouping
doesn't bring them back. Add when there's a real user need.
* Incremental updates on individual photo writes — currently we just
re-run the whole job after each scan, which is fine while the cost is
bounded.
"""
from __future__ import annotations
import logging
import uuid
from typing import Optional
from sqlalchemy import select, update
from app.database import AsyncSessionLocal
from app.models.photos import Photo
logger = logging.getLogger(__name__)
# Hamming distance threshold under which two phashes are considered
# "the same image". 6 bits out of 64 is the rule-of-thumb sweet spot for
# pHash — tight enough to avoid false positives between unrelated photos,
# loose enough to catch JPEG re-encodes, slight crops, and a screenshot
# of a screenshot.
DEFAULT_THRESHOLD = 6
def _hex_to_int(h: str) -> int:
"""Parse a 16-char hex pHash to a Python int. Returns -1 on bad input
so the pairwise loop can skip the row without raising."""
try:
return int(h, 16)
except (TypeError, ValueError):
return -1
def _hamming(a: int, b: int) -> int:
"""Population count of XOR — the canonical hash distance metric.
`int.bit_count()` is C-implemented in CPython 3.10+ and is by far
the fastest path; the `bin(...).count('1')` fallback is here only
so the function still works on older interpreters."""
x = a ^ b
try:
return x.bit_count() # type: ignore[attr-defined]
except AttributeError:
return bin(x).count('1')
class _UnionFind:
"""Tiny union-find / disjoint-set used to merge similar phashes into
connected components. Inlined here (rather than pulled from a dep)
because it's ~15 lines and we don't need anything fancy."""
def __init__(self, n: int) -> None:
self.parent = list(range(n))
self.rank = [0] * n
def find(self, x: int) -> int:
# Path compression — flattens the tree on lookup so subsequent
# finds are amortized O(α(N)) ≈ O(1).
while self.parent[x] != x:
self.parent[x] = self.parent[self.parent[x]]
x = self.parent[x]
return x
def union(self, a: int, b: int) -> None:
ra, rb = self.find(a), self.find(b)
if ra == rb:
return
if self.rank[ra] < self.rank[rb]:
ra, rb = rb, ra
self.parent[rb] = ra
if self.rank[ra] == self.rank[rb]:
self.rank[ra] += 1
async def regroup_duplicates(threshold: int = DEFAULT_THRESHOLD) -> dict:
"""Recompute every photo's duplicate_group_id from current phashes.
Idempotent — safe to call as often as you like. Returns a small
summary dict the maintenance endpoint surfaces back to the UI.
Photos that end up alone in a component (size 1) get
`duplicate_group_id=NULL` and `is_duplicate=False`. This is what
cleans up "dead" groups after the user discards N-1 members from
one.
"""
async with AsyncSessionLocal() as session:
# Pull (id, phash) for every visible photo with a hash. Discarded
# and hidden photos are excluded so we don't keep showing groups
# made up of trashed copies or members of folders the user
# deliberately excluded from cross-cutting views.
rows = (
await session.execute(
select(Photo.id, Photo.phash)
.where(Photo.phash.is_not(None))
.where(Photo.is_discarded.is_(False))
.where(Photo.is_hidden.is_(False))
)
).all()
n = len(rows)
if n == 0:
# Still need to clear stale group_ids in case the user just
# discarded the last surviving member of every group.
await _clear_all_groups(session)
await session.commit()
return {
'photos_considered': 0,
'groups': 0,
'members': 0,
}
ids: list[str] = [row[0] for row in rows]
hashes: list[int] = [_hex_to_int(row[1]) for row in rows]
uf = _UnionFind(n)
# O(N²) pairwise comparison. See module docstring for the
# scaling analysis and the BK-tree upgrade path.
for i in range(n):
hi = hashes[i]
if hi < 0:
continue
for j in range(i + 1, n):
hj = hashes[j]
if hj < 0:
continue
if _hamming(hi, hj) <= threshold:
uf.union(i, j)
# Collect components. Each connected component of size >= 2 gets
# a fresh group id; size-1 components are intentionally dropped.
components: dict[int, list[int]] = {}
for i in range(n):
root = uf.find(i)
components.setdefault(root, []).append(i)
# First pass: clear EVERY photo's group_id so survivors of an
# earlier grouping that no longer match anyone end up clean. This
# is one bulk UPDATE rather than per-photo to keep the cost low
# even on big libraries.
await _clear_all_groups(session)
# Second pass: write the new group ids for components of size 2+.
groups_created = 0
members_total = 0
for members in components.values():
if len(members) < 2:
continue
group_id = str(uuid.uuid4())
member_ids = [ids[i] for i in members]
await session.execute(
update(Photo)
.where(Photo.id.in_(member_ids))
.values(
duplicate_group_id=group_id,
is_duplicate=True,
)
)
groups_created += 1
members_total += len(member_ids)
await session.commit()
logger.info(
f"regroup_duplicates: considered {n} photos, "
f"created {groups_created} group(s) covering {members_total} member(s)"
)
return {
'photos_considered': n,
'groups': groups_created,
'members': members_total,
}
async def _clear_all_groups(session) -> None:
"""Reset duplicate_group_id / is_duplicate on every photo. Used as
the first half of a regroup pass so photos that no longer cluster
with anyone end up clean instead of carrying a stale group id."""
await session.execute(
update(Photo).values(duplicate_group_id=None, is_duplicate=False)
)