library: scraper options — per-source + per-field auto-apply, review-queue order (R1) (#726)

* library: scraper options — per-source + per-field auto-apply toggles, review-queue order (R1)

Grows the Settings→Library "Metadata matching" card into the full
scraper-options panel — not everybody needs the same things out of a
scraper:

- Sources: enrich_src_musicbrainz gates the background matcher (phase 2;
  identity hashes still stamp, and manual Fix-match/search stays
  available — same contract as the master toggle). enrich_src_caa gates
  the Cover Art Archive fetch (phase 3). Rows skipped by an off toggle
  stay unevaluated, so re-enabling picks them up on the next pass —
  nothing is permanently forfeited.
- Auto-apply fields: enrich_apply_names/year/genres filter what an
  AUTOMATIC match may canonicalize (_enrich_field_filter, applied on all
  three automatic paths: cache copy, mbid/isrc exact keys, text auto).
  MusicBrainz ids always stamp — they're identity, not display; the art
  fetch and future re-matching need them. A match the user confirms in
  the review modal applies in full. enrich_apply_art gates the art fetch
  alongside the CAA source toggle (two axes, one behaviour today —
  future art sources slot in without re-teaching the panel).
- Review queue order: enrich_review_order = missing_first (default,
  today's behaviour) | artist | recent, read by GET /api/enrichment/review;
  unknown stored values degrade to the default.
- Settings card: Sources / Auto-apply / Review-queue-order groups wired
  in match-review.js; the master toggle is relabelled "Match songs
  automatically" so it doesn't read the same as the new MusicBrainz
  source toggle. No tailwind rebuild needed — every class was already
  scanned from core source.

Tests: tests/test_scraper_options.py (9) — settings validation,
MB-source-off stamps-without-matching + re-enable, per-field stripping
on auto matches with ids preserved, review-accept full-apply despite
toggles, CAA gating on both axes, review-order modes incl. the
unknown-value fallback. Full-suite failure set A/B-identical to the
base (39 env/pre-existing).

Stacked on feat/enrichment-art (#715) — merge that first.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nm7tHs1Yvjjtnnu4nzJgdN

* library: per-field auto-apply honours "nothing forfeited" (backfill + no partial seeding)

The R1 per-FIELD auto-apply toggles settled a `matched` row with the
disabled fields stripped, but enrichment_pending() never revisits an
unchanged-hash matched row — so re-enabling a field never backfilled it,
and enrichment_cache_lookup() (which gates only on mb_recording_id) could
seed sibling charts with the stripped blanks. This broke the same
"nothing is permanently forfeited" contract the source (mb_on) and art
(art_on) toggles already keep.

Fix: persist an `apply_mask` marker (sorted blocked apply-keys) on every
AUTOMATIC match:
- migration: additive `apply_mask TEXT` column (idempotent ALTER).
- enrichment_pending(allowed_keys=...): re-queues a `matched` row whose
  apply_mask names a field that is now re-enabled → backfill on re-enable,
  converges (a fully-applied row is never re-queued).
- enrichment_cache_lookup: only fully-applied donors (apply_mask empty/NULL)
  may seed siblings; a partial row is skipped and the sibling falls through
  to its own re-filtered match.
- _enrich_apply_mask()/_enrich_blocked_apply_keys() helpers; threaded through
  _enrich_one → apply_enrichment_match. Review/manual writers leave it NULL
  (a confirmed pick applies in full).

Tests: re-enable-backfills-and-converges; partial row is not a cache donor
(fully-applied one is). 13 scraper-options tests pass; 170 enrichment/
settings tests green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: ChrisBeWithYou <christian.a.cowan@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Byron Gamatos
2026-07-02 20:57:42 +02:00
committed by GitHub
co-authored by Claude Opus 4.8 ChrisBeWithYou
parent 7ca736d525
commit f5c9c34291
4 changed files with 579 additions and 49 deletions
+194 -44
View File
@@ -893,9 +893,18 @@ class MetadataDB:
# MusicBrainz just to render; `last_attempt_at` anchors the failed-row # MusicBrainz just to render; `last_attempt_at` anchors the failed-row
# retry backoff (epoch seconds). Idempotent ALTERs, same pattern as # retry backoff (epoch seconds). Idempotent ALTERs, same pattern as
# the `songs` migrations above. # the `songs` migrations above.
# R1 scraper options: `apply_mask` records which per-field auto-apply
# toggles were OFF (suppressed) when an AUTOMATIC match settled the row,
# as a canonical sorted comma-joined marker of blocked keys (''/NULL =
# nothing suppressed). It keeps the per-field toggles to the same
# "nothing forfeited" contract as the source/art toggles: re-enabling a
# field re-queues affected `matched` rows for backfill (enrichment_pending)
# and a partially-applied row is barred from seeding siblings
# (enrichment_cache_lookup). Idempotent ALTER, same pattern as above.
for ddl in ( for ddl in (
"ALTER TABLE song_enrichment ADD COLUMN candidates TEXT", "ALTER TABLE song_enrichment ADD COLUMN candidates TEXT",
"ALTER TABLE song_enrichment ADD COLUMN last_attempt_at REAL", "ALTER TABLE song_enrichment ADD COLUMN last_attempt_at REAL",
"ALTER TABLE song_enrichment ADD COLUMN apply_mask TEXT",
): ):
try: try:
self.conn.execute(ddl) self.conn.execute(ddl)
@@ -2642,7 +2651,8 @@ class MetadataDB:
raw = "|".join([norm(artist), norm(title), norm(album), dur]) raw = "|".join([norm(artist), norm(title), norm(album), dur])
return hashlib.sha1(raw.encode("utf-8")).hexdigest() return hashlib.sha1(raw.encode("utf-8")).hexdigest()
def enrichment_pending(self, limit: int = 500) -> list[dict]: def enrichment_pending(self, limit: int = 500,
allowed_keys: frozenset | None = None) -> list[dict]:
"""Songs whose enrichment row needs (re)matching: no row yet, or a """Songs whose enrichment row needs (re)matching: no row yet, or a
row whose content_hash no longer matches the song's current metadata row whose content_hash no longer matches the song's current metadata
(an edit changed the identity re-match), or an `unscanned` row. (an edit changed the identity re-match), or an `unscanned` row.
@@ -2652,24 +2662,37 @@ class MetadataDB:
retries only via the matcher's backoff policy (enrichment_failed_rows) retries only via the matcher's backoff policy (enrichment_failed_rows)
rather than being re-queued every pass. An identity edit (say, the rather than being re-queued every pass. An identity edit (say, the
user fixes the typo that made matching fail) re-queues any of them user fixes the typo that made matching fail) re-queues any of them
immediately via the hash mismatch.""" immediately via the hash mismatch.
`allowed_keys` is the set of per-field auto-apply toggle keys that are
currently ON. A `matched` row stamped while one of those fields was
suppressed (its key in `apply_mask`) is re-queued for backfill, so
re-enabling a field honours the same "nothing forfeited" contract the
source/art toggles already keep. None = don't apply the mask rule (the
caller isn't the field-aware matcher, e.g. a plain count)."""
# Read under _lock: the worker commits on this shared connection under # Read under _lock: the worker commits on this shared connection under
# _lock, so an unlocked SELECT could interleave with its execute+commit. # _lock, so an unlocked SELECT could interleave with its execute+commit.
with self._lock: with self._lock:
rows = self.conn.execute( rows = self.conn.execute(
"SELECT s.filename, s.artist, s.title, s.album, s.year, s.duration, " "SELECT s.filename, s.artist, s.title, s.album, s.year, s.duration, "
"e.content_hash, e.match_state " "e.content_hash, e.match_state, e.apply_mask "
"FROM songs s LEFT JOIN song_enrichment e ON e.filename = s.filename " "FROM songs s LEFT JOIN song_enrichment e ON e.filename = s.filename "
"WHERE s.title != '' AND (e.filename IS NULL " "WHERE s.title != '' AND (e.filename IS NULL "
"OR e.match_state IN ('unscanned', 'matched', 'review', 'failed')) " "OR e.match_state IN ('unscanned', 'matched', 'review', 'failed')) "
"ORDER BY s.filename LIMIT ?", (max(1, int(limit)),)).fetchall() "ORDER BY s.filename LIMIT ?", (max(1, int(limit)),)).fetchall()
out = [] out = []
for fn, artist, title, album, year, duration, ehash, state in rows: for fn, artist, title, album, year, duration, ehash, state, mask in rows:
h = self.enrichment_content_hash(artist, title, album, duration) h = self.enrichment_content_hash(artist, title, album, duration)
# No row yet, still unmatched, or the identity changed under a # No row yet, still unmatched, or the identity changed under a
# settled row → needs the matcher. A settled row with an # settled row → needs the matcher. A settled row with an
# unchanged hash stays settled (idempotence). # unchanged hash stays settled (idempotence)
if state is None or state == "unscanned" or ehash != h: needs = state is None or state == "unscanned" or ehash != h
# …EXCEPT a `matched` row that suppressed a field now re-enabled:
# re-queue it so the newly-allowed field gets backfilled.
if not needs and state == "matched" and allowed_keys is not None and mask:
if {k for k in mask.split(",") if k} & allowed_keys:
needs = True
if needs:
out.append({"filename": fn, "artist": artist, "title": title, out.append({"filename": fn, "artist": artist, "title": title,
"album": album, "year": year, "duration": duration, "album": album, "year": year, "duration": duration,
"content_hash": h, "match_state": state}) "content_hash": h, "match_state": state})
@@ -2724,7 +2747,8 @@ class MetadataDB:
"SELECT filename, content_hash, match_state, match_source, match_score, attempts, " "SELECT filename, content_hash, match_state, match_source, match_score, attempts, "
"mb_recording_id, mb_release_id, mb_artist_id, isrc, " "mb_recording_id, mb_release_id, mb_artist_id, isrc, "
"canon_artist, canon_album, canon_title, canon_year, canon_artist_sort, " "canon_artist, canon_album, canon_title, canon_year, canon_artist_sort, "
"genres, art_cache_path, art_state, fetched_at, candidates, last_attempt_at " "genres, art_cache_path, art_state, fetched_at, candidates, last_attempt_at, "
"apply_mask "
"FROM song_enrichment WHERE filename = ?", (filename,)).fetchone() "FROM song_enrichment WHERE filename = ?", (filename,)).fetchone()
if not row: if not row:
return None return None
@@ -2732,7 +2756,7 @@ class MetadataDB:
"attempts", "mb_recording_id", "mb_release_id", "mb_artist_id", "isrc", "attempts", "mb_recording_id", "mb_release_id", "mb_artist_id", "isrc",
"canon_artist", "canon_album", "canon_title", "canon_year", "canon_artist", "canon_album", "canon_title", "canon_year",
"canon_artist_sort", "genres", "art_cache_path", "art_state", "fetched_at", "canon_artist_sort", "genres", "art_cache_path", "art_state", "fetched_at",
"candidates", "last_attempt_at") "candidates", "last_attempt_at", "apply_mask")
out = dict(zip(keys, row)) out = dict(zip(keys, row))
for k in ("genres", "candidates"): for k in ("genres", "candidates"):
try: try:
@@ -2784,12 +2808,17 @@ class MetadataDB:
def enrichment_cache_lookup(self, content_hash: str, exclude_filename: str = "") -> dict | None: def enrichment_cache_lookup(self, content_hash: str, exclude_filename: str = "") -> dict | None:
"""A settled match for the same identity hash — another chart of the """A settled match for the same identity hash — another chart of the
same recording already matched/pinned copy it, no network (design same recording already matched/pinned copy it, no network (design
§5 step 1: the local match-cache).""" §5 step 1: the local match-cache). Only FULLY-applied donors qualify
(apply_mask empty/NULL): a row that suppressed a display field under an
auto-apply toggle would otherwise seed siblings with its blanks even
when the reader's own toggles want that field — so a partial row is
skipped and the sibling falls through to its own (re-filtered) match."""
row = self.conn.execute( row = self.conn.execute(
"SELECT match_score, mb_recording_id, mb_release_id, mb_artist_id, isrc, " "SELECT match_score, mb_recording_id, mb_release_id, mb_artist_id, isrc, "
"canon_artist, canon_album, canon_title, canon_year, canon_artist_sort, genres " "canon_artist, canon_album, canon_title, canon_year, canon_artist_sort, genres "
"FROM song_enrichment WHERE content_hash = ? AND filename != ? " "FROM song_enrichment WHERE content_hash = ? AND filename != ? "
"AND match_state IN ('matched', 'manual') AND mb_recording_id IS NOT NULL " "AND match_state IN ('matched', 'manual') AND mb_recording_id IS NOT NULL "
"AND COALESCE(apply_mask, '') = '' "
"LIMIT 1", (content_hash, exclude_filename or "")).fetchone() "LIMIT 1", (content_hash, exclude_filename or "")).fetchone()
if not row: if not row:
return None return None
@@ -2809,7 +2838,8 @@ class MetadataDB:
source: str | None = None, score: float | None = None, source: str | None = None, score: float | None = None,
cand: dict | None = None, candidates: list | None = None, cand: dict | None = None, candidates: list | None = None,
bump_attempts: bool = False, bump_attempts: bool = False,
allow_manual_overwrite: bool = False) -> bool: allow_manual_overwrite: bool = False,
apply_mask: str | None = None) -> bool:
"""The single writer for every matcher/review outcome. Writes the """The single writer for every matcher/review outcome. Writes the
full lifecycle row: state + source + score, the canonical fields a full lifecycle row: state + source + score, the canonical fields a
confident match supplies (`cand`), and/or the review tier's ranked confident match supplies (`cand`), and/or the review tier's ranked
@@ -2817,7 +2847,11 @@ class MetadataDB:
`manual` and the caller isn't explicitly acting for the user — the `manual` and the caller isn't explicitly acting for the user — the
never-overwrite-manual contract lives HERE so no future call path never-overwrite-manual contract lives HERE so no future call path
can forget it. Art-cache fields are preserved verbatim (they belong can forget it. Art-cache fields are preserved verbatim (they belong
to the art slice, not the matcher).""" to the art slice, not the matcher). `apply_mask` (blocked per-field
keys, from the matcher) is stamped verbatim so enrichment_pending /
enrichment_cache_lookup can tell a fully-applied match from a
field-suppressed one; the review/manual writers leave it NULL (a
confirmed pick applies in full)."""
cand = cand or {} cand = cand or {}
now = time.time() now = time.time()
with self._lock: with self._lock:
@@ -2840,8 +2874,9 @@ class MetadataDB:
"match_state, match_source, match_score, attempts, " "match_state, match_source, match_score, attempts, "
"mb_recording_id, mb_release_id, mb_artist_id, isrc, " "mb_recording_id, mb_release_id, mb_artist_id, isrc, "
"canon_artist, canon_album, canon_title, canon_year, canon_artist_sort, " "canon_artist, canon_album, canon_title, canon_year, canon_artist_sort, "
"genres, art_cache_path, art_state, fetched_at, candidates, last_attempt_at) " "genres, art_cache_path, art_state, fetched_at, candidates, last_attempt_at, "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)", "apply_mask) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(filename, content_hash, state, source, score, attempts, (filename, content_hash, state, source, score, attempts,
cand.get("recording_id") or None, cand.get("release_id") or None, cand.get("recording_id") or None, cand.get("release_id") or None,
cand.get("artist_id") or None, cand.get("isrc") or None, cand.get("artist_id") or None, cand.get("isrc") or None,
@@ -2852,7 +2887,8 @@ class MetadataDB:
cur[2] if cur else None, cur[3] if cur else None, cur[2] if cur else None, cur[3] if cur else None,
fetched_at, fetched_at,
json.dumps(candidates) if candidates else None, json.dumps(candidates) if candidates else None,
now if state == "failed" else None)) now if state == "failed" else None,
apply_mask or None))
self.conn.commit() self.conn.commit()
return True return True
@@ -2882,19 +2918,26 @@ class MetadataDB:
filename, row["content_hash"], "failed", source="rejected", filename, row["content_hash"], "failed", source="rejected",
score=None, candidates=row.get("candidates") or None) score=None, candidates=row.get("candidates") or None)
def enrichment_review_queue(self, limit: int = 200) -> list[dict]: def enrichment_review_queue(self, limit: int = 200,
order: str = "missing_first") -> list[dict]:
"""The Match-Review drawer's queue: review-tier rows joined to their """The Match-Review drawer's queue: review-tier rows joined to their
(still-existing) songs, with the stored candidate list parsed.""" (still-existing) songs, with the stored candidate list parsed.
`order` is the user's review-queue preference: 'missing_first'
(default charts missing album/year surface first, they gain the
most from a confirm; complete charts only stand to be re-labelled),
'artist' (AZ), or 'recent' (newest files first). Unknown values
fall back to missing_first."""
order_sql = {
"artist": "s.artist COLLATE NOCASE, s.title COLLATE NOCASE, e.filename",
"recent": "s.mtime DESC, e.filename",
}.get(order, "((COALESCE(s.album, '') = '') + (COALESCE(s.year, '') = '')) DESC, "
"s.artist COLLATE NOCASE, s.title COLLATE NOCASE, e.filename")
rows = self.conn.execute( rows = self.conn.execute(
"SELECT e.filename, s.title, s.artist, s.album, s.year, s.duration, s.mtime, " "SELECT e.filename, s.title, s.artist, s.album, s.year, s.duration, s.mtime, "
"e.match_score, e.candidates, e.attempts " "e.match_score, e.candidates, e.attempts "
"FROM song_enrichment e JOIN songs s ON s.filename = e.filename " "FROM song_enrichment e JOIN songs s ON s.filename = e.filename "
"WHERE e.match_state = 'review' " "WHERE e.match_state = 'review' "
# Charts that are MISSING data (no album / no year) surface first — "ORDER BY " + order_sql + " "
# confirming those has the most to gain; complete charts only
# stand to be re-labelled.
"ORDER BY ((COALESCE(s.album, '') = '') + (COALESCE(s.year, '') = '')) DESC, "
"s.artist COLLATE NOCASE, s.title COLLATE NOCASE, e.filename "
"LIMIT ?", (max(1, int(limit)),)).fetchall() "LIMIT ?", (max(1, int(limit)),)).fetchall()
out = [] out = []
for fn, title, artist, album, year, duration, mtime, score, cands, attempts in rows: for fn, title, artist, album, year, duration, mtime, score, cands, attempts in rows:
@@ -5932,7 +5975,48 @@ def _enrich_art_one(row: dict) -> bool:
return True return True
def _enrich_one(row: dict, auto_min: float | None = None) -> None: _ENRICH_APPLY_FIELDS = {
# Per-field auto-apply toggle → the candidate fields it governs. The
# MusicBrainz ids + isrc are deliberately NOT here: they're identity,
# not display — the art fetch and any future re-match need them stamped
# even when every display field is toggled off.
"enrich_apply_names": ("artist", "title", "album", "artist_sort"),
"enrich_apply_year": ("year",),
"enrich_apply_genres": ("genres",),
}
def _enrich_blocked_apply_keys(cfg: dict) -> frozenset:
"""The per-field auto-apply toggle keys that are currently OFF (suppressed).
Its complement (`_ENRICH_APPLY_FIELDS` minus these) is what an automatic
match may canonicalize."""
return frozenset(k for k in _ENRICH_APPLY_FIELDS if cfg.get(k, True) is False)
def _enrich_apply_mask(cfg: dict) -> str:
"""Canonical marker of the suppressed apply keys, persisted on each
automatic match so re-enabling a field re-queues the row for backfill
(enrichment_pending) and a partial match can't seed siblings
(enrichment_cache_lookup). '' = nothing suppressed (the default)."""
return ",".join(sorted(_enrich_blocked_apply_keys(cfg)))
def _enrich_field_filter(cfg: dict):
"""Build the cand filter for AUTOMATIC matches from the per-field
auto-apply settings: strips the display fields whose toggle is off
before they're stamped as canonical. Returns None when everything is
on (the default) so the common path stays zero-copy. Review candidates
and user-confirmed picks bypass this a match the user confirms in
the modal applies in full."""
blocked = {f for key in _enrich_blocked_apply_keys(cfg)
for f in _ENRICH_APPLY_FIELDS[key]}
if not blocked:
return None
return lambda cand: {k: v for k, v in cand.items() if k not in blocked}
def _enrich_one(row: dict, auto_min: float | None = None, field_filter=None,
apply_mask: str = "") -> None:
"""The matcher (P8; replaces P7's no-op). Precedence per design §5: """The matcher (P8; replaces P7's no-op). Precedence per design §5:
1. local match-cache by content_hash another chart of the same 1. local match-cache by content_hash another chart of the same
@@ -5945,17 +6029,25 @@ def _enrich_one(row: dict, auto_min: float | None = None) -> None:
`auto_min` is the user's auto-apply confidence setting (None → the `auto_min` is the user's auto-apply confidence setting (None → the
engine default); it moves only the auto/review boundary of step 3 engine default); it moves only the auto/review boundary of step 3
the per-field floors and exact-key tiers are unaffected. Never touches the per-field floors and exact-key tiers are unaffected. `field_filter`
a `manual` row (the writer enforces it). Network errors raise (from _enrich_field_filter) strips per-field-disabled display values
EnrichTransportError so the pass pauses instead of burning attempts from every AUTOMATIC stamp all three steps here are automatic, so it
while offline.""" applies to each; the review tier stores candidates unfiltered because
accepting one is a user action. Never touches a `manual` row (the
writer enforces it). `apply_mask` (the suppressed keys, from
_enrich_apply_mask) is stamped on each AUTOMATIC match so a later
re-enable re-queues the row for backfill and a partial match can't seed
siblings. Network errors raise EnrichTransportError so the pass pauses
instead of burning attempts while offline."""
fn, chash = row["filename"], row["content_hash"] fn, chash = row["filename"], row["content_hash"]
cached = meta_db.enrichment_cache_lookup(chash, exclude_filename=fn) cached = meta_db.enrichment_cache_lookup(chash, exclude_filename=fn)
if cached: if cached:
score = cached.pop("score", None) score = cached.pop("score", None)
if field_filter:
cached = field_filter(cached)
meta_db.apply_enrichment_match(fn, chash, "matched", source="cache", meta_db.apply_enrichment_match(fn, chash, "matched", source="cache",
score=score, cand=cached) score=score, cand=cached, apply_mask=apply_mask)
return return
ids = _manifest_exact_ids(fn) ids = _manifest_exact_ids(fn)
@@ -5963,14 +6055,16 @@ def _enrich_one(row: dict, auto_min: float | None = None) -> None:
cand = _mb_lookup_recording(ids["mbid"]) cand = _mb_lookup_recording(ids["mbid"])
if cand: if cand:
meta_db.apply_enrichment_match(fn, chash, "matched", source="mbid", meta_db.apply_enrichment_match(fn, chash, "matched", source="mbid",
score=1.0, cand=cand) score=1.0, apply_mask=apply_mask,
cand=field_filter(cand) if field_filter else cand)
return return
# A 404'd mbid (typo'd manifest) falls through to the text tiers. # A 404'd mbid (typo'd manifest) falls through to the text tiers.
if ids.get("isrc"): if ids.get("isrc"):
cands = mb_match.rank_candidates(row, _mb_lookup_isrc(ids["isrc"])) cands = mb_match.rank_candidates(row, _mb_lookup_isrc(ids["isrc"]))
if cands: if cands:
meta_db.apply_enrichment_match(fn, chash, "matched", source="isrc", meta_db.apply_enrichment_match(fn, chash, "matched", source="isrc",
score=1.0, cand=cands[0]) score=1.0, apply_mask=apply_mask,
cand=field_filter(cands[0]) if field_filter else cands[0])
return return
ranked = mb_match.rank_candidates(row, _mb_search_recordings(row.get("artist"), row.get("title"))) ranked = mb_match.rank_candidates(row, _mb_search_recordings(row.get("artist"), row.get("title")))
@@ -5978,7 +6072,8 @@ def _enrich_one(row: dict, auto_min: float | None = None) -> None:
tier = mb_match.classify(row, best, best["score"], auto_min=auto_min) if best else "none" tier = mb_match.classify(row, best, best["score"], auto_min=auto_min) if best else "none"
if tier == "auto": if tier == "auto":
meta_db.apply_enrichment_match(fn, chash, "matched", source="text", meta_db.apply_enrichment_match(fn, chash, "matched", source="text",
score=best["score"], cand=best) score=best["score"], apply_mask=apply_mask,
cand=field_filter(best) if field_filter else best)
elif tier == "review": elif tier == "review":
meta_db.apply_enrichment_match(fn, chash, "review", source="text", meta_db.apply_enrichment_match(fn, chash, "review", source="text",
score=best["score"], score=best["score"],
@@ -6000,8 +6095,15 @@ def _background_enrich():
(kill-switch or the test env) skips phase 2 entirely. Never drains in a (kill-switch or the test env) skips phase 2 entirely. Never drains in a
loop a dead network would make that spin forever.""" loop a dead network would make that spin forever."""
_enrich_status["processed"] = 0 _enrich_status["processed"] = 0
# User settings gate the BACKGROUND matcher only (the review modal's
# manual search/fix stays available when it's off); read once per pass,
# up front so the pending query can honour the per-field apply mask
# (a re-enabled field re-queues its `matched` rows for backfill).
cfg = _load_config(CONFIG_DIR / "config.json") or {}
allowed_keys = frozenset(_ENRICH_APPLY_FIELDS) - _enrich_blocked_apply_keys(cfg)
apply_mask = _enrich_apply_mask(cfg)
try: try:
pending = meta_db.enrichment_pending(limit=100000) pending = meta_db.enrichment_pending(limit=100000, allowed_keys=allowed_keys)
except Exception: except Exception:
log.exception("enrichment: pending query failed") log.exception("enrichment: pending query failed")
return return
@@ -6013,9 +6115,6 @@ def _background_enrich():
_enrich_status["processed"] += 1 _enrich_status["processed"] += 1
_enrich_status["last_pass_at"] = time.time() _enrich_status["last_pass_at"] = time.time()
# User settings gate the BACKGROUND matcher only (the review modal's
# manual search/fix stays available when it's off); read once per pass.
cfg = _load_config(CONFIG_DIR / "config.json") or {}
if cfg.get("enrich_enabled", True) is False: if cfg.get("enrich_enabled", True) is False:
if pending: if pending:
log.info("Enrichment pass: %d rows stamped (matching disabled in Settings)", len(pending)) log.info("Enrichment pass: %d rows stamped (matching disabled in Settings)", len(pending))
@@ -6030,19 +6129,31 @@ def _background_enrich():
log.info("Enrichment pass: %d rows stamped (network disabled — matching skipped)", len(pending)) log.info("Enrichment pass: %d rows stamped (network disabled — matching skipped)", len(pending))
return return
# Scraper options (R1), read from the same per-pass cfg: `mb_on` gates
# the matcher (phase 2), `art_on` the cover-art fetch (phase 3 — the
# Cover Art Archive is the only automatic art source today, so the
# source toggle and the cover-art apply toggle both have to be on).
mb_on = cfg.get("enrich_src_musicbrainz", True) is not False
art_on = (cfg.get("enrich_src_caa", True) is not False
and cfg.get("enrich_apply_art", True) is not False)
field_filter = _enrich_field_filter(cfg)
now = time.time() now = time.time()
retriable = [] retriable = []
try: if mb_on:
retriable = [r for r in meta_db.enrichment_failed_rows(limit=100000) try:
if _enrich_backoff_elapsed(r.get("attempts"), r.get("last_attempt_at"), now)] retriable = [r for r in meta_db.enrichment_failed_rows(limit=100000)
except Exception: if _enrich_backoff_elapsed(r.get("attempts"), r.get("last_attempt_at"), now)]
log.exception("enrichment: failed-row query failed") except Exception:
log.exception("enrichment: failed-row query failed")
elif pending:
log.info("Enrichment pass: %d rows stamped (MusicBrainz source disabled in Settings)", len(pending))
matched = 0 matched = 0
# A `failed` row with a changed identity hash can surface in BOTH lists; # A `failed` row with a changed identity hash can surface in BOTH lists;
# de-dup by filename so each row consumes the rate budget only once. # de-dup by filename so each row consumes the rate budget only once.
seen_filenames = set() seen_filenames = set()
queue = [] queue = []
for row in pending + retriable: for row in (pending + retriable) if mb_on else []:
fn = row.get("filename") fn = row.get("filename")
if fn in seen_filenames: if fn in seen_filenames:
continue continue
@@ -6050,7 +6161,8 @@ def _background_enrich():
queue.append(row) queue.append(row)
for row in queue: for row in queue:
try: try:
_enrich_one(row, auto_min=auto_min) _enrich_one(row, auto_min=auto_min, field_filter=field_filter,
apply_mask=apply_mask)
matched += 1 matched += 1
except EnrichTransportError as e: except EnrichTransportError as e:
log.info("enrichment: network unavailable, pass paused (%s)", e) log.info("enrichment: network unavailable, pass paused (%s)", e)
@@ -6065,7 +6177,7 @@ def _background_enrich():
source="error", bump_attempts=True) source="error", bump_attempts=True)
except Exception: except Exception:
pass pass
if pending or retriable: if mb_on and (pending or retriable):
log.info("Enrichment pass: %d rows stamped, %d matched", len(pending), matched) log.info("Enrichment pass: %d rows stamped, %d matched", len(pending), matched)
# Phase 3 — cover art (R3/P9). For freshly-matched songs, resolve the art # Phase 3 — cover art (R3/P9). For freshly-matched songs, resolve the art
@@ -6073,6 +6185,10 @@ def _background_enrich():
# marked and skipped; the rest fetch the release's front cover from the # marked and skipped; the rest fetch the release's front cover from the
# Cover Art Archive into the size-capped cache. Same pause-on-transport- # Cover Art Archive into the size-capped cache. Same pause-on-transport-
# error rule as matching — a dead network never burns a row's evaluation. # error rule as matching — a dead network never burns a row's evaluation.
# Rows skipped here stay art_state NULL, so re-enabling the toggles picks
# them up on the next pass — nothing is permanently forfeited.
if not art_on:
return
try: try:
art_rows = meta_db.enrichment_art_pending(limit=100000) art_rows = meta_db.enrichment_art_pending(limit=100000)
except Exception: except Exception:
@@ -6636,10 +6752,13 @@ def api_enrichment_refresh(filename: str):
def api_enrichment_review(limit: int = 200): def api_enrichment_review(limit: int = 200):
"""The Match-Review queue: songs whose text match landed in the medium- """The Match-Review queue: songs whose text match landed in the medium-
confidence review tier, each with its stored candidate list the drawer confidence review tier, each with its stored candidate list the drawer
renders straight from this, no MusicBrainz round-trip.""" renders straight from this, no MusicBrainz round-trip. Ordered by the
user's enrich_review_order setting."""
limit = max(1, min(int(limit), 500)) limit = max(1, min(int(limit), 500))
cfg = _load_config(CONFIG_DIR / "config.json") or {}
order = cfg.get("enrich_review_order", "missing_first")
return { return {
"songs": meta_db.enrichment_review_queue(limit=limit), "songs": meta_db.enrichment_review_queue(limit=limit, order=order),
"total_review": meta_db.enrichment_state_counts().get("review", 0), "total_review": meta_db.enrichment_state_counts().get("review", 0),
} }
@@ -8938,6 +9057,22 @@ def _default_settings():
# review" option) sends every text match to review. # review" option) sends every text match to review.
"enrich_enabled": True, "enrich_enabled": True,
"enrich_auto_threshold": 0.9, "enrich_auto_threshold": 0.9,
# Scraper options (R1). Two axes, media-server style: sources say WHO
# may be contacted (MusicBrainz = the matcher, Cover Art Archive = the
# art fetch); the apply toggles say WHICH fields an AUTOMATIC match may
# canonicalize. A match the user confirms in the review modal always
# applies in full — these gate only what happens without them. All of
# it is display-side cache; nothing here ever writes to a pack file.
"enrich_src_musicbrainz": True,
"enrich_src_caa": True,
"enrich_apply_names": True,
"enrich_apply_year": True,
"enrich_apply_genres": True,
"enrich_apply_art": True,
# Review-queue ordering: missing_first = charts lacking album/year
# surface first (they gain the most), artist = AZ, recent = newest
# files first.
"enrich_review_order": "missing_first",
} }
@@ -9108,6 +9243,21 @@ def save_settings(data: dict):
if not math.isfinite(t) or not (0.5 <= t <= 1.01): if not math.isfinite(t) or not (0.5 <= t <= 1.01):
return {"error": "enrich_auto_threshold must be a number between 0.5 and 1.01"} return {"error": "enrich_auto_threshold must be a number between 0.5 and 1.01"}
updates["enrich_auto_threshold"] = t updates["enrich_auto_threshold"] = t
for _bool_key in ("enrich_src_musicbrainz", "enrich_src_caa",
"enrich_apply_names", "enrich_apply_year",
"enrich_apply_genres", "enrich_apply_art"):
if _bool_key in data:
raw = data[_bool_key]
if raw is not None:
if not isinstance(raw, bool):
return {"error": f"{_bool_key} must be a boolean"}
updates[_bool_key] = raw
if "enrich_review_order" in data:
raw = data["enrich_review_order"]
if raw is not None:
if not isinstance(raw, str) or raw not in ("missing_first", "artist", "recent"):
return {"error": "enrich_review_order must be one of missing_first, artist, recent"}
updates["enrich_review_order"] = raw
if "miss_penalty" in data: if "miss_penalty" in data:
raw = data["miss_penalty"] raw = data["miss_penalty"]
if raw is not None: if raw is not None:
+29 -1
View File
@@ -742,7 +742,7 @@
<div class="fb-srow-desc">Matches your charts against MusicBrainz in the background to tidy names, years and genres — display only, your files are never modified. Matches at or above the confidence level apply automatically; the rest wait in the library's review queue. Turning this off never disables manual match fixes.</div> <div class="fb-srow-desc">Matches your charts against MusicBrainz in the background to tidy names, years and genres — display only, your files are never modified. Matches at or above the confidence level apply automatically; the rest wait in the library's review queue. Turning this off never disables manual match fixes.</div>
</div> </div>
<div class="grid grid-cols-2 gap-2 mb-1 text-xs text-gray-400 fb-srow-wide"> <div class="grid grid-cols-2 gap-2 mb-1 text-xs text-gray-400 fb-srow-wide">
<label class="flex items-center gap-2"><input type="checkbox" id="enrich-enabled" checked class="rounded border-gray-600 bg-dark-700 text-accent"> Match songs against MusicBrainz</label> <label class="flex items-center gap-2"><input type="checkbox" id="enrich-enabled" checked class="rounded border-gray-600 bg-dark-700 text-accent"> Match songs automatically</label>
<label class="flex items-center gap-2">Auto-apply confidence <label class="flex items-center gap-2">Auto-apply confidence
<select id="enrich-threshold" class="bg-dark-700 border border-gray-800 rounded-xl px-2 py-1.5 text-xs text-gray-300 outline-none"> <select id="enrich-threshold" class="bg-dark-700 border border-gray-800 rounded-xl px-2 py-1.5 text-xs text-gray-300 outline-none">
<option value="0.85">85% — more auto-matches</option> <option value="0.85">85% — more auto-matches</option>
@@ -752,6 +752,34 @@
</select> </select>
</label> </label>
</div> </div>
<!-- R1 scraper options: sources (who may be contacted) ×
auto-apply fields (what a confident match fills in). -->
<div class="fb-srow-wide mb-1">
<div class="text-[10px] uppercase tracking-wide text-gray-500 mb-1">Sources</div>
<div class="grid grid-cols-2 gap-2 text-xs text-gray-400">
<label class="flex items-center gap-2"><input type="checkbox" id="enrich-src-musicbrainz" checked class="rounded border-gray-600 bg-dark-700 text-accent"> MusicBrainz — names, years, genres</label>
<label class="flex items-center gap-2"><input type="checkbox" id="enrich-src-caa" checked class="rounded border-gray-600 bg-dark-700 text-accent"> Cover Art Archive — album covers</label>
</div>
</div>
<div class="fb-srow-wide mb-1">
<div class="text-[10px] uppercase tracking-wide text-gray-500 mb-1">Auto-apply</div>
<div class="grid grid-cols-4 gap-2 text-xs text-gray-400">
<label class="flex items-center gap-2"><input type="checkbox" id="enrich-apply-names" checked class="rounded border-gray-600 bg-dark-700 text-accent"> Names</label>
<label class="flex items-center gap-2"><input type="checkbox" id="enrich-apply-year" checked class="rounded border-gray-600 bg-dark-700 text-accent"> Year</label>
<label class="flex items-center gap-2"><input type="checkbox" id="enrich-apply-genres" checked class="rounded border-gray-600 bg-dark-700 text-accent"> Genres</label>
<label class="flex items-center gap-2"><input type="checkbox" id="enrich-apply-art" checked class="rounded border-gray-600 bg-dark-700 text-accent"> Cover art</label>
</div>
<div class="text-[11px] text-gray-600 mt-1">What a confident match may fill in on its own — matches you confirm in the review queue always apply in full.</div>
</div>
<div class="grid grid-cols-2 gap-2 mb-1 text-xs text-gray-400 fb-srow-wide">
<label class="flex items-center gap-2">Review queue order
<select id="enrich-review-order" class="bg-dark-700 border border-gray-800 rounded-xl px-2 py-1.5 text-xs text-gray-300 outline-none">
<option value="missing_first" selected>Missing info first</option>
<option value="artist">Artist AZ</option>
<option value="recent">Recently added</option>
</select>
</label>
</div>
<div class="fb-srow-control"> <div class="fb-srow-control">
<button id="enrich-match-now" class="bg-dark-600 hover:bg-dark-500 px-5 py-2.5 rounded-xl text-sm text-gray-300 transition">Match Now</button> <button id="enrich-match-now" class="bg-dark-600 hover:bg-dark-500 px-5 py-2.5 rounded-xl text-sm text-gray-300 transition">Match Now</button>
<span id="enrich-status" class="text-xs text-gray-500"></span> <span id="enrich-status" class="text-xs text-gray-500"></span>
+23 -4
View File
@@ -401,16 +401,28 @@
// Markup lives statically in index.html (the v3 settings pattern); this // Markup lives statically in index.html (the v3 settings pattern); this
// wires it. All null-guarded so v2 (which lacks the elements) no-ops. // wires it. All null-guarded so v2 (which lacks the elements) no-ops.
function wireSettingsCard() { function wireSettingsCard() {
const toggle = document.getElementById('enrich-enabled');
const sel = document.getElementById('enrich-threshold'); const sel = document.getElementById('enrich-threshold');
const order = document.getElementById('enrich-review-order');
const btn = document.getElementById('enrich-match-now'); const btn = document.getElementById('enrich-match-now');
if (!toggle && !sel && !btn) return; // Boolean toggles, element id → settings key. enrich-enabled is the
// master background switch; the rest are the R1 scraper options
// (per-source + per-field auto-apply).
const toggles = [
['enrich-enabled', 'enrich_enabled'],
['enrich-src-musicbrainz', 'enrich_src_musicbrainz'],
['enrich-src-caa', 'enrich_src_caa'],
['enrich-apply-names', 'enrich_apply_names'],
['enrich-apply-year', 'enrich_apply_year'],
['enrich-apply-genres', 'enrich_apply_genres'],
['enrich-apply-art', 'enrich_apply_art'],
].map(([id, key]) => [document.getElementById(id), key]).filter(([el]) => el);
if (!toggles.length && !sel && !btn) return;
(async () => { (async () => {
try { try {
const r = await fetch('/api/settings'); const r = await fetch('/api/settings');
if (r.ok) { if (r.ok) {
const cfg = await r.json(); const cfg = await r.json();
if (toggle) toggle.checked = cfg.enrich_enabled !== false; for (const [el, key] of toggles) el.checked = cfg[key] !== false;
if (sel) { if (sel) {
const t = Number(cfg.enrich_auto_threshold); const t = Number(cfg.enrich_auto_threshold);
const want = Number.isFinite(t) ? t : 0.9; const want = Number.isFinite(t) ? t : 0.9;
@@ -421,13 +433,20 @@
} }
if (best) sel.value = best.value; if (best) sel.value = best.value;
} }
if (order) {
const v = String(cfg.enrich_review_order || 'missing_first');
order.value = ['missing_first', 'artist', 'recent'].includes(v) ? v : 'missing_first';
}
} }
} catch (_) { /* leave markup defaults */ } } catch (_) { /* leave markup defaults */ }
refreshChip(); // also fills #enrich-status refreshChip(); // also fills #enrich-status
})(); })();
const save = (key, value) => post('/api/settings', { [key]: value }); const save = (key, value) => post('/api/settings', { [key]: value });
toggle?.addEventListener('change', () => save('enrich_enabled', !!toggle.checked)); for (const [el, key] of toggles) {
el.addEventListener('change', () => save(key, !!el.checked));
}
sel?.addEventListener('change', () => save('enrich_auto_threshold', Number(sel.value))); sel?.addEventListener('change', () => save('enrich_auto_threshold', Number(sel.value)));
order?.addEventListener('change', () => save('enrich_review_order', order.value));
btn?.addEventListener('click', async () => { btn?.addEventListener('click', async () => {
await post('/api/enrichment/kick'); await post('/api/enrichment/kick');
const line = document.getElementById('enrich-status'); const line = document.getElementById('enrich-status');
+333
View File
@@ -0,0 +1,333 @@
"""Tests for the R1 scraper options: per-source toggles (MusicBrainz /
Cover Art Archive), per-field auto-apply toggles (names / year / genres /
cover art), and the review-queue order preference.
Same no-network contract as the P8/R3 suites: both transports are faked
over their single seams (`_mb_http_get`, `_caa_http_get`) and the network
flag is only force-enabled where a test needs the pipeline to run.
"""
import importlib
import io as _io
import sys
import pytest
from fastapi.testclient import TestClient
from PIL import Image
@pytest.fixture()
def server(tmp_path, monkeypatch, isolate_logging):
monkeypatch.setenv("CONFIG_DIR", str(tmp_path / "config"))
dlc = tmp_path / "dlc"
dlc.mkdir()
monkeypatch.setenv("DLC_DIR", str(dlc))
monkeypatch.setenv("FEEDBACK_SKIP_STARTUP_TASKS", "1")
sys.modules.pop("server", None)
srv = importlib.import_module("server")
try:
yield srv
finally:
conn = getattr(getattr(srv, "meta_db", None), "conn", None)
if conn is not None:
conn.close()
sys.modules.pop("server", None)
@pytest.fixture()
def client(server):
return TestClient(server.app)
class FakeMB:
"""Canned MusicBrainz over the `_mb_http_get` seam (the P8 fixture)."""
def __init__(self):
self.calls = []
self.search_response = {"recordings": []}
def __call__(self, path, params):
self.calls.append((path, dict(params)))
if path == "recording":
return self.search_response
raise AssertionError(f"unexpected MB path {path!r}")
@pytest.fixture()
def mb(server, monkeypatch):
fake = FakeMB()
monkeypatch.setattr(server, "_mb_http_get", fake)
monkeypatch.setattr(server, "_enrich_network_enabled", lambda: True)
return fake
def _put(server, fn, title="Thunderstruck (v2)", artist="ACDC", album="",
duration=292, year="1990", mtime=0):
server.meta_db.put(fn, mtime, 0, {
"title": title, "artist": artist, "album": album, "year": year,
"duration": duration, "arrangements": [{"name": "Lead", "index": 0}],
})
def mb_doc(rid="rec-1", title="Thunderstruck", artist="AC/DC", artist_id="art-1",
album="The Razors Edge", date="1990-09-24", length_ms=292000, score=100):
return {
"id": rid, "score": score, "title": title, "length": length_ms,
"isrcs": ["AUAP09000045"],
"artist-credit": [{"name": artist, "artist": {
"id": artist_id, "name": artist, "sort-name": artist}}],
"releases": [{"id": "rel-1", "title": album, "status": "Official",
"date": date, "release-group": {"primary-type": "Album"}}],
"tags": [{"name": "hard rock", "count": 7}],
}
# ── settings keys ─────────────────────────────────────────────────────────────
BOOL_KEYS = ("enrich_src_musicbrainz", "enrich_src_caa", "enrich_apply_names",
"enrich_apply_year", "enrich_apply_genres", "enrich_apply_art")
def test_scraper_option_keys_validate_and_persist(client):
for key in BOOL_KEYS:
bad = client.post("/api/settings", json={key: "yes"}).json()
assert "error" in bad
client.post("/api/settings", json={key: False})
assert client.get("/api/settings").json()[key] is False
assert "error" in client.post(
"/api/settings", json={"enrich_review_order": "random"}).json()
assert "error" in client.post(
"/api/settings", json={"enrich_review_order": 42}).json()
client.post("/api/settings", json={"enrich_review_order": "artist"})
assert client.get("/api/settings").json()["enrich_review_order"] == "artist"
def test_defaults_present(server):
d = server._default_settings()
for key in BOOL_KEYS:
assert d[key] is True
assert d["enrich_review_order"] == "missing_first"
# ── per-source: MusicBrainz ───────────────────────────────────────────────────
def test_musicbrainz_source_off_stamps_without_matching(server, mb, client):
client.post("/api/settings", json={"enrich_src_musicbrainz": False})
_put(server, "a.sloppak")
mb.search_response = {"recordings": [mb_doc()]}
server._background_enrich()
assert mb.calls == []
row = server.meta_db.get_enrichment("a.sloppak")
assert row["match_state"] == "unscanned" # hash stamped, no match
# Re-enabling picks the same row up on the next pass.
client.post("/api/settings", json={"enrich_src_musicbrainz": True})
server._background_enrich()
assert server.meta_db.get_enrichment("a.sloppak")["match_state"] == "matched"
# ── per-field auto-apply ──────────────────────────────────────────────────────
def test_field_toggles_strip_auto_applied_fields(server, mb, client):
client.post("/api/settings", json={"enrich_apply_year": False,
"enrich_apply_genres": False})
_put(server, "a.sloppak")
mb.search_response = {"recordings": [mb_doc()]}
server._background_enrich()
row = server.meta_db.get_enrichment("a.sloppak")
assert row["match_state"] == "matched"
assert row["canon_artist"] == "AC/DC"
assert row["canon_title"] == "Thunderstruck"
assert row["canon_album"] == "The Razors Edge"
assert row["canon_year"] is None # toggled off
assert row["genres"] == [] # toggled off
# Identity ids always stamp — the art fetch and re-matching need them.
assert row["mb_recording_id"] == "rec-1"
assert row["mb_release_id"] == "rel-1"
def test_names_toggle_keeps_ids_and_other_fields(server, mb, client):
client.post("/api/settings", json={"enrich_apply_names": False})
_put(server, "a.sloppak")
mb.search_response = {"recordings": [mb_doc()]}
server._background_enrich()
row = server.meta_db.get_enrichment("a.sloppak")
assert row["match_state"] == "matched"
assert row["canon_artist"] is None
assert row["canon_title"] is None
assert row["canon_album"] is None
assert row["canon_year"] == "1990"
assert row["genres"] == ["hard rock"]
assert row["mb_recording_id"] == "rec-1"
def test_review_accept_applies_all_fields_despite_toggles(server, mb, client):
"""A match the USER confirms is their intent — the auto-apply toggles
gate only what happens without them."""
client.post("/api/settings", json={"enrich_apply_names": False,
"enrich_apply_year": False,
"enrich_apply_genres": False})
# Partial artist agreement → review tier (candidates stored unfiltered).
_put(server, "a.sloppak", artist="AC/DC ft Nobody")
mb.search_response = {"recordings": [mb_doc()]}
server._background_enrich()
assert server.meta_db.get_enrichment("a.sloppak")["match_state"] == "review"
r = client.post("/api/enrichment/review/a.sloppak/accept",
json={"recording_id": "rec-1"})
assert r.status_code == 200
row = server.meta_db.get_enrichment("a.sloppak")
assert row["match_state"] == "manual"
assert row["canon_artist"] == "AC/DC"
assert row["canon_year"] == "1990"
assert row["genres"] == ["hard rock"]
# ── per-field: nothing-forfeited (re-enable backfills, no partial seeding) ─────
def test_reenabling_field_backfills_matched_row(server, mb, client):
"""A field toggled OFF strips the value AND records it in apply_mask, so
turning the field back on re-queues the (unchanged-hash) matched row and
backfills — the same "nothing forfeited" contract the source/art toggles
keep."""
client.post("/api/settings", json={"enrich_apply_year": False})
_put(server, "a.sloppak")
mb.search_response = {"recordings": [mb_doc()]}
server._background_enrich()
row = server.meta_db.get_enrichment("a.sloppak")
assert row["match_state"] == "matched"
assert row["canon_year"] is None # suppressed
assert row["apply_mask"] == "enrich_apply_year" # …and remembered
# Re-enable → next pass re-queues and backfills the year (hash unchanged).
client.post("/api/settings", json={"enrich_apply_year": True})
server._background_enrich()
row = server.meta_db.get_enrichment("a.sloppak")
assert row["match_state"] == "matched"
assert row["canon_year"] == "1990" # backfilled
assert row["apply_mask"] in (None, "") # fully applied now
# Converged: a fully-applied row is not re-queued again.
assert server.meta_db.enrichment_pending(
allowed_keys=frozenset(server._ENRICH_APPLY_FIELDS)) == []
def test_partial_match_is_not_a_cache_donor(server):
"""A row that suppressed a display field must not seed a sibling chart of
the same recording with its blank — enrichment_cache_lookup skips it, so
the sibling falls through to its own (correctly-filtered) match. A fully-
applied row IS a donor."""
h = server.meta_db.enrichment_content_hash("ACDC", "Thunderstruck (v2)", "", 292)
# Partial donor: matched with the year suppressed (apply_mask set).
server.meta_db.apply_enrichment_match(
"a.sloppak", h, "matched", source="text", score=1.0,
cand={"recording_id": "rec-1", "artist": "AC/DC", "title": "Thunderstruck"},
apply_mask="enrich_apply_year")
assert server.meta_db.enrichment_cache_lookup(
h, exclude_filename="b.sloppak") is None
# Fully-applied donor (no apply_mask): offered, with its year intact.
server.meta_db.apply_enrichment_match(
"c.sloppak", h, "matched", source="text", score=1.0,
cand={"recording_id": "rec-1", "artist": "AC/DC",
"title": "Thunderstruck", "year": "1990"})
donor = server.meta_db.enrichment_cache_lookup(h, exclude_filename="b.sloppak")
assert donor is not None and donor["year"] == "1990"
# ── per-source / per-field: cover art ─────────────────────────────────────────
def png_bytes(color=(200, 30, 30)):
buf = _io.BytesIO()
Image.new("RGB", (4, 4), color).save(buf, "PNG")
return buf.getvalue()
def make_sloppak(server, name, title="Song", artist="Artist"):
d = server.DLC_DIR / name
d.mkdir(parents=True)
(d / "manifest.yaml").write_text(
f"title: {title}\nartist: {artist}\nduration: 100\n"
"arrangements: []\nstems: []\n", encoding="utf-8")
server.meta_db.put(name, 0, 0, {
"title": title, "artist": artist, "album": "", "year": "",
"duration": 100, "arrangements": [{"name": "Lead", "index": 0}],
})
return d
def _match_row(server, fn, release_id="rel-1"):
song = server.meta_db.enrichment_song_row(fn)
h = server.meta_db.enrichment_content_hash(
song["artist"], song["title"], song["album"], song["duration"])
server.meta_db.apply_enrichment_match(
fn, h, "matched", source="text", score=1.0,
cand={"recording_id": "rec-1", "release_id": release_id,
"title": song["title"], "artist": song["artist"]})
@pytest.fixture()
def caa(server, monkeypatch):
calls = []
art = {"rel-1": png_bytes((60, 60, 200))}
def fake(release_id):
calls.append(release_id)
return art.get(release_id)
fake.calls = calls
monkeypatch.setattr(server, "_caa_http_get", fake)
monkeypatch.setattr(server, "_enrich_network_enabled", lambda: True)
return fake
def test_caa_source_toggle_gates_art_fetch(server, client, caa):
make_sloppak(server, "a.sloppak")
_match_row(server, "a.sloppak")
client.post("/api/settings", json={"enrich_src_caa": False})
server._background_enrich()
assert caa.calls == []
row = server.meta_db.get_enrichment("a.sloppak")
assert row["art_state"] is None # not forfeited, just skipped
# Re-enable → the same row is picked up.
client.post("/api/settings", json={"enrich_src_caa": True})
server._background_enrich()
assert server.meta_db.get_enrichment("a.sloppak")["art_state"] == "caa"
def test_apply_art_toggle_gates_art_fetch(server, client, caa):
make_sloppak(server, "a.sloppak")
_match_row(server, "a.sloppak")
client.post("/api/settings", json={"enrich_apply_art": False})
server._background_enrich()
assert caa.calls == []
assert server.meta_db.get_enrichment("a.sloppak")["art_state"] is None
# ── review-queue order ────────────────────────────────────────────────────────
def _review_row(server, fn):
server.meta_db.apply_enrichment_match(
fn, "h", "review", source="text", score=0.7,
candidates=[{"recording_id": "rec-1", "title": "T", "artist": "A"}])
def _review_filenames(client):
return [s["filename"] for s in
client.get("/api/enrichment/review").json()["songs"]]
def test_review_queue_order_setting(server, client):
# a: complete, artist Alpha, oldest; b: missing album+year, artist Zeta,
# middle; c: missing year, artist Mid, newest — the three orders differ.
_put(server, "a.sloppak", title="Song A", artist="Alpha",
album="Full", year="1990", mtime=100)
_put(server, "b.sloppak", title="Song B", artist="Zeta",
album="", year="", mtime=200)
_put(server, "c.sloppak", title="Song C", artist="Mid",
album="Full", year="", mtime=300)
for fn in ("a.sloppak", "b.sloppak", "c.sloppak"):
_review_row(server, fn)
# Default: missing-data-first.
assert _review_filenames(client) == ["b.sloppak", "c.sloppak", "a.sloppak"]
client.post("/api/settings", json={"enrich_review_order": "artist"})
assert _review_filenames(client) == ["a.sloppak", "c.sloppak", "b.sloppak"]
client.post("/api/settings", json={"enrich_review_order": "recent"})
assert _review_filenames(client) == ["c.sloppak", "b.sloppak", "a.sloppak"]
# An unknown stored value degrades to the default order, never an error.
assert server.meta_db.enrichment_review_queue(order="bogus")[0]["filename"] == "b.sloppak"