mirror of
https://github.com/got-feedBack/feedBack.git
synced 2026-08-11 03:09:57 +00:00
library: MusicBrainz text matching + Match-Review UI (P8) (#710)
* library: MusicBrainz text matching + Match-Review UI — P8
Replaces the enrichment plumbing's no-op matcher (P7) with the real
pipeline, per the library-metadata design: a wrong match is worse than a
slow one, so medium confidence goes to a human review queue and never
straight to canonical values.
- lib/mb_match.py (new, pure — no network/DB/server imports): denoise
(author credits, (440Hz)/(Live)/(No Lead)/(v2) parentheticals,
diacritics/punctuation, ACDC / AC DC / AC/DC folding via compacted
token equality), token-set similarity, scoring with year/duration
corroboration bonuses, tier classification (auto needs combined
>= 0.95 AND per-field floors — a perfect-title cover by the wrong
artist, or a chart with no artist, can never auto-match), Lucene
query building, MusicBrainz response normalization.
- Matcher precedence in _enrich_one: content-hash cache copy (another
chart of the same recording matches with no network) -> manifest
mbid (tier 0) / isrc (tier 1) exact keys, feature-detected and
strictly shape-validated, read-only -> text search tiers
(auto / review / failed).
- Lifecycle: review rows store their ranked candidate list (JSON) and
write NO canonical fields until a human accepts; failed rows retry on
an exponential backoff (1 h doubling, 7 d cap) via the attempts
column; user-rejected rows never auto-retry; an identity edit
re-queues anything and resets the backoff; never-overwrite-manual is
enforced inside the single writer (apply_enrichment_match) so no call
path can forget it.
- Network: _mb_http_get is the one transport seam — throttled to
<= 1 req/s through P7's _enrich_throttle, identified with a real
User-Agent from VERSION, and a 503 pauses the whole pass without
burning attempts. Offline guard: no sockets under
FEEDBACK_ENRICH_OFFLINE or FEEDBACK_SKIP_STARTUP_TASKS, so pytest can
never reach MusicBrainz; the pass still stamps identity hashes
(two-phase), which is why every P7 test passes unchanged.
- Routes: GET /api/enrichment/review, POST
/api/enrichment/review/{filename}/accept|reject|pick, GET
/api/enrichment/search (throttled manual-search proxy). All four are
demo-mode blocked.
- Match facet: match= CSV accepted by /api/library AND
/api/library/stats (the A-Z rail's letter counts stay lockstep with
the grid) — review / matched (incl. manual) / unmatched / pending,
the same EXISTS idiom as the mastery facet.
- UI: static/v3/match-review.js (new, self-contained) — an ambient
"N to review" chip beside the song count (rendered only when
non-zero; silent on success, no toasts), and a review drawer on the
filter-drawer slide idiom (Escape + focus trap; row click accepts,
"Not a match" rejects, "Search instead" is the fix-match escape
hatch). songs.js gets the chip mount, a Match filter section, and
session-only match state; also fixes the latent applySavedPrefs bug
where restored filters dropped the mastery key, which made the
filter drawer throw for anyone with saved prefs.
- static/tailwind.min.css regenerated (scripts/build-tailwind.sh) for
the new utility classes; conflicts with sibling PRs resolve by
re-running the script.
Nothing is ever written to pack files — canonical values live only in
the song_enrichment display cache. Cover art caching and acoustic
fingerprinting are follow-up slices.
22 pure unit tests + 19 server tests (fake transport injected over the
_mb_http_get seam) + demo-mode route cases; full-suite failure set
A/B-identical with the change stashed vs applied.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nm7tHs1Yvjjtnnu4nzJgdN
* library: match-review modal + configurable auto-apply confidence (P8 R0)
Follow-up to the initial P8 commit, folding in the first round of tester
feedback on the review surface and the matcher's knobs:
- Review GUI is a centred MODAL now, not a sidebar — one chart at a
time (the scraper-review model from media-server / emulation-frontend
apps): the chart's current metadata with explicit amber
"Missing: album / year / cover art" chips (art detected via the art
request failing), candidates each carrying "Adds: year - genres -
ISRC" / "Shows as: ACDC -> AC/DC" per-field chips, and Skip /
Not a match / Search instead / Use selected with prev-next + arrow-key
navigation. Chip + window API surface unchanged, so songs.js needed no
edits for the rework.
- Auto-apply confidence is a SETTING: default drops 0.95 -> 0.90
(mb_match.AUTO_MIN; classify() takes an auto_min override). The
per-field floors are untouched and threshold-independent — a
perfect-title cover by the wrong artist still can't auto-match at any
setting. New validated settings keys: enrich_enabled (bool) +
enrich_auto_threshold (0.5–1.01; >1.0 = "Always review", since a
capped score can equal exactly 1.0). Read once per pass; disabling
gates only the BACKGROUND matcher — manual search/fix stays available.
- Settings -> Library -> "Metadata matching" card: enable toggle,
confidence select (85 / 90 / 95 / Always review), a Match Now button
(new POST /api/enrichment/kick, single-flight like every other kick,
demo-mode blocked), and a live status line fed by the same fetch as
the review chip. Markup in index.html per the v3 settings pattern,
wired by match-review.js, null-guarded so v2 no-ops.
- Review queue orders missing-data charts first — confirming those has
the most to gain; complete charts only stand to be re-labelled.
Tests: threshold moves the auto/review boundary via settings; the
enable toggle gates matching but not the manual proxy; settings
validation; kick route; queue ordering; classify(auto_min=...) floors.
Full-suite failure set byte-identical to the pre-change baseline.
tailwind.min.css regenerated for the modal's utility classes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nm7tHs1Yvjjtnnu4nzJgdN
* fix(library): lock MusicBrainz throttle across sleep + de-dup enrich queue (PR #710 review)
Hold a module-level lock across _enrich_throttle's read/sleep/write so the
background daemon and threadpooled sync search route serialize outbound MB
requests instead of bursting past the 1 req/s limit. De-dup the enrich queue
by filename so a changed-hash failed row isn't processed twice per pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: byrongamatos <xasiklas@gmail.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
byrongamatos
parent
7c15cdda66
commit
74cd08f765
+295
@@ -0,0 +1,295 @@
|
|||||||
|
"""Text-matching engine for MusicBrainz metadata enrichment (P8).
|
||||||
|
|
||||||
|
Pure functions only — no network, no database, no server imports — so the
|
||||||
|
whole matching pipeline is unit-testable in isolation. server.py owns the
|
||||||
|
throttled HTTP transport and the song_enrichment writes; this module owns:
|
||||||
|
|
||||||
|
* denoise/tokenize: fold community chart-title noise (author suffixes,
|
||||||
|
``(440Hz)``/``(Live)``/``(No Lead)``/``(v2)`` parentheticals, punctuation,
|
||||||
|
diacritics, ``AC DC``/``ACDC``/``AC/DC`` spelling drift) into a comparable
|
||||||
|
token form,
|
||||||
|
* similarity + scoring: token-set similarity on artist+title with year and
|
||||||
|
duration proximity as corroborating bonuses,
|
||||||
|
* tier classification: auto (high) / review (medium) / none (low) — the
|
||||||
|
design rule is that a WRONG match is worse than no match, so the auto
|
||||||
|
tier is deliberately strict and medium confidence goes to a human,
|
||||||
|
* MusicBrainz JSON parsing: normalize ``/ws/2`` recording documents into
|
||||||
|
the flat candidate dicts the review UI and song_enrichment store.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import re
|
||||||
|
import unicodedata
|
||||||
|
|
||||||
|
# ── Tier thresholds ───────────────────────────────────────────────────────────
|
||||||
|
# Combined score = 0.5*artist_sim + 0.5*title_sim + corroboration bonuses
|
||||||
|
# (capped at 1.0). Wrong-match is worse than slow (design §5), so `auto`
|
||||||
|
# additionally requires BOTH fields to individually agree — a perfect title
|
||||||
|
# with a mismatched artist (a cover) must never auto-canonicalize, whatever
|
||||||
|
# the combined threshold is set to. AUTO_MIN is only the DEFAULT: the host
|
||||||
|
# surfaces it as the user-configurable "auto-apply confidence" setting and
|
||||||
|
# passes the chosen value into classify(auto_min=…).
|
||||||
|
AUTO_MIN = 0.90
|
||||||
|
AUTO_ARTIST_MIN = 0.8
|
||||||
|
AUTO_TITLE_MIN = 0.6
|
||||||
|
REVIEW_MIN = 0.65
|
||||||
|
|
||||||
|
YEAR_BONUS = 0.05 # candidate year within ±1 of the chart's year
|
||||||
|
DURATION_BONUS = 0.05 # candidate length within 5s of the chart's audio
|
||||||
|
DURATION_BONUS_LOOSE = 0.025 # …within 15s
|
||||||
|
_DURATION_TIGHT = 5
|
||||||
|
_DURATION_LOOSE = 15
|
||||||
|
|
||||||
|
# ── Denoise ───────────────────────────────────────────────────────────────────
|
||||||
|
# A parenthetical/bracketed group is dropped when it contains any of these
|
||||||
|
# noise terms as a whole word (chart-variant markers, tuning/pitch notes,
|
||||||
|
# performance qualifiers) or when it reads as an author credit ("by X",
|
||||||
|
# "charted by X"). Both sides of a comparison are denoised symmetrically, so
|
||||||
|
# over-stripping a meaningful group costs a little precision but never
|
||||||
|
# produces an asymmetric mismatch.
|
||||||
|
_NOISE_TERMS = (
|
||||||
|
r"440\s*hz", r"a440", r"432\s*hz",
|
||||||
|
r"live", r"acoustic", r"instrumental",
|
||||||
|
r"no\s+(?:lead|rhythm|bass|vocals?|drums)",
|
||||||
|
r"(?:lead|rhythm|bass)\s+only",
|
||||||
|
r"v\d+", r"ver(?:sion)?\s*\d+",
|
||||||
|
r"remaster(?:ed)?(?:\s*\d{4})?", r"re-?recorded?",
|
||||||
|
r"fix(?:ed)?", r"updated?",
|
||||||
|
r"bonus", r"custom",
|
||||||
|
)
|
||||||
|
_NOISE_GROUP_RE = re.compile(
|
||||||
|
r"[(\[][^)\]]*\b(?:" + "|".join(_NOISE_TERMS) + r")\b[^)\]]*[)\]]",
|
||||||
|
re.IGNORECASE,
|
||||||
|
)
|
||||||
|
# Author credits: "(by SomeCharter)", "[charted by X]", "(chart by X)".
|
||||||
|
_AUTHOR_GROUP_RE = re.compile(
|
||||||
|
r"[(\[]\s*(?:chart(?:ed)?\s+)?by\s+[^)\]]*[)\]]", re.IGNORECASE)
|
||||||
|
# Trailing "- by SomeCharter" outside parens.
|
||||||
|
_AUTHOR_TAIL_RE = re.compile(r"\s+-\s+(?:chart(?:ed)?\s+)?by\s+.+$", re.IGNORECASE)
|
||||||
|
|
||||||
|
_PUNCT_RE = re.compile(r"[^\w\s]|_")
|
||||||
|
_WS_RE = re.compile(r"\s+")
|
||||||
|
|
||||||
|
|
||||||
|
def _strip_diacritics(s: str) -> str:
|
||||||
|
return "".join(
|
||||||
|
ch for ch in unicodedata.normalize("NFKD", s)
|
||||||
|
if not unicodedata.combining(ch)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def denoise(s, *, strip_leading_the: bool = False) -> str:
|
||||||
|
"""Fold a community metadata string into its comparable form:
|
||||||
|
lowercase, diacritics stripped, noise parentheticals and author credits
|
||||||
|
removed, punctuation collapsed to spaces. ``strip_leading_the`` drops a
|
||||||
|
leading "The " — used for ARTIST comparison only ("The Beatles" ==
|
||||||
|
"Beatles"), never titles ("The Trooper" must keep its "the")."""
|
||||||
|
s = str(s or "")
|
||||||
|
s = _NOISE_GROUP_RE.sub(" ", s)
|
||||||
|
s = _AUTHOR_GROUP_RE.sub(" ", s)
|
||||||
|
s = _AUTHOR_TAIL_RE.sub(" ", s)
|
||||||
|
s = _strip_diacritics(s).casefold()
|
||||||
|
s = s.replace("&", " and ")
|
||||||
|
s = _PUNCT_RE.sub(" ", s)
|
||||||
|
s = _WS_RE.sub(" ", s).strip()
|
||||||
|
if strip_leading_the and s.startswith("the "):
|
||||||
|
s = s[4:]
|
||||||
|
return s
|
||||||
|
|
||||||
|
|
||||||
|
def tokens(s, **kw) -> list[str]:
|
||||||
|
d = denoise(s, **kw)
|
||||||
|
return d.split() if d else []
|
||||||
|
|
||||||
|
|
||||||
|
def _compact(toks: list[str]) -> str:
|
||||||
|
return "".join(toks)
|
||||||
|
|
||||||
|
|
||||||
|
def similarity(a, b, *, artist: bool = False) -> float:
|
||||||
|
"""Token-set similarity in [0, 1]. Dice coefficient over the denoised
|
||||||
|
token sets, with a compacted-string equality fold so spelling drift that
|
||||||
|
only moves token boundaries ("ACDC" / "AC DC" / "AC/DC", "Greenday" /
|
||||||
|
"Green Day") counts as identical."""
|
||||||
|
kw = {"strip_leading_the": artist}
|
||||||
|
ta, tb = tokens(a, **kw), tokens(b, **kw)
|
||||||
|
if not ta or not tb:
|
||||||
|
return 0.0
|
||||||
|
if _compact(ta) == _compact(tb):
|
||||||
|
return 1.0
|
||||||
|
sa, sb = set(ta), set(tb)
|
||||||
|
return 2.0 * len(sa & sb) / (len(sa) + len(sb))
|
||||||
|
|
||||||
|
|
||||||
|
def _year_int(v):
|
||||||
|
try:
|
||||||
|
y = int(str(v)[:4])
|
||||||
|
return y if y > 0 else None
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _duration_int(v):
|
||||||
|
try:
|
||||||
|
d = int(round(float(v)))
|
||||||
|
return d if d > 0 else None
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def score_candidate(song: dict, cand: dict) -> float:
|
||||||
|
"""Combined confidence that MusicBrainz candidate `cand` is the song the
|
||||||
|
chart transcribes. 0.5*artist + 0.5*title, plus small year/duration
|
||||||
|
corroboration bonuses, capped at 1.0. Missing fields score 0 on their
|
||||||
|
half — classify() separately refuses to auto-match without both."""
|
||||||
|
artist_sim = similarity(song.get("artist"), cand.get("artist"), artist=True)
|
||||||
|
title_sim = similarity(song.get("title"), cand.get("title"))
|
||||||
|
score = 0.5 * artist_sim + 0.5 * title_sim
|
||||||
|
sy, cy = _year_int(song.get("year")), _year_int(cand.get("year"))
|
||||||
|
if sy and cy and abs(sy - cy) <= 1:
|
||||||
|
score += YEAR_BONUS
|
||||||
|
sd, cd = _duration_int(song.get("duration")), _duration_int(cand.get("duration"))
|
||||||
|
if sd and cd:
|
||||||
|
diff = abs(sd - cd)
|
||||||
|
if diff <= _DURATION_TIGHT:
|
||||||
|
score += DURATION_BONUS
|
||||||
|
elif diff <= _DURATION_LOOSE:
|
||||||
|
score += DURATION_BONUS_LOOSE
|
||||||
|
return min(score, 1.0)
|
||||||
|
|
||||||
|
|
||||||
|
def classify(song: dict, cand: dict, score: float, auto_min: float | None = None) -> str:
|
||||||
|
"""Tier for a scored candidate: 'auto' | 'review' | 'none'.
|
||||||
|
|
||||||
|
`auto` (tier-2) needs the combined score AND per-field agreement AND
|
||||||
|
both fields present — a perfect-title/wrong-artist cover, or a chart
|
||||||
|
with no artist at all, is at best a review item, never an auto match.
|
||||||
|
`auto_min` overrides the default combined-score threshold (the user's
|
||||||
|
"auto-apply confidence" setting); the per-field floors always apply.
|
||||||
|
"""
|
||||||
|
if auto_min is None:
|
||||||
|
auto_min = AUTO_MIN
|
||||||
|
artist_sim = similarity(song.get("artist"), cand.get("artist"), artist=True)
|
||||||
|
title_sim = similarity(song.get("title"), cand.get("title"))
|
||||||
|
if (score >= auto_min and artist_sim >= AUTO_ARTIST_MIN
|
||||||
|
and title_sim >= AUTO_TITLE_MIN):
|
||||||
|
return "auto"
|
||||||
|
if score >= REVIEW_MIN:
|
||||||
|
return "review"
|
||||||
|
return "none"
|
||||||
|
|
||||||
|
|
||||||
|
def rank_candidates(song: dict, candidates: list[dict]) -> list[dict]:
|
||||||
|
"""Score every candidate against the song and return them sorted by our
|
||||||
|
score (MusicBrainz's own search score is only a tiebreak). Each returned
|
||||||
|
dict is a copy carrying `score` (rounded — it's displayed and stored)."""
|
||||||
|
ranked = []
|
||||||
|
for cand in candidates or []:
|
||||||
|
c = dict(cand)
|
||||||
|
c["score"] = round(score_candidate(song, cand), 4)
|
||||||
|
ranked.append(c)
|
||||||
|
ranked.sort(key=lambda c: (c["score"], c.get("mb_score") or 0), reverse=True)
|
||||||
|
return ranked
|
||||||
|
|
||||||
|
|
||||||
|
# ── MusicBrainz query + response parsing ──────────────────────────────────────
|
||||||
|
|
||||||
|
def _lucene_escape_phrase(s: str) -> str:
|
||||||
|
"""Escape a string for use inside a quoted Lucene phrase."""
|
||||||
|
return s.replace("\\", "\\\\").replace('"', '\\"')
|
||||||
|
|
||||||
|
|
||||||
|
def build_recording_query(artist, title) -> str:
|
||||||
|
"""Lucene query for /ws/2/recording. Built from the DENOISED fields —
|
||||||
|
the noise we strip (author credits, "(Live)", "(v2)") would otherwise
|
||||||
|
poison the search server's own scoring."""
|
||||||
|
t = denoise(title)
|
||||||
|
a = denoise(artist)
|
||||||
|
parts = []
|
||||||
|
if t:
|
||||||
|
parts.append('recording:"%s"' % _lucene_escape_phrase(t))
|
||||||
|
if a:
|
||||||
|
parts.append('artist:"%s"' % _lucene_escape_phrase(a))
|
||||||
|
return " AND ".join(parts)
|
||||||
|
|
||||||
|
|
||||||
|
def _artist_credit(doc: dict) -> tuple[str, str, str]:
|
||||||
|
"""(display name, artist mbid, sort name) from an artist-credit array."""
|
||||||
|
credits = doc.get("artist-credit") or []
|
||||||
|
name = ""
|
||||||
|
for part in credits:
|
||||||
|
if isinstance(part, dict):
|
||||||
|
name += str(part.get("name", "")) + str(part.get("joinphrase", "") or "")
|
||||||
|
else: # ws/2 can emit bare join strings in older serializations
|
||||||
|
name += str(part)
|
||||||
|
first = next((p for p in credits if isinstance(p, dict)), None) or {}
|
||||||
|
artist = first.get("artist") or {}
|
||||||
|
return name, str(artist.get("id", "") or ""), str(artist.get("sort-name", "") or "")
|
||||||
|
|
||||||
|
|
||||||
|
def _best_release(doc: dict) -> dict:
|
||||||
|
"""Pick the release used for canon album/year: prefer Official status and
|
||||||
|
an Album release-group, then the earliest date. Returns {} if none."""
|
||||||
|
releases = [r for r in (doc.get("releases") or []) if isinstance(r, dict)]
|
||||||
|
if not releases:
|
||||||
|
return {}
|
||||||
|
|
||||||
|
def sort_key(r):
|
||||||
|
status_ok = 0 if str(r.get("status", "")).lower() == "official" else 1
|
||||||
|
rg = r.get("release-group") or {}
|
||||||
|
album_ok = 0 if str(rg.get("primary-type", "")).lower() == "album" else 1
|
||||||
|
date = str(r.get("date", "") or "9999")
|
||||||
|
return (status_ok, album_ok, date)
|
||||||
|
|
||||||
|
return sorted(releases, key=sort_key)[0]
|
||||||
|
|
||||||
|
|
||||||
|
def _genres(doc: dict, limit: int = 5) -> list[str]:
|
||||||
|
"""Genre names from a recording doc. Search results carry folksonomy
|
||||||
|
`tags`; lookups with inc=genres carry curated `genres`. Both are
|
||||||
|
[{name, count}] — take the most-voted few."""
|
||||||
|
raw = doc.get("genres") or doc.get("tags") or []
|
||||||
|
entries = [e for e in raw if isinstance(e, dict) and e.get("name")]
|
||||||
|
entries.sort(key=lambda e: e.get("count") or 0, reverse=True)
|
||||||
|
return [str(e["name"]) for e in entries[:limit]]
|
||||||
|
|
||||||
|
|
||||||
|
def parse_recording_doc(doc: dict) -> dict | None:
|
||||||
|
"""Normalize one /ws/2 recording document (search hit or direct lookup)
|
||||||
|
into the flat candidate dict stored in song_enrichment.candidates and
|
||||||
|
rendered by the review drawer. Returns None for malformed docs."""
|
||||||
|
if not isinstance(doc, dict) or not doc.get("id") or not doc.get("title"):
|
||||||
|
return None
|
||||||
|
artist_name, artist_id, artist_sort = _artist_credit(doc)
|
||||||
|
release = _best_release(doc)
|
||||||
|
length = doc.get("length")
|
||||||
|
try:
|
||||||
|
duration = int(round(float(length) / 1000.0)) if length else None
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
duration = None
|
||||||
|
isrcs = doc.get("isrcs") or []
|
||||||
|
isrcs = [str(i) for i in isrcs if isinstance(i, (str,))]
|
||||||
|
return {
|
||||||
|
"recording_id": str(doc["id"]),
|
||||||
|
"title": str(doc.get("title", "")),
|
||||||
|
"artist": artist_name,
|
||||||
|
"artist_id": artist_id,
|
||||||
|
"artist_sort": artist_sort,
|
||||||
|
"release_id": str(release.get("id", "") or ""),
|
||||||
|
"album": str(release.get("title", "") or ""),
|
||||||
|
"year": str(release.get("date", "") or "")[:4],
|
||||||
|
"duration": duration,
|
||||||
|
"isrc": isrcs[0] if isrcs else "",
|
||||||
|
"genres": _genres(doc),
|
||||||
|
"mb_score": int(doc.get("score") or 0),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_search_response(body: dict) -> list[dict]:
|
||||||
|
"""Candidates from a /ws/2/recording search response."""
|
||||||
|
docs = (body or {}).get("recordings") or []
|
||||||
|
out = []
|
||||||
|
for doc in docs:
|
||||||
|
cand = parse_recording_doc(doc)
|
||||||
|
if cand:
|
||||||
|
out.append(cand)
|
||||||
|
return out
|
||||||
@@ -47,6 +47,10 @@ import sloppak as sloppak_mod
|
|||||||
import drums as drums_mod
|
import drums as drums_mod
|
||||||
import notation as notation_mod
|
import notation as notation_mod
|
||||||
import loosefolder as loosefolder_mod
|
import loosefolder as loosefolder_mod
|
||||||
|
# Pure text-matching engine for MusicBrainz enrichment (P8): denoise/score/
|
||||||
|
# tier classification + response parsing. No network/DB in there — the
|
||||||
|
# throttled transport and the song_enrichment writes live in this module.
|
||||||
|
import mb_match
|
||||||
# Metadata extraction lives in a side-effect-free module so ProcessPool
|
# Metadata extraction lives in a side-effect-free module so ProcessPool
|
||||||
# scan workers can import + unpickle _scan_one without re-running this
|
# scan workers can import + unpickle _scan_one without re-running this
|
||||||
# module's import-time side effects (see lib/scan_worker.py).
|
# module's import-time side effects (see lib/scan_worker.py).
|
||||||
@@ -230,6 +234,12 @@ _DEMO_BLOCKED: list[tuple[str, re.Pattern]] = [
|
|||||||
("POST", re.compile(r"^/api/progression/events$")),
|
("POST", re.compile(r"^/api/progression/events$")),
|
||||||
("POST", re.compile(r"^/api/shop/buy$")),
|
("POST", re.compile(r"^/api/shop/buy$")),
|
||||||
("POST", re.compile(r"^/api/shop/equip$")),
|
("POST", re.compile(r"^/api/shop/equip$")),
|
||||||
|
# Enrichment (P8): review writes mutate the local match cache, and the
|
||||||
|
# search proxy / manual kick relay to MusicBrainz — none of it belongs to
|
||||||
|
# anonymous demo visitors (they'd spend the shared rate limit).
|
||||||
|
("POST", re.compile(r"^/api/enrichment/review/.+$")),
|
||||||
|
("POST", re.compile(r"^/api/enrichment/kick$")),
|
||||||
|
("GET", re.compile(r"^/api/enrichment/search$")),
|
||||||
]
|
]
|
||||||
|
|
||||||
|
|
||||||
@@ -850,6 +860,19 @@ class MetadataDB:
|
|||||||
""")
|
""")
|
||||||
self.conn.execute("CREATE INDEX IF NOT EXISTS idx_enrichment_hash ON song_enrichment(content_hash)")
|
self.conn.execute("CREATE INDEX IF NOT EXISTS idx_enrichment_hash ON song_enrichment(content_hash)")
|
||||||
self.conn.execute("CREATE INDEX IF NOT EXISTS idx_enrichment_state ON song_enrichment(match_state)")
|
self.conn.execute("CREATE INDEX IF NOT EXISTS idx_enrichment_state ON song_enrichment(match_state)")
|
||||||
|
# P8 (the matcher): `candidates` holds the review tier's ranked
|
||||||
|
# candidate list (JSON) so the Match-Review drawer never re-queries
|
||||||
|
# MusicBrainz just to render; `last_attempt_at` anchors the failed-row
|
||||||
|
# retry backoff (epoch seconds). Idempotent ALTERs, same pattern as
|
||||||
|
# the `songs` migrations above.
|
||||||
|
for ddl in (
|
||||||
|
"ALTER TABLE song_enrichment ADD COLUMN candidates TEXT",
|
||||||
|
"ALTER TABLE song_enrichment ADD COLUMN last_attempt_at REAL",
|
||||||
|
):
|
||||||
|
try:
|
||||||
|
self.conn.execute(ddl)
|
||||||
|
except sqlite3.OperationalError:
|
||||||
|
pass
|
||||||
# Progression (spec 010): instrument paths, challenges, quests, the
|
# Progression (spec 010): instrument paths, challenges, quests, the
|
||||||
# Decibels wallet, and the cosmetics shop. Targets/titles live in the
|
# Decibels wallet, and the cosmetics shop. Targets/titles live in the
|
||||||
# bundled content (data/progression/); these tables hold only player
|
# bundled content (data/progression/); these tables hold only player
|
||||||
@@ -2586,30 +2609,35 @@ class MetadataDB:
|
|||||||
return hashlib.sha1(raw.encode("utf-8")).hexdigest()
|
return hashlib.sha1(raw.encode("utf-8")).hexdigest()
|
||||||
|
|
||||||
def enrichment_pending(self, limit: int = 500) -> list[dict]:
|
def enrichment_pending(self, limit: int = 500) -> list[dict]:
|
||||||
"""Songs whose enrichment row needs (re)matching: no row yet, or an
|
"""Songs whose enrichment row needs (re)matching: no row yet, or a
|
||||||
`unscanned`/`matched` row whose content_hash no longer matches the
|
row whose content_hash no longer matches the song's current metadata
|
||||||
song's current metadata (an edit changed the identity → re-match).
|
(an edit changed the identity → re-match), or an `unscanned` row.
|
||||||
`manual` rows are the user's pinned pick and are NEVER re-queued;
|
`manual` rows are the user's pinned pick and are NEVER re-queued.
|
||||||
`failed` rows wait for the matcher's backoff policy (next slice) rather
|
`matched`/`review`/`failed` rows with an UNCHANGED hash are settled
|
||||||
than being re-queued here every pass."""
|
here — a review row stands until the user acts, and a failed row
|
||||||
|
retries only via the matcher's backoff policy (enrichment_failed_rows)
|
||||||
|
rather than being re-queued every pass. An identity edit (say, the
|
||||||
|
user fixes the typo that made matching fail) re-queues any of them
|
||||||
|
immediately via the hash mismatch."""
|
||||||
# Read under _lock: the worker commits on this shared connection under
|
# Read under _lock: the worker commits on this shared connection under
|
||||||
# _lock, so an unlocked SELECT could interleave with its execute+commit.
|
# _lock, so an unlocked SELECT could interleave with its execute+commit.
|
||||||
with self._lock:
|
with self._lock:
|
||||||
rows = self.conn.execute(
|
rows = self.conn.execute(
|
||||||
"SELECT s.filename, s.artist, s.title, s.album, s.duration, "
|
"SELECT s.filename, s.artist, s.title, s.album, s.year, s.duration, "
|
||||||
"e.content_hash, e.match_state "
|
"e.content_hash, e.match_state "
|
||||||
"FROM songs s LEFT JOIN song_enrichment e ON e.filename = s.filename "
|
"FROM songs s LEFT JOIN song_enrichment e ON e.filename = s.filename "
|
||||||
"WHERE s.title != '' AND (e.filename IS NULL OR e.match_state IN ('unscanned', 'matched')) "
|
"WHERE s.title != '' AND (e.filename IS NULL "
|
||||||
|
"OR e.match_state IN ('unscanned', 'matched', 'review', 'failed')) "
|
||||||
"ORDER BY s.filename LIMIT ?", (max(1, int(limit)),)).fetchall()
|
"ORDER BY s.filename LIMIT ?", (max(1, int(limit)),)).fetchall()
|
||||||
out = []
|
out = []
|
||||||
for fn, artist, title, album, duration, ehash, state in rows:
|
for fn, artist, title, album, year, duration, ehash, state in rows:
|
||||||
h = self.enrichment_content_hash(artist, title, album, duration)
|
h = self.enrichment_content_hash(artist, title, album, duration)
|
||||||
# No row yet, still unmatched, or the identity changed under a
|
# No row yet, still unmatched, or the identity changed under a
|
||||||
# match → needs the matcher. A matched row with an unchanged hash
|
# settled row → needs the matcher. A settled row with an
|
||||||
# is settled (idempotence).
|
# unchanged hash stays settled (idempotence).
|
||||||
if state is None or state == "unscanned" or ehash != h:
|
if state is None or state == "unscanned" or ehash != h:
|
||||||
out.append({"filename": fn, "artist": artist, "title": title,
|
out.append({"filename": fn, "artist": artist, "title": title,
|
||||||
"album": album, "duration": duration,
|
"album": album, "year": year, "duration": duration,
|
||||||
"content_hash": h, "match_state": state})
|
"content_hash": h, "match_state": state})
|
||||||
return out
|
return out
|
||||||
|
|
||||||
@@ -2641,6 +2669,14 @@ class MetadataDB:
|
|||||||
" WHEN song_enrichment.content_hash IS NOT excluded.content_hash "
|
" WHEN song_enrichment.content_hash IS NOT excluded.content_hash "
|
||||||
" THEN 'unscanned' "
|
" THEN 'unscanned' "
|
||||||
" ELSE song_enrichment.match_state END, "
|
" ELSE song_enrichment.match_state END, "
|
||||||
|
# An identity change restarts the failure backoff too — the
|
||||||
|
# accumulated attempts belonged to the OLD identity (e.g. the
|
||||||
|
# user just fixed the typo that made matching fail).
|
||||||
|
" attempts = CASE WHEN song_enrichment.match_state = 'manual' "
|
||||||
|
" THEN song_enrichment.attempts "
|
||||||
|
" WHEN song_enrichment.content_hash IS NOT excluded.content_hash "
|
||||||
|
" THEN 0 "
|
||||||
|
" ELSE song_enrichment.attempts END, "
|
||||||
" content_hash = CASE WHEN song_enrichment.match_state = 'manual' "
|
" content_hash = CASE WHEN song_enrichment.match_state = 'manual' "
|
||||||
" THEN song_enrichment.content_hash "
|
" THEN song_enrichment.content_hash "
|
||||||
" ELSE excluded.content_hash END",
|
" ELSE excluded.content_hash END",
|
||||||
@@ -2654,19 +2690,21 @@ class MetadataDB:
|
|||||||
"SELECT filename, content_hash, match_state, match_source, match_score, attempts, "
|
"SELECT filename, content_hash, match_state, match_source, match_score, attempts, "
|
||||||
"mb_recording_id, mb_release_id, mb_artist_id, isrc, "
|
"mb_recording_id, mb_release_id, mb_artist_id, isrc, "
|
||||||
"canon_artist, canon_album, canon_title, canon_year, canon_artist_sort, "
|
"canon_artist, canon_album, canon_title, canon_year, canon_artist_sort, "
|
||||||
"genres, art_cache_path, art_state, fetched_at "
|
"genres, art_cache_path, art_state, fetched_at, candidates, last_attempt_at "
|
||||||
"FROM song_enrichment WHERE filename = ?", (filename,)).fetchone()
|
"FROM song_enrichment WHERE filename = ?", (filename,)).fetchone()
|
||||||
if not row:
|
if not row:
|
||||||
return None
|
return None
|
||||||
keys = ("filename", "content_hash", "match_state", "match_source", "match_score",
|
keys = ("filename", "content_hash", "match_state", "match_source", "match_score",
|
||||||
"attempts", "mb_recording_id", "mb_release_id", "mb_artist_id", "isrc",
|
"attempts", "mb_recording_id", "mb_release_id", "mb_artist_id", "isrc",
|
||||||
"canon_artist", "canon_album", "canon_title", "canon_year",
|
"canon_artist", "canon_album", "canon_title", "canon_year",
|
||||||
"canon_artist_sort", "genres", "art_cache_path", "art_state", "fetched_at")
|
"canon_artist_sort", "genres", "art_cache_path", "art_state", "fetched_at",
|
||||||
|
"candidates", "last_attempt_at")
|
||||||
out = dict(zip(keys, row))
|
out = dict(zip(keys, row))
|
||||||
try:
|
for k in ("genres", "candidates"):
|
||||||
out["genres"] = json.loads(out["genres"]) if out["genres"] else []
|
try:
|
||||||
except (ValueError, TypeError):
|
out[k] = json.loads(out[k]) if out[k] else []
|
||||||
out["genres"] = []
|
except (ValueError, TypeError):
|
||||||
|
out[k] = []
|
||||||
return out
|
return out
|
||||||
|
|
||||||
def enrichment_state_counts(self) -> dict:
|
def enrichment_state_counts(self) -> dict:
|
||||||
@@ -2679,6 +2717,160 @@ class MetadataDB:
|
|||||||
"JOIN songs s ON s.filename = e.filename GROUP BY e.match_state").fetchall()
|
"JOIN songs s ON s.filename = e.filename GROUP BY e.match_state").fetchall()
|
||||||
return {r[0]: r[1] for r in rows}
|
return {r[0]: r[1] for r in rows}
|
||||||
|
|
||||||
|
def enrichment_song_row(self, filename: str) -> dict | None:
|
||||||
|
"""The identity fields the matcher/scorer keys on, for one song."""
|
||||||
|
row = self.conn.execute(
|
||||||
|
"SELECT filename, artist, title, album, year, duration "
|
||||||
|
"FROM songs WHERE filename = ?", (filename,)).fetchone()
|
||||||
|
if not row:
|
||||||
|
return None
|
||||||
|
return dict(zip(("filename", "artist", "title", "album", "year", "duration"), row))
|
||||||
|
|
||||||
|
def enrichment_failed_rows(self, limit: int = 500) -> list[dict]:
|
||||||
|
"""`failed` rows that MAY retry, with the fields the backoff policy
|
||||||
|
(worker-side) needs to decide eligibility. `rejected` rows are the
|
||||||
|
user's explicit "none of these" — never auto-retried (an identity
|
||||||
|
edit re-queues them through enrichment_pending's hash mismatch
|
||||||
|
instead)."""
|
||||||
|
rows = self.conn.execute(
|
||||||
|
"SELECT s.filename, s.artist, s.title, s.album, s.year, s.duration, "
|
||||||
|
"e.attempts, e.last_attempt_at "
|
||||||
|
"FROM songs s JOIN song_enrichment e ON e.filename = s.filename "
|
||||||
|
"WHERE s.title != '' AND e.match_state = 'failed' "
|
||||||
|
"AND COALESCE(e.match_source, '') != 'rejected' "
|
||||||
|
"ORDER BY s.filename LIMIT ?", (max(1, int(limit)),)).fetchall()
|
||||||
|
out = []
|
||||||
|
for fn, artist, title, album, year, duration, attempts, last_at in rows:
|
||||||
|
out.append({"filename": fn, "artist": artist, "title": title,
|
||||||
|
"album": album, "year": year, "duration": duration,
|
||||||
|
"content_hash": self.enrichment_content_hash(artist, title, album, duration),
|
||||||
|
"attempts": attempts or 0, "last_attempt_at": last_at})
|
||||||
|
return out
|
||||||
|
|
||||||
|
def enrichment_cache_lookup(self, content_hash: str, exclude_filename: str = "") -> dict | None:
|
||||||
|
"""A settled match for the same identity hash — another chart of the
|
||||||
|
same recording already matched/pinned → copy it, no network (design
|
||||||
|
§5 step 1: the local match-cache)."""
|
||||||
|
row = self.conn.execute(
|
||||||
|
"SELECT match_score, mb_recording_id, mb_release_id, mb_artist_id, isrc, "
|
||||||
|
"canon_artist, canon_album, canon_title, canon_year, canon_artist_sort, genres "
|
||||||
|
"FROM song_enrichment WHERE content_hash = ? AND filename != ? "
|
||||||
|
"AND match_state IN ('matched', 'manual') AND mb_recording_id IS NOT NULL "
|
||||||
|
"LIMIT 1", (content_hash, exclude_filename or "")).fetchone()
|
||||||
|
if not row:
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
genres = json.loads(row[10]) if row[10] else []
|
||||||
|
except (ValueError, TypeError):
|
||||||
|
genres = []
|
||||||
|
return {
|
||||||
|
"score": row[0],
|
||||||
|
"recording_id": row[1], "release_id": row[2] or "", "artist_id": row[3] or "",
|
||||||
|
"isrc": row[4] or "", "artist": row[5] or "", "album": row[6] or "",
|
||||||
|
"title": row[7] or "", "year": row[8] or "", "artist_sort": row[9] or "",
|
||||||
|
"genres": genres,
|
||||||
|
}
|
||||||
|
|
||||||
|
def apply_enrichment_match(self, filename: str, content_hash: str, state: str,
|
||||||
|
source: str | None = None, score: float | None = None,
|
||||||
|
cand: dict | None = None, candidates: list | None = None,
|
||||||
|
bump_attempts: bool = False,
|
||||||
|
allow_manual_overwrite: bool = False) -> bool:
|
||||||
|
"""The single writer for every matcher/review outcome. Writes the
|
||||||
|
full lifecycle row: state + source + score, the canonical fields a
|
||||||
|
confident match supplies (`cand`), and/or the review tier's ranked
|
||||||
|
`candidates`. Returns False without touching anything when the row is
|
||||||
|
`manual` and the caller isn't explicitly acting for the user — the
|
||||||
|
never-overwrite-manual contract lives HERE so no future call path
|
||||||
|
can forget it. Art-cache fields are preserved verbatim (they belong
|
||||||
|
to the art slice, not the matcher)."""
|
||||||
|
cand = cand or {}
|
||||||
|
now = time.time()
|
||||||
|
with self._lock:
|
||||||
|
cur = self.conn.execute(
|
||||||
|
"SELECT match_state, attempts, art_cache_path, art_state, fetched_at "
|
||||||
|
"FROM song_enrichment WHERE filename = ?", (filename,)).fetchone()
|
||||||
|
if cur and cur[0] == "manual" and not allow_manual_overwrite:
|
||||||
|
return False
|
||||||
|
attempts = int(cur[1] or 0) if cur else 0
|
||||||
|
if bump_attempts:
|
||||||
|
attempts += 1
|
||||||
|
fetched_at = (time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
|
||||||
|
if state in ("matched", "manual", "review")
|
||||||
|
else (cur[4] if cur else None))
|
||||||
|
self.conn.execute(
|
||||||
|
"INSERT OR REPLACE INTO song_enrichment (filename, content_hash, "
|
||||||
|
"match_state, match_source, match_score, attempts, "
|
||||||
|
"mb_recording_id, mb_release_id, mb_artist_id, isrc, "
|
||||||
|
"canon_artist, canon_album, canon_title, canon_year, canon_artist_sort, "
|
||||||
|
"genres, art_cache_path, art_state, fetched_at, candidates, last_attempt_at) "
|
||||||
|
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
|
||||||
|
(filename, content_hash, state, source, score, attempts,
|
||||||
|
cand.get("recording_id") or None, cand.get("release_id") or None,
|
||||||
|
cand.get("artist_id") or None, cand.get("isrc") or None,
|
||||||
|
cand.get("artist") or None, cand.get("album") or None,
|
||||||
|
cand.get("title") or None, cand.get("year") or None,
|
||||||
|
cand.get("artist_sort") or None,
|
||||||
|
json.dumps(cand.get("genres") or []) if cand else "[]",
|
||||||
|
cur[2] if cur else None, cur[3] if cur else None,
|
||||||
|
fetched_at,
|
||||||
|
json.dumps(candidates) if candidates else None,
|
||||||
|
now if state == "failed" else None))
|
||||||
|
self.conn.commit()
|
||||||
|
return True
|
||||||
|
|
||||||
|
def set_enrichment_manual(self, filename: str, cand: dict, source: str = "search") -> bool:
|
||||||
|
"""User-pinned match (review Accept / manual search-and-pick). The
|
||||||
|
highest-authority state: never auto-reset, survives identity edits.
|
||||||
|
`source` records HOW it was pinned ('review' = accepted a proposed
|
||||||
|
candidate, 'search' = picked from a manual search)."""
|
||||||
|
song = self.enrichment_song_row(filename)
|
||||||
|
if not song:
|
||||||
|
return False
|
||||||
|
h = self.enrichment_content_hash(
|
||||||
|
song["artist"], song["title"], song["album"], song["duration"])
|
||||||
|
return self.apply_enrichment_match(
|
||||||
|
filename, h, "manual", source=source, score=1.0, cand=cand,
|
||||||
|
allow_manual_overwrite=True)
|
||||||
|
|
||||||
|
def set_enrichment_rejected(self, filename: str) -> bool:
|
||||||
|
"""User said "none of these candidates" — clear any canonical values
|
||||||
|
and park the row as failed/rejected (never auto-retried; an identity
|
||||||
|
edit re-queues it). Refused for `manual` rows: un-pinning a pick the
|
||||||
|
user explicitly made is not a review-drawer action."""
|
||||||
|
row = self.get_enrichment(filename)
|
||||||
|
if not row or row["match_state"] not in ("review", "matched"):
|
||||||
|
return False
|
||||||
|
return self.apply_enrichment_match(
|
||||||
|
filename, row["content_hash"], "failed", source="rejected",
|
||||||
|
score=None, candidates=row.get("candidates") or None)
|
||||||
|
|
||||||
|
def enrichment_review_queue(self, limit: int = 200) -> list[dict]:
|
||||||
|
"""The Match-Review drawer's queue: review-tier rows joined to their
|
||||||
|
(still-existing) songs, with the stored candidate list parsed."""
|
||||||
|
rows = self.conn.execute(
|
||||||
|
"SELECT e.filename, s.title, s.artist, s.album, s.year, s.duration, s.mtime, "
|
||||||
|
"e.match_score, e.candidates, e.attempts "
|
||||||
|
"FROM song_enrichment e JOIN songs s ON s.filename = e.filename "
|
||||||
|
"WHERE e.match_state = 'review' "
|
||||||
|
# Charts that are MISSING data (no album / no year) surface first —
|
||||||
|
# confirming those has the most to gain; complete charts only
|
||||||
|
# stand to be re-labelled.
|
||||||
|
"ORDER BY ((COALESCE(s.album, '') = '') + (COALESCE(s.year, '') = '')) DESC, "
|
||||||
|
"s.artist COLLATE NOCASE, s.title COLLATE NOCASE, e.filename "
|
||||||
|
"LIMIT ?", (max(1, int(limit)),)).fetchall()
|
||||||
|
out = []
|
||||||
|
for fn, title, artist, album, year, duration, mtime, score, cands, attempts in rows:
|
||||||
|
try:
|
||||||
|
candidates = json.loads(cands) if cands else []
|
||||||
|
except (ValueError, TypeError):
|
||||||
|
candidates = []
|
||||||
|
out.append({"filename": fn, "title": title, "artist": artist,
|
||||||
|
"album": album, "year": year, "duration": duration,
|
||||||
|
"mtime": mtime, "match_score": score,
|
||||||
|
"candidates": candidates, "attempts": attempts or 0})
|
||||||
|
return out
|
||||||
|
|
||||||
def _estd_set(self) -> set[str]:
|
def _estd_set(self) -> set[str]:
|
||||||
"""Get set of filenames that have a retuned variant (_EStd_ or _DropD_) in the DB."""
|
"""Get set of filenames that have a retuned variant (_EStd_ or _DropD_) in the DB."""
|
||||||
rows = self.conn.execute(
|
rows = self.conn.execute(
|
||||||
@@ -2733,6 +2925,7 @@ class MetadataDB:
|
|||||||
mastery: list[str] | None = None,
|
mastery: list[str] | None = None,
|
||||||
tags_has: list[str] | None = None,
|
tags_has: list[str] | None = None,
|
||||||
user_difficulty_in: list[str] | None = None,
|
user_difficulty_in: list[str] | None = None,
|
||||||
|
match_states: list[str] | None = None,
|
||||||
naming_mode: str = "legacy",
|
naming_mode: str = "legacy",
|
||||||
include_intrinsic: bool = True) -> tuple[str, list]:
|
include_intrinsic: bool = True) -> tuple[str, list]:
|
||||||
"""Shared WHERE-clause builder for query_page / query_artists /
|
"""Shared WHERE-clause builder for query_page / query_artists /
|
||||||
@@ -2798,6 +2991,21 @@ class MetadataDB:
|
|||||||
where += (" AND filename IN (SELECT filename FROM song_user_meta "
|
where += (" AND filename IN (SELECT filename FROM song_user_meta "
|
||||||
f"WHERE user_difficulty IN ({ph}))")
|
f"WHERE user_difficulty IN ({ph}))")
|
||||||
params += _diffs
|
params += _diffs
|
||||||
|
# Match facet (P8) = the song's enrichment lifecycle state, from the
|
||||||
|
# separate song_enrichment table (same EXISTS idiom as mastery above).
|
||||||
|
# 'matched' folds in 'manual' (a user pin IS a match); 'pending' means
|
||||||
|
# no verdict yet (no row, or still unscanned). OR within the set.
|
||||||
|
if match_states:
|
||||||
|
_esub = "SELECT 1 FROM song_enrichment e WHERE e.filename = songs.filename"
|
||||||
|
_mstates = {
|
||||||
|
"review": f"EXISTS ({_esub} AND e.match_state = 'review')",
|
||||||
|
"matched": f"EXISTS ({_esub} AND e.match_state IN ('matched', 'manual'))",
|
||||||
|
"unmatched": f"EXISTS ({_esub} AND e.match_state = 'failed')",
|
||||||
|
"pending": f"NOT EXISTS ({_esub} AND e.match_state != 'unscanned')",
|
||||||
|
}
|
||||||
|
_msel = [_mstates[b] for b in match_states if b in _mstates]
|
||||||
|
if _msel:
|
||||||
|
where += " AND (" + " OR ".join(_msel) + ")"
|
||||||
if q:
|
if q:
|
||||||
where += " AND (title LIKE ? COLLATE NOCASE OR artist LIKE ? COLLATE NOCASE OR album LIKE ? COLLATE NOCASE)"
|
where += " AND (title LIKE ? COLLATE NOCASE OR artist LIKE ? COLLATE NOCASE OR album LIKE ? COLLATE NOCASE)"
|
||||||
params += [f"%{q}%"] * 3
|
params += [f"%{q}%"] * 3
|
||||||
@@ -3228,6 +3436,7 @@ class MetadataDB:
|
|||||||
mastery: list[str] | None = None,
|
mastery: list[str] | None = None,
|
||||||
tags_has: list[str] | None = None,
|
tags_has: list[str] | None = None,
|
||||||
user_difficulty_in: list[str] | None = None,
|
user_difficulty_in: list[str] | None = None,
|
||||||
|
match_states: list[str] | None = None,
|
||||||
after: str | None = None,
|
after: str | None = None,
|
||||||
group: bool = False,
|
group: bool = False,
|
||||||
naming_mode: str = "legacy") -> tuple[list[dict], int]:
|
naming_mode: str = "legacy") -> tuple[list[dict], int]:
|
||||||
@@ -3258,6 +3467,7 @@ class MetadataDB:
|
|||||||
stems_has=stems_has, stems_lacks=stems_lacks,
|
stems_has=stems_has, stems_lacks=stems_lacks,
|
||||||
has_lyrics=has_lyrics, tunings=tunings, mastery=mastery,
|
has_lyrics=has_lyrics, tunings=tunings, mastery=mastery,
|
||||||
tags_has=tags_has, user_difficulty_in=user_difficulty_in,
|
tags_has=tags_has, user_difficulty_in=user_difficulty_in,
|
||||||
|
match_states=match_states,
|
||||||
naming_mode=naming_mode, include_intrinsic=not group,
|
naming_mode=naming_mode, include_intrinsic=not group,
|
||||||
)
|
)
|
||||||
ifrag, iparams = "", []
|
ifrag, iparams = "", []
|
||||||
@@ -3595,6 +3805,7 @@ class MetadataDB:
|
|||||||
stems_lacks: list[str] | None = None,
|
stems_lacks: list[str] | None = None,
|
||||||
has_lyrics: int | None = None,
|
has_lyrics: int | None = None,
|
||||||
tunings: list[str] | None = None,
|
tunings: list[str] | None = None,
|
||||||
|
match_states: list[str] | None = None,
|
||||||
sort: str = "artist",
|
sort: str = "artist",
|
||||||
want_sort_letters: bool = False,
|
want_sort_letters: bool = False,
|
||||||
group: bool = False,
|
group: bool = False,
|
||||||
@@ -3624,7 +3835,8 @@ class MetadataDB:
|
|||||||
artist_filter=artist_filter, album_filter=album_filter,
|
artist_filter=artist_filter, album_filter=album_filter,
|
||||||
arrangements_has=arrangements_has, arrangements_lacks=arrangements_lacks,
|
arrangements_has=arrangements_has, arrangements_lacks=arrangements_lacks,
|
||||||
stems_has=stems_has, stems_lacks=stems_lacks,
|
stems_has=stems_has, stems_lacks=stems_lacks,
|
||||||
has_lyrics=has_lyrics, tunings=tunings, naming_mode=naming_mode,
|
has_lyrics=has_lyrics, tunings=tunings, match_states=match_states,
|
||||||
|
naming_mode=naming_mode,
|
||||||
include_intrinsic=not group,
|
include_intrinsic=not group,
|
||||||
)
|
)
|
||||||
if group:
|
if group:
|
||||||
@@ -4390,7 +4602,8 @@ def _require_library_provider_capability(provider: object, capability: str) -> N
|
|||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
_OPTIONAL_NEW_PROVIDER_KWARGS = ("naming_mode", "sort", "want_sort_letters", "after")
|
_OPTIONAL_NEW_PROVIDER_KWARGS = ("naming_mode", "sort", "want_sort_letters", "after",
|
||||||
|
"mastery", "match_states")
|
||||||
|
|
||||||
|
|
||||||
def _filter_provider_kwargs(method: object, kwargs: dict) -> dict:
|
def _filter_provider_kwargs(method: object, kwargs: dict) -> dict:
|
||||||
@@ -5229,13 +5442,14 @@ def _scan_runner():
|
|||||||
_kick_enrich()
|
_kick_enrich()
|
||||||
|
|
||||||
|
|
||||||
# ── Metadata enrichment worker (P7 — plumbing) ────────────────────────────────
|
# ── Metadata enrichment worker (P7 plumbing + P8 matcher) ─────────────────────
|
||||||
# A single throttled daemon thread + queue, mirroring _kick_scan/_scan_runner
|
# A single throttled daemon thread + queue, mirroring _kick_scan/_scan_runner
|
||||||
# (single-flight + coalescing; NOT a pool — external lookups are rate-limited
|
# (single-flight + coalescing; NOT a pool — external lookups are rate-limited
|
||||||
# to ~1/s, which makes a pool pointless). This slice ships the full lifecycle
|
# to ~1/s, which makes a pool pointless). P7 shipped the lifecycle; P8 fills
|
||||||
# around a NO-OP matcher: the queue walk, the identity hashing, the throttle
|
# in the matcher (_enrich_one): local cache → manifest mbid/isrc exact keys →
|
||||||
# seam, and the status surface — so the real text matcher (next slice) replaces
|
# MusicBrainz text search, scored into auto/review/failed tiers by
|
||||||
# exactly one function (_enrich_one) and inherits everything else.
|
# lib/mb_match.py. Wrong-match is worse than slow (design §5): medium
|
||||||
|
# confidence goes to the Match-Review queue, never straight to canonical.
|
||||||
|
|
||||||
_enrich_kick_lock = threading.Lock()
|
_enrich_kick_lock = threading.Lock()
|
||||||
_enrich_pending_pass = False
|
_enrich_pending_pass = False
|
||||||
@@ -5243,6 +5457,9 @@ _enrich_status = {"running": False, "processed": 0, "last_pass_at": None}
|
|||||||
# Minimum spacing between EXTERNAL lookups (design: ≤1 req/s + local cache).
|
# Minimum spacing between EXTERNAL lookups (design: ≤1 req/s + local cache).
|
||||||
_ENRICH_MIN_INTERVAL = 1.1
|
_ENRICH_MIN_INTERVAL = 1.1
|
||||||
_enrich_last_fetch = 0.0
|
_enrich_last_fetch = 0.0
|
||||||
|
# Serializes throttling across the background daemon thread AND the sync
|
||||||
|
# /api/enrichment/search route (FastAPI runs sync routes in a threadpool).
|
||||||
|
_enrich_throttle_lock = threading.Lock()
|
||||||
|
|
||||||
|
|
||||||
def _enrichment_art_dir() -> Path:
|
def _enrichment_art_dir() -> Path:
|
||||||
@@ -5259,25 +5476,233 @@ def _enrich_throttle():
|
|||||||
before every network request — and must NOT hold meta_db._lock across the
|
before every network request — and must NOT hold meta_db._lock across the
|
||||||
request (fetch outside the lock, write inside)."""
|
request (fetch outside the lock, write inside)."""
|
||||||
global _enrich_last_fetch
|
global _enrich_last_fetch
|
||||||
wait = _ENRICH_MIN_INTERVAL - (time.monotonic() - _enrich_last_fetch)
|
# Hold the lock across the read, sleep, and write so concurrent callers
|
||||||
if wait > 0:
|
# serialize instead of all reading the same stale timestamp and firing
|
||||||
time.sleep(wait)
|
# together (which would burst past MusicBrainz's 1 req/s limit).
|
||||||
_enrich_last_fetch = time.monotonic()
|
with _enrich_throttle_lock:
|
||||||
|
wait = _ENRICH_MIN_INTERVAL - (time.monotonic() - _enrich_last_fetch)
|
||||||
|
if wait > 0:
|
||||||
|
time.sleep(wait)
|
||||||
|
_enrich_last_fetch = time.monotonic()
|
||||||
|
|
||||||
|
|
||||||
def _enrich_one(row: dict) -> None:
|
class EnrichTransportError(Exception):
|
||||||
"""P7's NO-OP matcher: stamp/refresh the row's identity hash (which also
|
"""Network-level enrichment failure — offline, DNS, MusicBrainz down or
|
||||||
drops a stale match back to `unscanned`, never a `manual` pick) and stop.
|
rate-limiting. Pauses the current pass (rows keep their state and no
|
||||||
The text-match pipeline replaces this function; anything that reaches the
|
attempt is consumed); the next kick (scan-complete / the 5-min periodic
|
||||||
network must go through _enrich_throttle() and must not hold meta_db._lock
|
rescan) retries naturally."""
|
||||||
across the fetch."""
|
|
||||||
meta_db.upsert_enrichment_stub(row["filename"], row["content_hash"])
|
|
||||||
|
_MB_API_ROOT = "https://musicbrainz.org/ws/2"
|
||||||
|
_enrich_ua_cache: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
def _enrich_user_agent() -> str:
|
||||||
|
"""MusicBrainz etiquette requires a real identifying User-Agent
|
||||||
|
(app/version + contact URL); anonymous defaults get throttled/blocked."""
|
||||||
|
global _enrich_ua_cache
|
||||||
|
if _enrich_ua_cache is None:
|
||||||
|
version = "unknown"
|
||||||
|
try:
|
||||||
|
vf = Path(__file__).parent / "VERSION"
|
||||||
|
if vf.exists():
|
||||||
|
version = vf.read_text().strip() or "unknown"
|
||||||
|
except (OSError, UnicodeDecodeError):
|
||||||
|
pass
|
||||||
|
_enrich_ua_cache = f"feedBack/{version} (https://github.com/got-feedback/feedBack)"
|
||||||
|
return _enrich_ua_cache
|
||||||
|
|
||||||
|
|
||||||
|
def _enrich_network_enabled() -> bool:
|
||||||
|
"""False = the matcher runs local-only (hash stamping, cache copies) and
|
||||||
|
never opens a socket. FEEDBACK_ENRICH_OFFLINE is the explicit user
|
||||||
|
kill-switch (privacy / air-gapped installs); FEEDBACK_SKIP_STARTUP_TASKS
|
||||||
|
marks the test/CI environment, where pytest must never reach the network
|
||||||
|
no matter what a test triggers."""
|
||||||
|
return not (_env_flag("FEEDBACK_ENRICH_OFFLINE")
|
||||||
|
or _env_flag("FEEDBACK_SKIP_STARTUP_TASKS"))
|
||||||
|
|
||||||
|
|
||||||
|
def _mb_http_get(path: str, params: dict) -> dict | None:
|
||||||
|
"""The ONE place enrichment touches the network (tests fake exactly this
|
||||||
|
seam). Throttled (≤1 req/s via _enrich_throttle), identified (real
|
||||||
|
User-Agent), offline-guarded. Returns the parsed JSON body, or None for
|
||||||
|
a 404 lookup; raises EnrichTransportError for anything network-shaped.
|
||||||
|
NEVER call this while holding meta_db._lock — fetch outside, write
|
||||||
|
inside."""
|
||||||
|
if not _enrich_network_enabled():
|
||||||
|
raise EnrichTransportError("enrichment network disabled")
|
||||||
|
import requests # declared in requirements.txt; lazy so tests never need it
|
||||||
|
_enrich_throttle()
|
||||||
|
try:
|
||||||
|
resp = requests.get(
|
||||||
|
f"{_MB_API_ROOT}/{path.lstrip('/')}",
|
||||||
|
params={**params, "fmt": "json"},
|
||||||
|
headers={"User-Agent": _enrich_user_agent()},
|
||||||
|
timeout=10,
|
||||||
|
)
|
||||||
|
except requests.RequestException as e:
|
||||||
|
raise EnrichTransportError(str(e)) from e
|
||||||
|
if resp.status_code == 404:
|
||||||
|
return None
|
||||||
|
if resp.status_code == 503:
|
||||||
|
# MusicBrainz signals rate-limit pressure with 503 — back the whole
|
||||||
|
# pass off rather than hammering on.
|
||||||
|
raise EnrichTransportError("musicbrainz 503 (rate limited)")
|
||||||
|
if resp.status_code != 200:
|
||||||
|
raise EnrichTransportError(f"musicbrainz HTTP {resp.status_code}")
|
||||||
|
try:
|
||||||
|
return resp.json()
|
||||||
|
except ValueError as e:
|
||||||
|
raise EnrichTransportError("bad JSON from musicbrainz") from e
|
||||||
|
|
||||||
|
|
||||||
|
def _mb_search_recordings(artist, title, limit: int = 8) -> list[dict]:
|
||||||
|
"""Text search (tier 2–4): denoised Lucene query over /recording."""
|
||||||
|
query = mb_match.build_recording_query(artist, title)
|
||||||
|
if not query:
|
||||||
|
return []
|
||||||
|
body = _mb_http_get("recording", {"query": query, "limit": limit})
|
||||||
|
return mb_match.parse_search_response(body or {})
|
||||||
|
|
||||||
|
|
||||||
|
def _mb_lookup_recording(mbid: str) -> dict | None:
|
||||||
|
"""Direct lookup for a manifest-carried recording MBID (tier 0)."""
|
||||||
|
body = _mb_http_get(
|
||||||
|
f"recording/{mbid}",
|
||||||
|
{"inc": "artist-credits+releases+release-groups+isrcs+genres"})
|
||||||
|
return mb_match.parse_recording_doc(body) if body else None
|
||||||
|
|
||||||
|
|
||||||
|
def _mb_lookup_isrc(isrc: str) -> list[dict]:
|
||||||
|
"""Recordings registered under a manifest-carried ISRC (tier 1)."""
|
||||||
|
body = _mb_http_get(
|
||||||
|
f"isrc/{isrc}", {"inc": "artist-credits+releases+release-groups"})
|
||||||
|
if not body:
|
||||||
|
return []
|
||||||
|
docs = body.get("recordings") or []
|
||||||
|
return [c for c in (mb_match.parse_recording_doc(d) for d in docs) if c]
|
||||||
|
|
||||||
|
|
||||||
|
# Strict shapes for the manifest's optional identity keys (feedpak spec §5.1).
|
||||||
|
# Validated before use — the mbid is interpolated into a URL path, so junk or
|
||||||
|
# hostile manifest values must never reach the request line.
|
||||||
|
_MBID_RE = re.compile(r"^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$")
|
||||||
|
_ISRC_RE = re.compile(r"^[A-Z]{2}[A-Z0-9]{3}[0-9]{7}$")
|
||||||
|
|
||||||
|
|
||||||
|
def _manifest_exact_ids(filename: str) -> dict:
|
||||||
|
"""Optional `mbid`/`isrc` from the pack manifest — the spec's additive
|
||||||
|
identity keys. Feature-detected: packs published before that spec
|
||||||
|
revision simply lack them and fall through to text matching. READ-only:
|
||||||
|
enrichment never writes anything into pack files."""
|
||||||
|
try:
|
||||||
|
dlc = _get_dlc_dir()
|
||||||
|
if not dlc:
|
||||||
|
return {}
|
||||||
|
p = _resolve_dlc_path(dlc, filename)
|
||||||
|
if p is None or not p.exists() or not sloppak_mod.is_sloppak(p):
|
||||||
|
return {}
|
||||||
|
manifest = sloppak_mod.load_manifest(p) or {}
|
||||||
|
except Exception:
|
||||||
|
return {}
|
||||||
|
out = {}
|
||||||
|
mbid = str(manifest.get("mbid", "") or "").strip().lower()
|
||||||
|
if _MBID_RE.match(mbid):
|
||||||
|
out["mbid"] = mbid
|
||||||
|
isrc = str(manifest.get("isrc", "") or "").strip().upper()
|
||||||
|
if _ISRC_RE.match(isrc):
|
||||||
|
out["isrc"] = isrc
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
# Failed-row retry backoff: 1 h after the first failed attempt, doubling per
|
||||||
|
# attempt, capped at a week — a permanently-unmatchable obscure chart must
|
||||||
|
# not re-hammer MusicBrainz on every scan kick.
|
||||||
|
_ENRICH_BACKOFF_BASE = 3600.0
|
||||||
|
_ENRICH_BACKOFF_CAP = 7 * 86400.0
|
||||||
|
|
||||||
|
|
||||||
|
def _enrich_backoff_elapsed(attempts, last_attempt_at, now: float) -> bool:
|
||||||
|
if not last_attempt_at:
|
||||||
|
return True
|
||||||
|
delay = min(_ENRICH_BACKOFF_BASE * (2 ** max(0, int(attempts or 1) - 1)),
|
||||||
|
_ENRICH_BACKOFF_CAP)
|
||||||
|
return (now - float(last_attempt_at)) >= delay
|
||||||
|
|
||||||
|
|
||||||
|
# Review tier keeps a short ranked candidate list for the drawer; more than a
|
||||||
|
# handful is noise the user has to scroll past.
|
||||||
|
_ENRICH_MAX_CANDIDATES = 5
|
||||||
|
|
||||||
|
|
||||||
|
def _enrich_one(row: dict, auto_min: float | None = None) -> None:
|
||||||
|
"""The matcher (P8; replaces P7's no-op). Precedence per design §5:
|
||||||
|
|
||||||
|
1. local match-cache by content_hash — another chart of the same
|
||||||
|
recording already matched/pinned → copy it, NO network;
|
||||||
|
2. manifest `mbid` (tier 0) / `isrc` (tier 1) exact keys → direct
|
||||||
|
lookup, auto;
|
||||||
|
3. text search → scored tiers: auto (high) / review (medium — a human
|
||||||
|
confirms before anything canonicalizes) / failed (low, retried on
|
||||||
|
backoff).
|
||||||
|
|
||||||
|
`auto_min` is the user's auto-apply confidence setting (None → the
|
||||||
|
engine default); it moves only the auto/review boundary of step 3 —
|
||||||
|
the per-field floors and exact-key tiers are unaffected. Never touches
|
||||||
|
a `manual` row (the writer enforces it). Network errors raise
|
||||||
|
EnrichTransportError so the pass pauses instead of burning attempts
|
||||||
|
while offline."""
|
||||||
|
fn, chash = row["filename"], row["content_hash"]
|
||||||
|
|
||||||
|
cached = meta_db.enrichment_cache_lookup(chash, exclude_filename=fn)
|
||||||
|
if cached:
|
||||||
|
score = cached.pop("score", None)
|
||||||
|
meta_db.apply_enrichment_match(fn, chash, "matched", source="cache",
|
||||||
|
score=score, cand=cached)
|
||||||
|
return
|
||||||
|
|
||||||
|
ids = _manifest_exact_ids(fn)
|
||||||
|
if ids.get("mbid"):
|
||||||
|
cand = _mb_lookup_recording(ids["mbid"])
|
||||||
|
if cand:
|
||||||
|
meta_db.apply_enrichment_match(fn, chash, "matched", source="mbid",
|
||||||
|
score=1.0, cand=cand)
|
||||||
|
return
|
||||||
|
# A 404'd mbid (typo'd manifest) falls through to the text tiers.
|
||||||
|
if ids.get("isrc"):
|
||||||
|
cands = mb_match.rank_candidates(row, _mb_lookup_isrc(ids["isrc"]))
|
||||||
|
if cands:
|
||||||
|
meta_db.apply_enrichment_match(fn, chash, "matched", source="isrc",
|
||||||
|
score=1.0, cand=cands[0])
|
||||||
|
return
|
||||||
|
|
||||||
|
ranked = mb_match.rank_candidates(row, _mb_search_recordings(row.get("artist"), row.get("title")))
|
||||||
|
best = ranked[0] if ranked else None
|
||||||
|
tier = mb_match.classify(row, best, best["score"], auto_min=auto_min) if best else "none"
|
||||||
|
if tier == "auto":
|
||||||
|
meta_db.apply_enrichment_match(fn, chash, "matched", source="text",
|
||||||
|
score=best["score"], cand=best)
|
||||||
|
elif tier == "review":
|
||||||
|
meta_db.apply_enrichment_match(fn, chash, "review", source="text",
|
||||||
|
score=best["score"],
|
||||||
|
candidates=ranked[:_ENRICH_MAX_CANDIDATES])
|
||||||
|
else:
|
||||||
|
meta_db.apply_enrichment_match(fn, chash, "failed", source="text",
|
||||||
|
score=(best["score"] if best else None),
|
||||||
|
candidates=ranked[:_ENRICH_MAX_CANDIDATES] or None,
|
||||||
|
bump_attempts=True)
|
||||||
|
|
||||||
|
|
||||||
def _background_enrich():
|
def _background_enrich():
|
||||||
"""One pass over the rows needing (re)matching. A single bounded pass —
|
"""One bounded pass, two phases. Phase 1 stamps/refreshes identity-hash
|
||||||
with the no-op matcher, `unscanned` rows legitimately stay unscanned, so
|
stubs for every song whose identity is new or changed — pure-local, so
|
||||||
looping until the queue drains would spin forever."""
|
hashes stay fresh (and stale matches drop back to `unscanned`) even
|
||||||
|
fully offline. Phase 2 runs the matcher over those rows plus any
|
||||||
|
`failed` rows whose backoff has elapsed; a transport failure pauses it
|
||||||
|
(state untouched, no attempt burned) and the next kick retries. Offline
|
||||||
|
(kill-switch or the test env) skips phase 2 entirely. Never drains in a
|
||||||
|
loop — a dead network would make that spin forever."""
|
||||||
_enrich_status["processed"] = 0
|
_enrich_status["processed"] = 0
|
||||||
try:
|
try:
|
||||||
pending = meta_db.enrichment_pending(limit=100000)
|
pending = meta_db.enrichment_pending(limit=100000)
|
||||||
@@ -5286,13 +5711,66 @@ def _background_enrich():
|
|||||||
return
|
return
|
||||||
for row in pending:
|
for row in pending:
|
||||||
try:
|
try:
|
||||||
_enrich_one(row)
|
meta_db.upsert_enrichment_stub(row["filename"], row["content_hash"])
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
log.warning("enrichment failed for %s: %s", row.get("filename"), e)
|
log.warning("enrichment stub failed for %s: %s", row.get("filename"), e)
|
||||||
_enrich_status["processed"] += 1
|
_enrich_status["processed"] += 1
|
||||||
_enrich_status["last_pass_at"] = time.time()
|
_enrich_status["last_pass_at"] = time.time()
|
||||||
if pending:
|
|
||||||
log.info("Enrichment pass: %d rows refreshed", len(pending))
|
# User settings gate the BACKGROUND matcher only (the review modal's
|
||||||
|
# manual search/fix stays available when it's off); read once per pass.
|
||||||
|
cfg = _load_config(CONFIG_DIR / "config.json") or {}
|
||||||
|
if cfg.get("enrich_enabled", True) is False:
|
||||||
|
if pending:
|
||||||
|
log.info("Enrichment pass: %d rows stamped (matching disabled in Settings)", len(pending))
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
auto_min = float(cfg.get("enrich_auto_threshold", 0.9))
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
auto_min = 0.9
|
||||||
|
|
||||||
|
if not _enrich_network_enabled():
|
||||||
|
if pending:
|
||||||
|
log.info("Enrichment pass: %d rows stamped (network disabled — matching skipped)", len(pending))
|
||||||
|
return
|
||||||
|
|
||||||
|
now = time.time()
|
||||||
|
retriable = []
|
||||||
|
try:
|
||||||
|
retriable = [r for r in meta_db.enrichment_failed_rows(limit=100000)
|
||||||
|
if _enrich_backoff_elapsed(r.get("attempts"), r.get("last_attempt_at"), now)]
|
||||||
|
except Exception:
|
||||||
|
log.exception("enrichment: failed-row query failed")
|
||||||
|
matched = 0
|
||||||
|
# A `failed` row with a changed identity hash can surface in BOTH lists;
|
||||||
|
# de-dup by filename so each row consumes the rate budget only once.
|
||||||
|
seen_filenames = set()
|
||||||
|
queue = []
|
||||||
|
for row in pending + retriable:
|
||||||
|
fn = row.get("filename")
|
||||||
|
if fn in seen_filenames:
|
||||||
|
continue
|
||||||
|
seen_filenames.add(fn)
|
||||||
|
queue.append(row)
|
||||||
|
for row in queue:
|
||||||
|
try:
|
||||||
|
_enrich_one(row, auto_min=auto_min)
|
||||||
|
matched += 1
|
||||||
|
except EnrichTransportError as e:
|
||||||
|
log.info("enrichment: network unavailable, pass paused (%s)", e)
|
||||||
|
break
|
||||||
|
except Exception as e:
|
||||||
|
log.warning("enrichment failed for %s: %s", row.get("filename"), e)
|
||||||
|
try:
|
||||||
|
# Park the row on the failure backoff instead of retrying a
|
||||||
|
# poisoned input every pass.
|
||||||
|
meta_db.apply_enrichment_match(
|
||||||
|
row["filename"], row["content_hash"], "failed",
|
||||||
|
source="error", bump_attempts=True)
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
if pending or retriable:
|
||||||
|
log.info("Enrichment pass: %d rows stamped, %d matched", len(pending), matched)
|
||||||
|
|
||||||
|
|
||||||
def _kick_enrich() -> bool:
|
def _kick_enrich() -> bool:
|
||||||
@@ -5807,6 +6285,112 @@ def enrichment_status():
|
|||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/api/enrichment/kick")
|
||||||
|
def api_enrichment_kick():
|
||||||
|
"""The Settings "Match now" button: request an enrichment pass without
|
||||||
|
waiting for a scan to complete. Single-flight + coalescing like every
|
||||||
|
other kick — spamming it queues at most one follow-up pass."""
|
||||||
|
return {"started": _kick_enrich()}
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/api/enrichment/review")
|
||||||
|
def api_enrichment_review(limit: int = 200):
|
||||||
|
"""The Match-Review queue: songs whose text match landed in the medium-
|
||||||
|
confidence review tier, each with its stored candidate list — the drawer
|
||||||
|
renders straight from this, no MusicBrainz round-trip."""
|
||||||
|
limit = max(1, min(int(limit), 500))
|
||||||
|
return {
|
||||||
|
"songs": meta_db.enrichment_review_queue(limit=limit),
|
||||||
|
"total_review": meta_db.enrichment_state_counts().get("review", 0),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/api/enrichment/review/{filename:path}/accept")
|
||||||
|
def api_enrichment_accept(filename: str, data: dict = Body(...)):
|
||||||
|
"""Accept one of the stored review candidates: the row becomes a
|
||||||
|
user-pinned `manual` match (never auto-reset). Display-only, like every
|
||||||
|
enrichment write — nothing touches the pack file."""
|
||||||
|
recording_id = str((data or {}).get("recording_id") or "")
|
||||||
|
row = meta_db.get_enrichment(filename)
|
||||||
|
if not row or row["match_state"] != "review":
|
||||||
|
raise HTTPException(status_code=404, detail="no review row for this song")
|
||||||
|
cand = next((c for c in (row.get("candidates") or [])
|
||||||
|
if c.get("recording_id") == recording_id), None)
|
||||||
|
if not cand:
|
||||||
|
raise HTTPException(status_code=404, detail="candidate not in the stored list")
|
||||||
|
if not meta_db.set_enrichment_manual(filename, cand, source="review"):
|
||||||
|
raise HTTPException(status_code=404, detail="unknown song")
|
||||||
|
return {"ok": True, "enrichment": meta_db.get_enrichment(filename)}
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/api/enrichment/review/{filename:path}/reject")
|
||||||
|
def api_enrichment_reject(filename: str):
|
||||||
|
""""None of these" — clears any canonical values and parks the row as
|
||||||
|
failed/rejected (never auto-retried; editing the song's metadata
|
||||||
|
re-queues it). Valid from `review` or `matched`, never from `manual`."""
|
||||||
|
if not meta_db.set_enrichment_rejected(filename):
|
||||||
|
raise HTTPException(status_code=404, detail="no rejectable match for this song")
|
||||||
|
return {"ok": True, "enrichment": meta_db.get_enrichment(filename)}
|
||||||
|
|
||||||
|
|
||||||
|
# The candidate fields a manual pick is allowed to carry — the payload comes
|
||||||
|
# from our own /api/enrichment/search proxy, but the route re-sanitizes so a
|
||||||
|
# hand-rolled client can't stuff arbitrary keys/types into the cache row.
|
||||||
|
_CAND_STR_FIELDS = ("recording_id", "title", "artist", "artist_id",
|
||||||
|
"artist_sort", "release_id", "album", "year", "isrc")
|
||||||
|
|
||||||
|
|
||||||
|
def _sanitize_candidate(raw: dict) -> dict | None:
|
||||||
|
if not isinstance(raw, dict):
|
||||||
|
return None
|
||||||
|
out = {k: str(raw.get(k) or "") for k in _CAND_STR_FIELDS}
|
||||||
|
if not out["recording_id"] or not out["title"]:
|
||||||
|
return None
|
||||||
|
genres = raw.get("genres") or []
|
||||||
|
out["genres"] = [str(g) for g in genres if isinstance(g, str)][:5] \
|
||||||
|
if isinstance(genres, list) else []
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/api/enrichment/review/{filename:path}/pick")
|
||||||
|
def api_enrichment_pick(filename: str, data: dict = Body(...)):
|
||||||
|
"""Fix-match / manual search-and-pick: pin a candidate the user found via
|
||||||
|
/api/enrichment/search (not limited to the stored review list — this is
|
||||||
|
the escape hatch for a wrong auto-match too). Sets `manual`, the
|
||||||
|
highest-authority state."""
|
||||||
|
cand = _sanitize_candidate((data or {}).get("candidate"))
|
||||||
|
if not cand:
|
||||||
|
raise HTTPException(status_code=400, detail="candidate needs recording_id + title")
|
||||||
|
if not meta_db.set_enrichment_manual(filename, cand, source="search"):
|
||||||
|
raise HTTPException(status_code=404, detail="unknown song")
|
||||||
|
return {"ok": True, "enrichment": meta_db.get_enrichment(filename)}
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/api/enrichment/search")
|
||||||
|
def api_enrichment_search(artist: str = "", title: str = "", limit: int = 8,
|
||||||
|
filename: str = ""):
|
||||||
|
"""Manual-search proxy to MusicBrainz (throttled + identified like the
|
||||||
|
background matcher — a user typing in the drawer must not sidestep the
|
||||||
|
rate limit). `filename` optionally scores results against that song's
|
||||||
|
stored identity (year/duration corroboration) instead of just the typed
|
||||||
|
text. Sync route on purpose: FastAPI runs it in the threadpool, so the
|
||||||
|
throttle's sleep never blocks the event loop."""
|
||||||
|
if not (artist.strip() or title.strip()):
|
||||||
|
raise HTTPException(status_code=400, detail="artist or title required")
|
||||||
|
limit = max(1, min(int(limit), 25))
|
||||||
|
try:
|
||||||
|
cands = _mb_search_recordings(artist, title, limit=limit)
|
||||||
|
except EnrichTransportError as e:
|
||||||
|
return JSONResponse({"error": "musicbrainz unavailable", "detail": str(e)},
|
||||||
|
status_code=503)
|
||||||
|
ref = None
|
||||||
|
if filename:
|
||||||
|
ref = meta_db.enrichment_song_row(filename)
|
||||||
|
if ref is None:
|
||||||
|
ref = {"artist": artist, "title": title}
|
||||||
|
return {"candidates": mb_match.rank_candidates(ref, cands)}
|
||||||
|
|
||||||
|
|
||||||
@app.get("/api/startup-status")
|
@app.get("/api/startup-status")
|
||||||
def startup_status():
|
def startup_status():
|
||||||
return _get_startup_status()
|
return _get_startup_status()
|
||||||
@@ -6405,7 +6989,8 @@ async def list_library(q: str = "", page: int = 0, size: int = 24, sort: str = "
|
|||||||
stems_has: str = "", stems_lacks: str = "",
|
stems_has: str = "", stems_lacks: str = "",
|
||||||
has_lyrics: str = "", tunings: str = "", provider: str = "local",
|
has_lyrics: str = "", tunings: str = "", provider: str = "local",
|
||||||
mastery: str = "", tags: str = "", user_difficulty: str = "",
|
mastery: str = "", tags: str = "", user_difficulty: str = "",
|
||||||
after: str = "", group: int = 0, naming_mode: str = "legacy"):
|
match: str = "", after: str = "", group: int = 0,
|
||||||
|
naming_mode: str = "legacy"):
|
||||||
"""Paginated library search through the selected library provider.
|
"""Paginated library search through the selected library provider.
|
||||||
|
|
||||||
`after` is an opaque keyset cursor (feedBack#636 item 3): pass back the
|
`after` is an opaque keyset cursor (feedBack#636 item 3): pass back the
|
||||||
@@ -6433,6 +7018,7 @@ async def list_library(q: str = "", page: int = 0, size: int = 24, sort: str = "
|
|||||||
mastery=_split_csv(mastery),
|
mastery=_split_csv(mastery),
|
||||||
tags_has=_split_csv(tags),
|
tags_has=_split_csv(tags),
|
||||||
user_difficulty_in=_split_csv(user_difficulty),
|
user_difficulty_in=_split_csv(user_difficulty),
|
||||||
|
match_states=_split_csv(match),
|
||||||
**_library_filter_args(
|
**_library_filter_args(
|
||||||
q=q, favorites=favorites, format=format,
|
q=q, favorites=favorites, format=format,
|
||||||
artist=artist, album=album,
|
artist=artist, album=album,
|
||||||
@@ -6544,6 +7130,7 @@ async def library_stats(favorites: int = 0, q: str = "", format: str = "",
|
|||||||
arrangements_has: str = "", arrangements_lacks: str = "",
|
arrangements_has: str = "", arrangements_lacks: str = "",
|
||||||
stems_has: str = "", stems_lacks: str = "",
|
stems_has: str = "", stems_lacks: str = "",
|
||||||
has_lyrics: str = "", tunings: str = "", provider: str = "local",
|
has_lyrics: str = "", tunings: str = "", provider: str = "local",
|
||||||
|
match: str = "",
|
||||||
sort: str = "artist", sort_letters: int = 0,
|
sort: str = "artist", sort_letters: int = 0,
|
||||||
group: int = 0, naming_mode: str = "legacy"):
|
group: int = 0, naming_mode: str = "legacy"):
|
||||||
"""Aggregate stats for the UI. Accepts the same filter params as
|
"""Aggregate stats for the UI. Accepts the same filter params as
|
||||||
@@ -6561,6 +7148,10 @@ async def library_stats(favorites: int = 0, q: str = "", format: str = "",
|
|||||||
sort=sort,
|
sort=sort,
|
||||||
want_sort_letters=bool(sort_letters),
|
want_sort_letters=bool(sort_letters),
|
||||||
group=bool(group),
|
group=bool(group),
|
||||||
|
# The match facet rides the stats call too — the A–Z rail's letter
|
||||||
|
# counts must agree with the grid under the facet or its cumulative
|
||||||
|
# seek + sizer geometry break.
|
||||||
|
match_states=_split_csv(match),
|
||||||
**_library_filter_args(
|
**_library_filter_args(
|
||||||
q=q, favorites=favorites, format=format,
|
q=q, favorites=favorites, format=format,
|
||||||
artist=artist, album=album,
|
artist=artist, album=album,
|
||||||
@@ -7862,6 +8453,17 @@ def _default_settings():
|
|||||||
# renderer to gate its saved-chain restore. Inert on the pure-web build,
|
# renderer to gate its saved-chain restore. Inert on the pure-web build,
|
||||||
# which has no native amp sims.
|
# which has no native amp sims.
|
||||||
"use_amp_sims": False,
|
"use_amp_sims": False,
|
||||||
|
# Metadata matching (P8). `enrich_enabled` gates only the BACKGROUND
|
||||||
|
# matcher — manual Fix-match/search in the review modal keeps working
|
||||||
|
# when it's off (the media-server model: scraper off ≠ no manual fix);
|
||||||
|
# the FEEDBACK_ENRICH_OFFLINE env var is the hard everything-off kill.
|
||||||
|
# `enrich_auto_threshold` is the auto-apply confidence — matches at or
|
||||||
|
# above it canonicalize automatically, below it queue for review. The
|
||||||
|
# per-field floors in lib/mb_match.py always apply on top, so lowering
|
||||||
|
# this can't make a wrong-artist cover auto-match. >1.0 (the "Always
|
||||||
|
# review" option) sends every text match to review.
|
||||||
|
"enrich_enabled": True,
|
||||||
|
"enrich_auto_threshold": 0.9,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
@@ -8010,6 +8612,28 @@ def save_settings(data: dict):
|
|||||||
if not isinstance(raw, bool):
|
if not isinstance(raw, bool):
|
||||||
return {"error": "use_amp_sims must be a boolean"}
|
return {"error": "use_amp_sims must be a boolean"}
|
||||||
updates["use_amp_sims"] = raw
|
updates["use_amp_sims"] = raw
|
||||||
|
if "enrich_enabled" in data:
|
||||||
|
raw = data["enrich_enabled"]
|
||||||
|
if raw is not None:
|
||||||
|
if not isinstance(raw, bool):
|
||||||
|
return {"error": "enrich_enabled must be a boolean"}
|
||||||
|
updates["enrich_enabled"] = raw
|
||||||
|
if "enrich_auto_threshold" in data:
|
||||||
|
# Auto-apply confidence for the metadata matcher. 0.5–1.0 are real
|
||||||
|
# thresholds; values just above 1.0 are the "Always review" option (a
|
||||||
|
# capped score can equal exactly 1.0, so "never auto" must sit above
|
||||||
|
# the cap). Same defensive coercion shape as av_offset_ms.
|
||||||
|
raw = data["enrich_auto_threshold"]
|
||||||
|
if raw is not None:
|
||||||
|
if isinstance(raw, bool):
|
||||||
|
return {"error": "enrich_auto_threshold must be a number between 0.5 and 1.01"}
|
||||||
|
try:
|
||||||
|
t = float(raw)
|
||||||
|
except (TypeError, ValueError, OverflowError):
|
||||||
|
return {"error": "enrich_auto_threshold must be a number between 0.5 and 1.01"}
|
||||||
|
if not math.isfinite(t) or not (0.5 <= t <= 1.01):
|
||||||
|
return {"error": "enrich_auto_threshold must be a number between 0.5 and 1.01"}
|
||||||
|
updates["enrich_auto_threshold"] = t
|
||||||
if "miss_penalty" in data:
|
if "miss_penalty" in data:
|
||||||
raw = data["miss_penalty"]
|
raw = data["miss_penalty"]
|
||||||
if raw is not None:
|
if raw is not None:
|
||||||
|
|||||||
Vendored
+1
-1
File diff suppressed because one or more lines are too long
@@ -681,6 +681,28 @@
|
|||||||
<span id="rescan-status" class="text-xs text-gray-500"></span>
|
<span id="rescan-status" class="text-xs text-gray-500"></span>
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
|
<!-- Metadata matching (P8 — wired by static/v3/match-review.js) -->
|
||||||
|
<div class="fb-srow fb-srow-stack">
|
||||||
|
<div class="fb-srow-main">
|
||||||
|
<div class="fb-srow-title">Metadata matching</div>
|
||||||
|
<div class="fb-srow-desc">Matches your charts against MusicBrainz in the background to tidy names, years and genres — display only, your files are never modified. Matches at or above the confidence level apply automatically; the rest wait in the library's review queue. Turning this off never disables manual match fixes.</div>
|
||||||
|
</div>
|
||||||
|
<div class="grid grid-cols-2 gap-2 mb-1 text-xs text-gray-400 fb-srow-wide">
|
||||||
|
<label class="flex items-center gap-2"><input type="checkbox" id="enrich-enabled" checked class="rounded border-gray-600 bg-dark-700 text-accent"> Match songs against MusicBrainz</label>
|
||||||
|
<label class="flex items-center gap-2">Auto-apply confidence
|
||||||
|
<select id="enrich-threshold" class="bg-dark-700 border border-gray-800 rounded-xl px-2 py-1.5 text-xs text-gray-300 outline-none">
|
||||||
|
<option value="0.85">85% — more auto-matches</option>
|
||||||
|
<option value="0.9" selected>90% (recommended)</option>
|
||||||
|
<option value="0.95">95% — cautious</option>
|
||||||
|
<option value="1.01">Always review</option>
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
|
</div>
|
||||||
|
<div class="fb-srow-control">
|
||||||
|
<button id="enrich-match-now" class="bg-dark-600 hover:bg-dark-500 px-5 py-2.5 rounded-xl text-sm text-gray-300 transition">Match Now</button>
|
||||||
|
<span id="enrich-status" class="text-xs text-gray-500"></span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
<!-- Backup -->
|
<!-- Backup -->
|
||||||
<div class="fb-srow fb-srow-stack">
|
<div class="fb-srow fb-srow-stack">
|
||||||
<div class="fb-srow-main">
|
<div class="fb-srow-main">
|
||||||
@@ -1118,6 +1140,9 @@
|
|||||||
<script src="/static/v3/pedal-cables.js"></script>
|
<script src="/static/v3/pedal-cables.js"></script>
|
||||||
<script src="/static/v3/plugins-page.js"></script>
|
<script src="/static/v3/plugins-page.js"></script>
|
||||||
<script src="/static/v3/card-actions-core.js"></script>
|
<script src="/static/v3/card-actions-core.js"></script>
|
||||||
|
<!-- Before songs.js: the songs toolbar calls the match-review chip hook
|
||||||
|
on build, so the module must already be registered. -->
|
||||||
|
<script src="/static/v3/match-review.js"></script>
|
||||||
<script src="/static/v3/songs.js"></script>
|
<script src="/static/v3/songs.js"></script>
|
||||||
<script src="/static/v3/lessons.js"></script>
|
<script src="/static/v3/lessons.js"></script>
|
||||||
<script src="/static/v3/dashboard.js"></script>
|
<script src="/static/v3/dashboard.js"></script>
|
||||||
|
|||||||
@@ -0,0 +1,414 @@
|
|||||||
|
// Match-Review UI (P8 — library-metadata design §5/§11). A self-contained
|
||||||
|
// module: the ambient "⚑ N to review" chip lives in the songs toolbar
|
||||||
|
// (songs.js renders the element and calls the hooks below); the review MODAL,
|
||||||
|
// the per-field available/missing detail, and the Settings → Library
|
||||||
|
// "Metadata matching" card behaviour all live here.
|
||||||
|
//
|
||||||
|
// The modal reviews ONE chart at a time (the scraper-review model from
|
||||||
|
// media-server / emulation-frontend apps): the chart's current metadata —
|
||||||
|
// with explicit "Missing: …" chips — above the candidate list, each
|
||||||
|
// candidate carrying "Adds / Shows as" chips, with Skip / Not a match /
|
||||||
|
// Search instead / Use selected plus ‹ › navigation.
|
||||||
|
//
|
||||||
|
// Engagement guardrails (§11): opt-in tool-state, not a score. The chip only
|
||||||
|
// appears when there is something to review, matching is silent on success
|
||||||
|
// (no toasts, no sounds — hearing-safe), and nothing here ever writes to
|
||||||
|
// pack files; a confirmed match only improves the local display cache.
|
||||||
|
(function () {
|
||||||
|
'use strict';
|
||||||
|
|
||||||
|
const esc = (s) => String(s == null ? '' : s).replace(/[&<>"']/g, (c) => (
|
||||||
|
{ '&': '&', '<': '<', '>': '>', '"': '"', "'": ''' }[c]));
|
||||||
|
const enc = encodeURIComponent;
|
||||||
|
|
||||||
|
function artUrl(song) {
|
||||||
|
const v = song.mtime ? ('?v=' + Math.floor(song.mtime)) : '';
|
||||||
|
return '/api/song/' + enc(song.filename) + '/art' + v;
|
||||||
|
}
|
||||||
|
|
||||||
|
function fmtDur(sec) {
|
||||||
|
if (!sec && sec !== 0) return '';
|
||||||
|
const s = Math.max(0, Math.round(sec));
|
||||||
|
return Math.floor(s / 60) + ':' + String(s % 60).padStart(2, '0');
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Ambient chip + the Settings card's status line ───────────────────────
|
||||||
|
// songs.js renders `#v3-songs-match-review` (hidden) in its toolbar and
|
||||||
|
// calls window.__fbMatchReviewChip() after each toolbar build; review
|
||||||
|
// actions here re-call it. The same fetch feeds the Settings status line.
|
||||||
|
// Silent on failure — surfaces just stay as they are.
|
||||||
|
let _chipBusy = false;
|
||||||
|
async function refreshChip() {
|
||||||
|
if (_chipBusy) return;
|
||||||
|
_chipBusy = true;
|
||||||
|
try {
|
||||||
|
const r = await fetch('/api/enrichment/status');
|
||||||
|
if (!r.ok) return;
|
||||||
|
const body = await r.json();
|
||||||
|
const st = body.states || {};
|
||||||
|
const n = st.review || 0;
|
||||||
|
const chip = document.getElementById('v3-songs-match-review');
|
||||||
|
if (chip) {
|
||||||
|
chip.textContent = '⚑ ' + n + ' to review';
|
||||||
|
chip.classList.toggle('hidden', !n);
|
||||||
|
}
|
||||||
|
const line = document.getElementById('enrich-status');
|
||||||
|
if (line) {
|
||||||
|
const parts = [
|
||||||
|
((st.matched || 0) + (st.manual || 0)) + ' matched',
|
||||||
|
n + ' to review',
|
||||||
|
(st.failed || 0) + ' unmatched',
|
||||||
|
];
|
||||||
|
if (st.unscanned) parts.push(st.unscanned + ' queued');
|
||||||
|
line.textContent = (body.running ? 'Matching… · ' : '') + parts.join(' · ');
|
||||||
|
}
|
||||||
|
} catch (_) { /* offline — leave as-is */ } finally {
|
||||||
|
_chipBusy = false;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Review modal (body-appended singleton, one chart at a time) ─────────
|
||||||
|
let _queue = [];
|
||||||
|
let _idx = 0;
|
||||||
|
let _lastFocus = null;
|
||||||
|
|
||||||
|
function ensureModal() {
|
||||||
|
let m = document.getElementById('v3-match-modal');
|
||||||
|
if (m) return m;
|
||||||
|
const overlay = document.createElement('div');
|
||||||
|
overlay.id = 'v3-match-overlay';
|
||||||
|
overlay.className = 'fixed inset-0 bg-black/60 z-40 hidden';
|
||||||
|
overlay.addEventListener('click', closeModal);
|
||||||
|
document.body.appendChild(overlay);
|
||||||
|
m = document.createElement('div');
|
||||||
|
m.id = 'v3-match-modal';
|
||||||
|
m.className = 'fixed inset-0 z-50 hidden flex items-center justify-center p-4 pointer-events-none';
|
||||||
|
m.innerHTML = '<div id="v3-match-panel" class="pointer-events-auto w-full max-w-2xl max-h-[85vh] bg-fb-sidebar border border-fb-border/50 rounded-xl shadow-2xl flex flex-col" role="dialog" aria-label="Match review"></div>';
|
||||||
|
m.addEventListener('keydown', onModalKeydown);
|
||||||
|
document.body.appendChild(m);
|
||||||
|
return m;
|
||||||
|
}
|
||||||
|
|
||||||
|
function isTyping(e) {
|
||||||
|
const t = e.target;
|
||||||
|
return t && (t.tagName === 'INPUT' || t.tagName === 'TEXTAREA');
|
||||||
|
}
|
||||||
|
|
||||||
|
function onModalKeydown(e) {
|
||||||
|
if (e.key === 'Escape') { e.stopPropagation(); closeModal(); return; }
|
||||||
|
if (e.key === 'ArrowLeft' && !isTyping(e)) { e.preventDefault(); nav(-1); return; }
|
||||||
|
if (e.key === 'ArrowRight' && !isTyping(e)) { e.preventDefault(); nav(1); return; }
|
||||||
|
if (e.key !== 'Tab') return;
|
||||||
|
// Light focus trap: cycle within the panel.
|
||||||
|
const panel = document.getElementById('v3-match-panel');
|
||||||
|
if (!panel) return;
|
||||||
|
const foci = panel.querySelectorAll('button, input, [tabindex="0"]');
|
||||||
|
if (!foci.length) return;
|
||||||
|
const first = foci[0], last = foci[foci.length - 1];
|
||||||
|
if (e.shiftKey && document.activeElement === first) { e.preventDefault(); last.focus(); }
|
||||||
|
else if (!e.shiftKey && document.activeElement === last) { e.preventDefault(); first.focus(); }
|
||||||
|
}
|
||||||
|
|
||||||
|
function openModal() {
|
||||||
|
_lastFocus = document.activeElement;
|
||||||
|
const m = ensureModal();
|
||||||
|
renderLoading();
|
||||||
|
m.classList.remove('hidden');
|
||||||
|
document.getElementById('v3-match-overlay')?.classList.remove('hidden');
|
||||||
|
loadQueue();
|
||||||
|
}
|
||||||
|
|
||||||
|
function closeModal() {
|
||||||
|
document.getElementById('v3-match-modal')?.classList.add('hidden');
|
||||||
|
document.getElementById('v3-match-overlay')?.classList.add('hidden');
|
||||||
|
refreshChip();
|
||||||
|
if (_lastFocus && _lastFocus.isConnected) { try { _lastFocus.focus(); } catch (_) { } }
|
||||||
|
_lastFocus = null;
|
||||||
|
}
|
||||||
|
|
||||||
|
function nav(step) {
|
||||||
|
if (!_queue.length) return;
|
||||||
|
_idx = Math.min(Math.max(_idx + step, 0), _queue.length - 1);
|
||||||
|
renderCurrent();
|
||||||
|
}
|
||||||
|
|
||||||
|
async function loadQueue() {
|
||||||
|
try {
|
||||||
|
const r = await fetch('/api/enrichment/review?limit=200');
|
||||||
|
_queue = r.ok ? ((await r.json()).songs || []) : [];
|
||||||
|
} catch (_) { _queue = []; }
|
||||||
|
_idx = 0;
|
||||||
|
renderCurrent();
|
||||||
|
}
|
||||||
|
|
||||||
|
function headerHtml() {
|
||||||
|
const counter = _queue.length
|
||||||
|
? '<span class="flex items-center gap-1 text-xs text-fb-textDim">' +
|
||||||
|
'<button data-mr-prev class="px-2 py-1 rounded hover:text-fb-text' + (_idx === 0 ? ' opacity-30' : '') + '" aria-label="Previous">‹</button>' +
|
||||||
|
(_idx + 1) + ' of ' + _queue.length +
|
||||||
|
'<button data-mr-next class="px-2 py-1 rounded hover:text-fb-text' + (_idx >= _queue.length - 1 ? ' opacity-30' : '') + '" aria-label="Next">›</button></span>'
|
||||||
|
: '';
|
||||||
|
return '<div class="flex items-center justify-between gap-3 p-5 pb-3 border-b border-fb-border/40 shrink-0">' +
|
||||||
|
'<h3 class="text-lg font-semibold text-fb-text">Match review</h3>' + counter +
|
||||||
|
'<button data-mr-close class="text-fb-textDim hover:text-fb-text" aria-label="Close">✕</button></div>';
|
||||||
|
}
|
||||||
|
|
||||||
|
function renderLoading() {
|
||||||
|
const panel = document.getElementById('v3-match-panel');
|
||||||
|
if (!panel) return;
|
||||||
|
panel.innerHTML = headerHtml() +
|
||||||
|
'<div class="p-5"><p class="text-sm text-fb-textDim">Loading…</p></div>';
|
||||||
|
panel.querySelector('[data-mr-close]')?.addEventListener('click', closeModal);
|
||||||
|
}
|
||||||
|
|
||||||
|
function renderDone() {
|
||||||
|
const panel = document.getElementById('v3-match-panel');
|
||||||
|
if (!panel) return;
|
||||||
|
panel.innerHTML = headerHtml() +
|
||||||
|
'<div class="p-5 space-y-2"><p class="text-sm text-fb-text">Nothing waiting for review.</p>' +
|
||||||
|
'<p class="text-xs text-fb-textDim">Medium-confidence matches queue here while the library is matched in the background. Matching options live in Settings → Library.</p></div>';
|
||||||
|
panel.querySelector('[data-mr-close]')?.addEventListener('click', closeModal);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Amber "what this chart lacks" chips. Album/year come from the library
|
||||||
|
// row; cover art is detected from the art request failing (flagged onto
|
||||||
|
// the song object by the <img> onerror handler, then re-rendered).
|
||||||
|
function missingChips(song) {
|
||||||
|
const missing = [];
|
||||||
|
if (!String(song.album || '').trim()) missing.push('album');
|
||||||
|
if (!String(song.year || '').trim()) missing.push('year');
|
||||||
|
if (song._artMissing) missing.push('cover art');
|
||||||
|
if (!missing.length) return '';
|
||||||
|
return '<div class="flex flex-wrap items-center gap-1 pt-1">' +
|
||||||
|
'<span class="text-xs text-fb-textDim">Missing:</span>' +
|
||||||
|
missing.map((f) => '<span class="text-xs px-1.5 py-0.5 rounded border border-amber-400/40 text-amber-300/90 bg-amber-400/10">' + esc(f) + '</span>').join('') +
|
||||||
|
'</div>';
|
||||||
|
}
|
||||||
|
|
||||||
|
// Per-candidate "what accepting this gets you": fields the chart lacks
|
||||||
|
// that the candidate supplies, and fields whose DISPLAYED value would
|
||||||
|
// change (never the file).
|
||||||
|
function diffChips(song, cand) {
|
||||||
|
const adds = [];
|
||||||
|
const changes = [];
|
||||||
|
const have = (v) => String(v == null ? '' : v).trim();
|
||||||
|
const differ = (a, b) => have(a) && have(b) && have(a).toLowerCase() !== have(b).toLowerCase();
|
||||||
|
if (have(cand.album)) { if (!have(song.album)) adds.push('album'); else if (differ(song.album, cand.album)) changes.push('album'); }
|
||||||
|
if (have(cand.year)) { if (!have(song.year)) adds.push('year'); else if (differ(song.year, cand.year)) changes.push('year'); }
|
||||||
|
if (cand.genres && cand.genres.length) adds.push('genres');
|
||||||
|
if (have(cand.isrc)) adds.push('ISRC');
|
||||||
|
if (differ(song.artist, cand.artist)) changes.push(have(song.artist) + ' → ' + have(cand.artist));
|
||||||
|
if (differ(song.title, cand.title)) changes.push('title');
|
||||||
|
let html = '';
|
||||||
|
if (adds.length) html += '<span class="text-xs text-fb-good">Adds: ' + esc(adds.join(' · ')) + '</span>';
|
||||||
|
if (changes.length) html += (html ? ' ' : '') + '<span class="text-xs text-fb-textDim">Shows as: ' + esc(changes.join(' · ')) + '</span>';
|
||||||
|
return html ? '<span class="block truncate pt-0.5">' + html + '</span>' : '';
|
||||||
|
}
|
||||||
|
|
||||||
|
function candRowHtml(song, c, i, selected) {
|
||||||
|
const meta = [c.artist, c.album, c.year, fmtDur(c.duration)].filter(Boolean).join(' · ');
|
||||||
|
const pct = c.score != null ? Math.round(c.score * 100) + '%' : '';
|
||||||
|
return '<button data-mr-cand="' + i + '" role="radio" aria-checked="' + (selected ? 'true' : 'false') + '" class="w-full text-left px-3 py-2 rounded-md border ' +
|
||||||
|
(selected ? 'border-fb-primary bg-fb-primary/10' : 'border-fb-border/50 bg-gray-800/50 hover:border-fb-primary/60') + '">' +
|
||||||
|
'<span class="flex items-baseline justify-between gap-2">' +
|
||||||
|
'<span class="text-sm text-fb-text truncate">' + esc(c.title) + '</span>' +
|
||||||
|
'<span class="text-xs text-fb-textDim shrink-0">' + esc(pct) + '</span></span>' +
|
||||||
|
'<span class="block text-xs text-fb-textDim truncate">' + esc(meta) + '</span>' +
|
||||||
|
diffChips(song, c) +
|
||||||
|
'</button>';
|
||||||
|
}
|
||||||
|
|
||||||
|
function renderCurrent() {
|
||||||
|
const panel = document.getElementById('v3-match-panel');
|
||||||
|
if (!panel) return;
|
||||||
|
if (!_queue.length) { renderDone(); return; }
|
||||||
|
_idx = Math.min(_idx, _queue.length - 1);
|
||||||
|
const song = _queue[_idx];
|
||||||
|
if (song._sel == null) song._sel = 0;
|
||||||
|
const sub = [song.artist, song.album, song.year, fmtDur(song.duration)].filter(Boolean).join(' · ');
|
||||||
|
|
||||||
|
panel.innerHTML = headerHtml() +
|
||||||
|
'<div class="p-5 space-y-4 overflow-y-auto v3-scroll">' +
|
||||||
|
// The chart being matched
|
||||||
|
'<div class="flex items-start gap-3">' +
|
||||||
|
'<img data-mr-art src="' + esc(artUrl(song)) + '" alt="" loading="lazy" class="w-16 h-16 rounded-lg object-cover bg-fb-card shrink-0">' +
|
||||||
|
'<div class="min-w-0">' +
|
||||||
|
'<div class="text-base text-fb-text font-medium truncate">' + esc(song.title) + '</div>' +
|
||||||
|
'<div class="text-xs text-fb-textDim truncate">' + esc(sub) + '</div>' +
|
||||||
|
'<div class="text-xs text-fb-textDim/70 truncate" title="' + esc(song.filename) + '">' + esc(song.filename) + '</div>' +
|
||||||
|
missingChips(song) +
|
||||||
|
'</div></div>' +
|
||||||
|
// Candidates
|
||||||
|
'<div class="space-y-1" role="radiogroup" aria-label="Candidates">' +
|
||||||
|
'<div class="text-xs font-semibold uppercase tracking-wider text-fb-textDim">Candidates (MusicBrainz)</div>' +
|
||||||
|
(song.candidates || []).map((c, i) => candRowHtml(song, c, i, i === song._sel)).join('') +
|
||||||
|
'</div>' +
|
||||||
|
// Search-instead panel
|
||||||
|
'<div data-mr-search-panel class="hidden space-y-2">' +
|
||||||
|
'<div class="flex gap-2">' +
|
||||||
|
'<input data-mr-search-input type="text" class="flex-1 bg-gray-800/50 border border-gray-700 rounded-md px-2 py-1 text-sm text-fb-text outline-none focus:border-fb-primary" placeholder="Artist – Title">' +
|
||||||
|
'<button data-mr-search-go class="text-sm text-fb-primary hover:text-fb-primaryHi border border-fb-primary/40 rounded-md px-3">Search</button></div>' +
|
||||||
|
'<div data-mr-search-results class="space-y-1"></div></div>' +
|
||||||
|
'</div>' +
|
||||||
|
// Footer actions
|
||||||
|
'<div class="flex items-center justify-between gap-3 p-5 pt-3 border-t border-fb-border/40 shrink-0">' +
|
||||||
|
'<div class="flex items-center gap-3">' +
|
||||||
|
'<button data-mr-reject class="text-sm text-fb-textDim hover:text-fb-text">Not a match</button>' +
|
||||||
|
'<button data-mr-search-toggle class="text-sm text-fb-textDim hover:text-fb-text">Search instead…</button></div>' +
|
||||||
|
'<div class="flex items-center gap-2">' +
|
||||||
|
'<button data-mr-skip class="text-sm text-fb-textDim hover:text-fb-text px-3 py-2">Skip</button>' +
|
||||||
|
'<button data-mr-accept class="bg-fb-primary hover:bg-fb-primaryHi text-white px-4 py-2 rounded-md text-sm">Use selected</button>' +
|
||||||
|
'</div></div>';
|
||||||
|
|
||||||
|
wireCurrent(panel, song);
|
||||||
|
}
|
||||||
|
|
||||||
|
function wireCurrent(panel, song) {
|
||||||
|
panel.querySelector('[data-mr-close]')?.addEventListener('click', closeModal);
|
||||||
|
panel.querySelector('[data-mr-prev]')?.addEventListener('click', () => nav(-1));
|
||||||
|
panel.querySelector('[data-mr-next]')?.addEventListener('click', () => nav(1));
|
||||||
|
panel.querySelector('[data-mr-skip]')?.addEventListener('click', () => nav(1));
|
||||||
|
// Art failure → flag + re-render once so the "cover art" chip shows.
|
||||||
|
const img = panel.querySelector('[data-mr-art]');
|
||||||
|
if (img) img.onerror = () => {
|
||||||
|
img.style.visibility = 'hidden';
|
||||||
|
if (!song._artMissing) { song._artMissing = true; renderCurrent(); }
|
||||||
|
};
|
||||||
|
panel.querySelectorAll('[data-mr-cand]').forEach((btn) => {
|
||||||
|
btn.addEventListener('click', () => {
|
||||||
|
song._sel = Number(btn.getAttribute('data-mr-cand'));
|
||||||
|
renderCurrent();
|
||||||
|
});
|
||||||
|
});
|
||||||
|
panel.querySelector('[data-mr-accept]')?.addEventListener('click', async () => {
|
||||||
|
const cand = (song.candidates || [])[song._sel || 0];
|
||||||
|
if (!cand) return;
|
||||||
|
await post('/api/enrichment/review/' + enc(song.filename) + '/accept',
|
||||||
|
{ recording_id: cand.recording_id });
|
||||||
|
settle(song);
|
||||||
|
});
|
||||||
|
panel.querySelector('[data-mr-reject]')?.addEventListener('click', async () => {
|
||||||
|
await post('/api/enrichment/review/' + enc(song.filename) + '/reject');
|
||||||
|
settle(song);
|
||||||
|
});
|
||||||
|
const sp = panel.querySelector('[data-mr-search-panel]');
|
||||||
|
const input = panel.querySelector('[data-mr-search-input]');
|
||||||
|
panel.querySelector('[data-mr-search-toggle]')?.addEventListener('click', () => {
|
||||||
|
sp?.classList.toggle('hidden');
|
||||||
|
if (sp && !sp.classList.contains('hidden') && input && !input.value) {
|
||||||
|
input.value = [song.artist, song.title].filter(Boolean).join(' – ');
|
||||||
|
input.focus();
|
||||||
|
}
|
||||||
|
});
|
||||||
|
const go = () => runSearch(panel, song);
|
||||||
|
panel.querySelector('[data-mr-search-go]')?.addEventListener('click', go);
|
||||||
|
input?.addEventListener('keydown', (e) => { if (e.key === 'Enter') { e.preventDefault(); go(); } });
|
||||||
|
}
|
||||||
|
|
||||||
|
// Silent-on-success: the chart just leaves the queue and the next one
|
||||||
|
// renders; the last one renders the done state. No toasts, no sounds.
|
||||||
|
function settle(song) {
|
||||||
|
const i = _queue.indexOf(song);
|
||||||
|
if (i >= 0) _queue.splice(i, 1);
|
||||||
|
if (_idx >= _queue.length) _idx = Math.max(0, _queue.length - 1);
|
||||||
|
refreshChip();
|
||||||
|
renderCurrent();
|
||||||
|
}
|
||||||
|
|
||||||
|
async function runSearch(panel, song) {
|
||||||
|
const input = panel.querySelector('[data-mr-search-input]');
|
||||||
|
const out = panel.querySelector('[data-mr-search-results]');
|
||||||
|
if (!input || !out) return;
|
||||||
|
const qRaw = input.value.trim();
|
||||||
|
if (!qRaw) return;
|
||||||
|
// "Artist – Title" splits on the first dash; a plain phrase searches
|
||||||
|
// as a title, which MusicBrainz handles well enough.
|
||||||
|
const m = qRaw.split(/\s+[–—-]\s+/);
|
||||||
|
const artist = m.length > 1 ? m[0] : '';
|
||||||
|
const title = m.length > 1 ? m.slice(1).join(' - ') : qRaw;
|
||||||
|
out.innerHTML = '<p class="text-xs text-fb-textDim">Searching…</p>';
|
||||||
|
let body = null;
|
||||||
|
try {
|
||||||
|
const r = await fetch('/api/enrichment/search?artist=' + enc(artist) +
|
||||||
|
'&title=' + enc(title) + '&filename=' + enc(song.filename));
|
||||||
|
if (r.status === 503) {
|
||||||
|
out.innerHTML = '<p class="text-xs text-fb-textDim">MusicBrainz is unavailable — try again later.</p>';
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (r.ok) body = await r.json();
|
||||||
|
} catch (_) { /* falls through to the no-results line */ }
|
||||||
|
const cands = (body && body.candidates) || [];
|
||||||
|
if (!cands.length) {
|
||||||
|
out.innerHTML = '<p class="text-xs text-fb-textDim">No results.</p>';
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
out.innerHTML = cands.map((c, i) => candRowHtml(song, c, i, false)).join('');
|
||||||
|
out.querySelectorAll('[data-mr-cand]').forEach((btn) => {
|
||||||
|
btn.addEventListener('click', async () => {
|
||||||
|
const cand = cands[Number(btn.getAttribute('data-mr-cand'))];
|
||||||
|
if (!cand) return;
|
||||||
|
await post('/api/enrichment/review/' + enc(song.filename) + '/pick',
|
||||||
|
{ candidate: cand });
|
||||||
|
settle(song);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async function post(url, payload) {
|
||||||
|
try {
|
||||||
|
await fetch(url, {
|
||||||
|
method: 'POST',
|
||||||
|
headers: { 'Content-Type': 'application/json' },
|
||||||
|
body: JSON.stringify(payload || {}),
|
||||||
|
});
|
||||||
|
} catch (_) { /* offline — the row simply stays queued */ }
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Settings → Library → "Metadata matching" card ────────────────────────
|
||||||
|
// Markup lives statically in index.html (the v3 settings pattern); this
|
||||||
|
// wires it. All null-guarded so v2 (which lacks the elements) no-ops.
|
||||||
|
function wireSettingsCard() {
|
||||||
|
const toggle = document.getElementById('enrich-enabled');
|
||||||
|
const sel = document.getElementById('enrich-threshold');
|
||||||
|
const btn = document.getElementById('enrich-match-now');
|
||||||
|
if (!toggle && !sel && !btn) return;
|
||||||
|
(async () => {
|
||||||
|
try {
|
||||||
|
const r = await fetch('/api/settings');
|
||||||
|
if (r.ok) {
|
||||||
|
const cfg = await r.json();
|
||||||
|
if (toggle) toggle.checked = cfg.enrich_enabled !== false;
|
||||||
|
if (sel) {
|
||||||
|
const t = Number(cfg.enrich_auto_threshold);
|
||||||
|
const want = Number.isFinite(t) ? t : 0.9;
|
||||||
|
// Snap to the nearest offered option.
|
||||||
|
let best = sel.options[0];
|
||||||
|
for (const o of sel.options) {
|
||||||
|
if (Math.abs(Number(o.value) - want) < Math.abs(Number(best.value) - want)) best = o;
|
||||||
|
}
|
||||||
|
if (best) sel.value = best.value;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
} catch (_) { /* leave markup defaults */ }
|
||||||
|
refreshChip(); // also fills #enrich-status
|
||||||
|
})();
|
||||||
|
const save = (key, value) => post('/api/settings', { [key]: value });
|
||||||
|
toggle?.addEventListener('change', () => save('enrich_enabled', !!toggle.checked));
|
||||||
|
sel?.addEventListener('change', () => save('enrich_auto_threshold', Number(sel.value)));
|
||||||
|
btn?.addEventListener('click', async () => {
|
||||||
|
await post('/api/enrichment/kick');
|
||||||
|
const line = document.getElementById('enrich-status');
|
||||||
|
if (line) line.textContent = 'Matching…';
|
||||||
|
setTimeout(refreshChip, 1500);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
if (document.readyState === 'loading') {
|
||||||
|
document.addEventListener('DOMContentLoaded', wireSettingsCard, { once: true });
|
||||||
|
} else {
|
||||||
|
wireSettingsCard();
|
||||||
|
}
|
||||||
|
|
||||||
|
window.__fbMatchReviewChip = refreshChip;
|
||||||
|
window.__fbOpenMatchReview = openModal;
|
||||||
|
})();
|
||||||
+24
-3
@@ -59,7 +59,7 @@
|
|||||||
artist: '', album: '',
|
artist: '', album: '',
|
||||||
grouping: true, // one card per song (multi-chart grouping); persisted
|
grouping: true, // one card per song (multi-chart grouping); persisted
|
||||||
|
|
||||||
filters: { arr_has: [], arr_lacks: [], stem_has: [], stem_lacks: [], lyrics: '', tunings: [], mastery: [] },
|
filters: { arr_has: [], arr_lacks: [], stem_has: [], stem_lacks: [], lyrics: '', tunings: [], mastery: [], match: [] },
|
||||||
page: 0, total: 0, loading: false, built: false, accuracy: {}, tuningNames: [],
|
page: 0, total: 0, loading: false, built: false, accuracy: {}, tuningNames: [],
|
||||||
artistCatalog: [], renderedHash: '',
|
artistCatalog: [], renderedHash: '',
|
||||||
scrollBound: false,
|
scrollBound: false,
|
||||||
@@ -110,6 +110,7 @@
|
|||||||
const f = state.filters;
|
const f = state.filters;
|
||||||
return f.arr_has.length + f.arr_lacks.length + f.stem_has.length + f.stem_lacks.length +
|
return f.arr_has.length + f.arr_lacks.length + f.stem_has.length + f.stem_lacks.length +
|
||||||
(f.lyrics ? 1 : 0) + f.tunings.length + (f.mastery ? f.mastery.length : 0) +
|
(f.lyrics ? 1 : 0) + f.tunings.length + (f.mastery ? f.mastery.length : 0) +
|
||||||
|
(f.match ? f.match.length : 0) +
|
||||||
(state.artist ? 1 : 0) + (state.album ? 1 : 0);
|
(state.artist ? 1 : 0) + (state.album ? 1 : 0);
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -133,6 +134,7 @@
|
|||||||
lyrics: f.lyrics || '',
|
lyrics: f.lyrics || '',
|
||||||
tunings: [...(f.tunings || [])].sort(),
|
tunings: [...(f.tunings || [])].sort(),
|
||||||
mastery: [...(f.mastery || [])].sort(),
|
mastery: [...(f.mastery || [])].sort(),
|
||||||
|
match: [...(f.match || [])].sort(),
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
@@ -155,10 +157,15 @@
|
|||||||
const f = saved.filters;
|
const f = saved.filters;
|
||||||
if (f && typeof f === 'object') {
|
if (f && typeof f === 'object') {
|
||||||
const arr = (x) => (Array.isArray(x) ? x.slice() : []);
|
const arr = (x) => (Array.isArray(x) ? x.slice() : []);
|
||||||
|
// mastery + match are session-only facets (deliberately not
|
||||||
|
// persisted), but the restored object must still CARRY the keys —
|
||||||
|
// the filter drawer indexes f.mastery/f.match unconditionally, so
|
||||||
|
// dropping them here breaks the drawer for anyone with saved prefs.
|
||||||
state.filters = {
|
state.filters = {
|
||||||
arr_has: arr(f.arr_has), arr_lacks: arr(f.arr_lacks),
|
arr_has: arr(f.arr_has), arr_lacks: arr(f.arr_lacks),
|
||||||
stem_has: arr(f.stem_has), stem_lacks: arr(f.stem_lacks),
|
stem_has: arr(f.stem_has), stem_lacks: arr(f.stem_lacks),
|
||||||
lyrics: f.lyrics || '', tunings: arr(f.tunings),
|
lyrics: f.lyrics || '', tunings: arr(f.tunings),
|
||||||
|
mastery: [], match: [],
|
||||||
};
|
};
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -263,6 +270,7 @@
|
|||||||
if (f.lyrics) p.set('has_lyrics', f.lyrics);
|
if (f.lyrics) p.set('has_lyrics', f.lyrics);
|
||||||
if (f.tunings.length) p.set('tunings', f.tunings.join(','));
|
if (f.tunings.length) p.set('tunings', f.tunings.join(','));
|
||||||
if (f.mastery && f.mastery.length) p.set('mastery', f.mastery.join(','));
|
if (f.mastery && f.mastery.length) p.set('mastery', f.mastery.join(','));
|
||||||
|
if (f.match && f.match.length) p.set('match', f.match.join(','));
|
||||||
Object.entries(extra || {}).forEach(([k, v]) => p.set(k, v));
|
Object.entries(extra || {}).forEach(([k, v]) => p.set(k, v));
|
||||||
return p;
|
return p;
|
||||||
}
|
}
|
||||||
@@ -2029,6 +2037,9 @@
|
|||||||
section('Lyrics', ['', '1', '0'].map((v) => '<button data-lyrics="' + v + '" class="px-2 py-1 rounded-md text-xs border ' + (f.lyrics === v ? 'bg-fb-primary text-white border-fb-primary' : 'bg-gray-800/50 text-fb-textDim border-gray-700') + '">' + (v === '' ? 'Any' : v === '1' ? 'Has lyrics' : 'No lyrics') + '</button>').join('')) +
|
section('Lyrics', ['', '1', '0'].map((v) => '<button data-lyrics="' + v + '" class="px-2 py-1 rounded-md text-xs border ' + (f.lyrics === v ? 'bg-fb-primary text-white border-fb-primary' : 'bg-gray-800/50 text-fb-textDim border-gray-700') + '">' + (v === '' ? 'Any' : v === '1' ? 'Has lyrics' : 'No lyrics') + '</button>').join('')) +
|
||||||
// Progress (mastery bands) — multi-select; server filters via song_stats.
|
// Progress (mastery bands) — multi-select; server filters via song_stats.
|
||||||
section('Progress', [['mastered', 'Mastered'], ['in_progress', 'In progress'], ['not_started', 'Not started']].map((it) => '<button data-mastery="' + it[0] + '" class="px-2 py-1 rounded-md text-xs border ' + (f.mastery.includes(it[0]) ? 'bg-fb-primary text-white border-fb-primary' : 'bg-gray-800/50 text-fb-textDim border-gray-700') + '">' + it[1] + '</button>').join('')) +
|
section('Progress', [['mastered', 'Mastered'], ['in_progress', 'In progress'], ['not_started', 'Not started']].map((it) => '<button data-mastery="' + it[0] + '" class="px-2 py-1 rounded-md text-xs border ' + (f.mastery.includes(it[0]) ? 'bg-fb-primary text-white border-fb-primary' : 'bg-gray-800/50 text-fb-textDim border-gray-700') + '">' + it[1] + '</button>').join('')) +
|
||||||
|
// Match (P8) — the song's metadata-match lifecycle state, a triage
|
||||||
|
// facet for the enrichment layer. Session-only, like Progress.
|
||||||
|
section('Match', [['review', 'To review'], ['matched', 'Matched'], ['unmatched', 'Unmatched'], ['pending', 'Not scanned']].map((it) => '<button data-match="' + it[0] + '" class="px-2 py-1 rounded-md text-xs border ' + (f.match.includes(it[0]) ? 'bg-fb-primary text-white border-fb-primary' : 'bg-gray-800/50 text-fb-textDim border-gray-700') + '">' + it[1] + '</button>').join('')) +
|
||||||
section('Tuning', (state.tuningNames || []).map((t) => {
|
section('Tuning', (state.tuningNames || []).map((t) => {
|
||||||
// Filter on the server's grouping key (raw offsets for customs)
|
// Filter on the server's grouping key (raw offsets for customs)
|
||||||
// so two "Custom Tuning" entries are distinct; show their target
|
// so two "Custom Tuning" entries are distinct; show their target
|
||||||
@@ -2080,11 +2091,12 @@
|
|||||||
renderDrawer();
|
renderDrawer();
|
||||||
reload(); // re-fetches grid + rail with/without group=1, saves prefs
|
reload(); // re-fetches grid + rail with/without group=1, saves prefs
|
||||||
});
|
});
|
||||||
|
d.querySelectorAll('[data-match]').forEach((b) => b.addEventListener('click', () => { const v = b.getAttribute('data-match'); const i = f.match.indexOf(v); if (i >= 0) f.match.splice(i, 1); else f.match.push(v); renderDrawer(); }));
|
||||||
d.querySelector('[data-drawer-save]')?.addEventListener('click', saveCurrentAsCollection);
|
d.querySelector('[data-drawer-save]')?.addEventListener('click', saveCurrentAsCollection);
|
||||||
d.querySelector('[data-drawer-tidy]')?.addEventListener('click', openArtistTidyUp);
|
d.querySelector('[data-drawer-tidy]')?.addEventListener('click', openArtistTidyUp);
|
||||||
d.querySelector('[data-drawer-close]')?.addEventListener('click', closeDrawer);
|
d.querySelector('[data-drawer-close]')?.addEventListener('click', closeDrawer);
|
||||||
d.querySelector('[data-drawer-clear]')?.addEventListener('click', async () => {
|
d.querySelector('[data-drawer-clear]')?.addEventListener('click', async () => {
|
||||||
state.filters = { arr_has: [], arr_lacks: [], stem_has: [], stem_lacks: [], lyrics: '', tunings: [], mastery: [] };
|
state.filters = { arr_has: [], arr_lacks: [], stem_has: [], stem_lacks: [], lyrics: '', tunings: [], mastery: [], match: [] };
|
||||||
state.artist = '';
|
state.artist = '';
|
||||||
state.album = '';
|
state.album = '';
|
||||||
renderDrawer();
|
renderDrawer();
|
||||||
@@ -2540,7 +2552,12 @@
|
|||||||
'<div class="max-w-7xl mx-auto px-6 md:px-8 pb-8">' +
|
'<div class="max-w-7xl mx-auto px-6 md:px-8 pb-8">' +
|
||||||
'<div id="v3-songs-toolbar" class="sticky z-20 -mx-6 md:-mx-8 px-6 md:px-8 py-3 mb-4 bg-fb-sidebar/95 backdrop-blur border-b border-fb-border/40">' +
|
'<div id="v3-songs-toolbar" class="sticky z-20 -mx-6 md:-mx-8 px-6 md:px-8 py-3 mb-4 bg-fb-sidebar/95 backdrop-blur border-b border-fb-border/40">' +
|
||||||
'<div class="flex flex-col md:flex-row md:items-end justify-between gap-4">' +
|
'<div class="flex flex-col md:flex-row md:items-end justify-between gap-4">' +
|
||||||
'<div><p class="text-fb-textDim text-sm" id="v3-songs-count"></p></div>' +
|
// The match-review chip is ambient tool-state (design §11): only
|
||||||
|
// rendered when matches are waiting, silent otherwise. Populated +
|
||||||
|
// shown by match-review.js (window.__fbMatchReviewChip), which
|
||||||
|
// also owns the drawer the click opens.
|
||||||
|
'<div class="flex items-baseline gap-3"><p class="text-fb-textDim text-sm" id="v3-songs-count"></p>' +
|
||||||
|
'<button id="v3-songs-match-review" class="hidden text-xs text-fb-primary hover:text-fb-primaryHi border border-fb-primary/40 rounded-full px-2.5 py-0.5"></button></div>' +
|
||||||
'<div class="flex flex-wrap gap-2">' +
|
'<div class="flex flex-wrap gap-2">' +
|
||||||
(providers.length > 1 ? '<select id="v3-songs-provider" class="' + ctrl + '">' + provOpts + '</select>' : '') +
|
(providers.length > 1 ? '<select id="v3-songs-provider" class="' + ctrl + '">' + provOpts + '</select>' : '') +
|
||||||
'<select id="v3-songs-artist" class="' + ctrl + ' max-w-[11rem]" aria-label="Artist">' + artistSelectHtml() + '</select>' +
|
'<select id="v3-songs-artist" class="' + ctrl + ' max-w-[11rem]" aria-label="Artist">' + artistSelectHtml() + '</select>' +
|
||||||
@@ -2599,6 +2616,10 @@
|
|||||||
});
|
});
|
||||||
byId('v3-songs-filters').addEventListener('click', openDrawer);
|
byId('v3-songs-filters').addEventListener('click', openDrawer);
|
||||||
byId('v3-songs-overlay').addEventListener('click', closeDrawer);
|
byId('v3-songs-overlay').addEventListener('click', closeDrawer);
|
||||||
|
// Match-review chip: feature-detected (match-review.js owns the
|
||||||
|
// drawer + the count; absent → the chip just stays hidden).
|
||||||
|
byId('v3-songs-match-review')?.addEventListener('click', () => { if (window.__fbOpenMatchReview) window.__fbOpenMatchReview(); });
|
||||||
|
if (window.__fbMatchReviewChip) window.__fbMatchReviewChip();
|
||||||
byId('v3-songs-upload').addEventListener('click', () => {
|
byId('v3-songs-upload').addEventListener('click', () => {
|
||||||
const legacy = document.getElementById('upload-songs-file');
|
const legacy = document.getElementById('upload-songs-file');
|
||||||
// Upload targets the LOCAL library + scan; watchUploadScan refreshes
|
// Upload targets the LOCAL library + scan; watchUploadScan refreshes
|
||||||
|
|||||||
@@ -91,6 +91,13 @@ def test_demo_off_settings_post_not_blocked(tmp_path, monkeypatch):
|
|||||||
("GET", "/api/plugins/updates"),
|
("GET", "/api/plugins/updates"),
|
||||||
("POST", "/api/plugins/highway_3d/files"),
|
("POST", "/api/plugins/highway_3d/files"),
|
||||||
("DELETE", "/api/plugins/highway_3d/files"),
|
("DELETE", "/api/plugins/highway_3d/files"),
|
||||||
|
# Enrichment (P8): review writes + the MusicBrainz search proxy (the
|
||||||
|
# proxy would spend the shared rate limit for anonymous demo visitors).
|
||||||
|
("POST", "/api/enrichment/review/some-file/accept"),
|
||||||
|
("POST", "/api/enrichment/review/some-file/reject"),
|
||||||
|
("POST", "/api/enrichment/review/some-file/pick"),
|
||||||
|
("POST", "/api/enrichment/kick"),
|
||||||
|
("GET", "/api/enrichment/search"),
|
||||||
])
|
])
|
||||||
def test_demo_on_blocked_routes_return_403(tmp_path, monkeypatch, method, path):
|
def test_demo_on_blocked_routes_return_403(tmp_path, monkeypatch, method, path):
|
||||||
server, client = _make_client(tmp_path, monkeypatch, demo=True)
|
server, client = _make_client(tmp_path, monkeypatch, demo=True)
|
||||||
|
|||||||
@@ -0,0 +1,512 @@
|
|||||||
|
"""Server-level tests for the P8 MusicBrainz matcher + Match-Review flow.
|
||||||
|
|
||||||
|
The HTTP transport is a fake installed over `server._mb_http_get` — the ONE
|
||||||
|
seam enrichment uses to reach the network — so nothing here ever opens a
|
||||||
|
socket. The offline default is itself under test: without explicitly
|
||||||
|
enabling the network flag, a pass must skip matching entirely (pytest can
|
||||||
|
never hit MusicBrainz, whatever a test triggers).
|
||||||
|
"""
|
||||||
|
|
||||||
|
import importlib
|
||||||
|
import sys
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture()
|
||||||
|
def server(tmp_path, monkeypatch, isolate_logging):
|
||||||
|
monkeypatch.setenv("CONFIG_DIR", str(tmp_path / "config"))
|
||||||
|
dlc = tmp_path / "dlc"
|
||||||
|
dlc.mkdir()
|
||||||
|
monkeypatch.setenv("DLC_DIR", str(dlc))
|
||||||
|
monkeypatch.setenv("FEEDBACK_SKIP_STARTUP_TASKS", "1")
|
||||||
|
sys.modules.pop("server", None)
|
||||||
|
srv = importlib.import_module("server")
|
||||||
|
try:
|
||||||
|
yield srv
|
||||||
|
finally:
|
||||||
|
conn = getattr(getattr(srv, "meta_db", None), "conn", None)
|
||||||
|
if conn is not None:
|
||||||
|
conn.close()
|
||||||
|
sys.modules.pop("server", None)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture()
|
||||||
|
def client(server):
|
||||||
|
return TestClient(server.app)
|
||||||
|
|
||||||
|
|
||||||
|
class FakeMB:
|
||||||
|
"""Canned MusicBrainz: records every call, serves per-path responses."""
|
||||||
|
|
||||||
|
def __init__(self):
|
||||||
|
self.calls = []
|
||||||
|
self.search_response = {"recordings": []}
|
||||||
|
self.recording_lookups = {} # mbid → recording doc
|
||||||
|
self.isrc_lookups = {} # isrc → {"recordings": [...]}
|
||||||
|
self.raise_transport = False
|
||||||
|
|
||||||
|
def __call__(self, path, params):
|
||||||
|
if self.raise_transport:
|
||||||
|
raise self._srv.EnrichTransportError("fake network down")
|
||||||
|
self.calls.append((path, dict(params)))
|
||||||
|
if path == "recording":
|
||||||
|
return self.search_response
|
||||||
|
if path.startswith("recording/"):
|
||||||
|
return self.recording_lookups.get(path.split("/", 1)[1])
|
||||||
|
if path.startswith("isrc/"):
|
||||||
|
return self.isrc_lookups.get(path.split("/", 1)[1])
|
||||||
|
raise AssertionError(f"unexpected MB path {path!r}")
|
||||||
|
|
||||||
|
@property
|
||||||
|
def search_calls(self):
|
||||||
|
return [c for c in self.calls if c[0] == "recording"]
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture()
|
||||||
|
def mb(server, monkeypatch):
|
||||||
|
"""Install the fake transport AND enable the network flag (the test env
|
||||||
|
disables it by default — see test_offline_default_skips_matching)."""
|
||||||
|
fake = FakeMB()
|
||||||
|
fake._srv = server
|
||||||
|
monkeypatch.setattr(server, "_mb_http_get", fake)
|
||||||
|
monkeypatch.setattr(server, "_enrich_network_enabled", lambda: True)
|
||||||
|
return fake
|
||||||
|
|
||||||
|
|
||||||
|
def _put(server, fn, title="Thunderstruck (v2)", artist="ACDC", album="",
|
||||||
|
duration=292, year="1990"):
|
||||||
|
server.meta_db.put(fn, 0, 0, {
|
||||||
|
"title": title, "artist": artist, "album": album, "year": year,
|
||||||
|
"duration": duration, "arrangements": [{"name": "Lead", "index": 0}],
|
||||||
|
})
|
||||||
|
|
||||||
|
|
||||||
|
def mb_doc(rid="rec-1", title="Thunderstruck", artist="AC/DC", artist_id="art-1",
|
||||||
|
album="The Razors Edge", date="1990-09-24", length_ms=292000, score=100):
|
||||||
|
return {
|
||||||
|
"id": rid, "score": score, "title": title, "length": length_ms,
|
||||||
|
"isrcs": ["AUAP09000045"],
|
||||||
|
"artist-credit": [{"name": artist, "artist": {
|
||||||
|
"id": artist_id, "name": artist, "sort-name": artist}}],
|
||||||
|
"releases": [{"id": "rel-1", "title": album, "status": "Official",
|
||||||
|
"date": date, "release-group": {"primary-type": "Album"}}],
|
||||||
|
"tags": [{"name": "hard rock", "count": 7}],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# ── offline safety (the pytest-never-hits-network contract) ──────────────────
|
||||||
|
|
||||||
|
def test_offline_default_skips_matching(server, monkeypatch):
|
||||||
|
"""Under the test env (FEEDBACK_SKIP_STARTUP_TASKS) a pass stamps hashes
|
||||||
|
but never matches — even with a transport installed."""
|
||||||
|
fake = FakeMB()
|
||||||
|
fake._srv = server
|
||||||
|
monkeypatch.setattr(server, "_mb_http_get", fake)
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
server._background_enrich()
|
||||||
|
assert fake.calls == []
|
||||||
|
assert server.meta_db.get_enrichment("a.sloppak")["match_state"] == "unscanned"
|
||||||
|
|
||||||
|
|
||||||
|
def test_real_transport_refuses_when_offline(server):
|
||||||
|
"""_mb_http_get itself raises (before any socket) when the network is
|
||||||
|
disabled — defence in depth under pytest."""
|
||||||
|
with pytest.raises(server.EnrichTransportError):
|
||||||
|
server._mb_http_get("recording", {"query": "x"})
|
||||||
|
|
||||||
|
|
||||||
|
def test_transport_error_pauses_pass_without_burning_attempts(server, mb):
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
_put(server, "b.sloppak", title="Other Song")
|
||||||
|
mb.raise_transport = True
|
||||||
|
server._background_enrich()
|
||||||
|
for fn in ("a.sloppak", "b.sloppak"):
|
||||||
|
row = server.meta_db.get_enrichment(fn)
|
||||||
|
assert row["match_state"] == "unscanned"
|
||||||
|
assert row["attempts"] == 0
|
||||||
|
# Network comes back → the next kick matches both.
|
||||||
|
mb.raise_transport = False
|
||||||
|
mb.search_response = {"recordings": [mb_doc()]}
|
||||||
|
server._background_enrich()
|
||||||
|
assert server.meta_db.get_enrichment("a.sloppak")["match_state"] == "matched"
|
||||||
|
|
||||||
|
|
||||||
|
# ── text tiers ────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
def test_high_confidence_auto_matches_and_settles(server, mb):
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
mb.search_response = {"recordings": [mb_doc()]}
|
||||||
|
server._background_enrich()
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["match_state"] == "matched"
|
||||||
|
assert row["match_source"] == "text"
|
||||||
|
assert row["match_score"] >= 0.95
|
||||||
|
assert row["mb_recording_id"] == "rec-1"
|
||||||
|
assert row["canon_artist"] == "AC/DC"
|
||||||
|
assert row["canon_title"] == "Thunderstruck"
|
||||||
|
assert row["canon_album"] == "The Razors Edge"
|
||||||
|
assert row["canon_year"] == "1990"
|
||||||
|
assert row["genres"] == ["hard rock"]
|
||||||
|
# Settled: another pass makes NO further network calls…
|
||||||
|
n = len(mb.calls)
|
||||||
|
server._background_enrich()
|
||||||
|
assert len(mb.calls) == n
|
||||||
|
# …until the identity changes, which re-matches.
|
||||||
|
_put(server, "a.sloppak", title="Back in Black")
|
||||||
|
mb.search_response = {"recordings": [mb_doc(rid="rec-2", title="Back in Black",
|
||||||
|
album="Back in Black", date="1980-07-25")]}
|
||||||
|
server._background_enrich()
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["mb_recording_id"] == "rec-2"
|
||||||
|
|
||||||
|
|
||||||
|
def test_medium_confidence_goes_to_review_not_canonical(server, mb):
|
||||||
|
# Partial artist agreement → medium confidence.
|
||||||
|
_put(server, "a.sloppak", artist="AC/DC ft Nobody")
|
||||||
|
mb.search_response = {"recordings": [mb_doc()]}
|
||||||
|
server._background_enrich()
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["match_state"] == "review"
|
||||||
|
assert row["match_source"] == "text"
|
||||||
|
# Review before auto-canonicalize: NO canonical values written yet.
|
||||||
|
assert row["canon_artist"] is None
|
||||||
|
assert row["mb_recording_id"] is None
|
||||||
|
assert row["candidates"] and row["candidates"][0]["recording_id"] == "rec-1"
|
||||||
|
# A review row is settled while its identity is unchanged — no re-query.
|
||||||
|
n = len(mb.calls)
|
||||||
|
server._background_enrich()
|
||||||
|
assert len(mb.calls) == n
|
||||||
|
|
||||||
|
|
||||||
|
def test_low_confidence_fails_with_backoff(server, mb):
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
mb.search_response = {"recordings": [mb_doc(rid="rec-x", title="Sunrise",
|
||||||
|
artist="Norah Jones")]}
|
||||||
|
server._background_enrich()
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["match_state"] == "failed"
|
||||||
|
assert row["attempts"] == 1
|
||||||
|
assert row["last_attempt_at"] is not None
|
||||||
|
# Immediately after, the backoff hasn't elapsed → no retry, no network.
|
||||||
|
n = len(mb.calls)
|
||||||
|
server._background_enrich()
|
||||||
|
assert len(mb.calls) == n
|
||||||
|
assert server.meta_db.get_enrichment("a.sloppak")["attempts"] == 1
|
||||||
|
# Rewind the clock two hours → eligible again, attempts increments.
|
||||||
|
with server.meta_db._lock:
|
||||||
|
server.meta_db.conn.execute(
|
||||||
|
"UPDATE song_enrichment SET last_attempt_at = last_attempt_at - 7200")
|
||||||
|
server.meta_db.conn.commit()
|
||||||
|
server._background_enrich()
|
||||||
|
assert len(mb.calls) == n + 1
|
||||||
|
assert server.meta_db.get_enrichment("a.sloppak")["attempts"] == 2
|
||||||
|
|
||||||
|
|
||||||
|
def test_no_results_fails(server, mb):
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
mb.search_response = {"recordings": []}
|
||||||
|
server._background_enrich()
|
||||||
|
assert server.meta_db.get_enrichment("a.sloppak")["match_state"] == "failed"
|
||||||
|
|
||||||
|
|
||||||
|
# ── the content-hash match cache ──────────────────────────────────────────────
|
||||||
|
|
||||||
|
def test_cache_hit_copies_match_without_network(server, mb):
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
_put(server, "b.sloppak") # identical identity → same content_hash
|
||||||
|
mb.search_response = {"recordings": [mb_doc()]}
|
||||||
|
server._background_enrich()
|
||||||
|
assert len(mb.search_calls) == 1 # ONE search covered both charts
|
||||||
|
a = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
b = server.meta_db.get_enrichment("b.sloppak")
|
||||||
|
assert a["match_state"] == b["match_state"] == "matched"
|
||||||
|
assert {a["match_source"], b["match_source"]} == {"text", "cache"}
|
||||||
|
assert a["mb_recording_id"] == b["mb_recording_id"] == "rec-1"
|
||||||
|
|
||||||
|
|
||||||
|
# ── exact keys from the manifest (tier 0 / tier 1) ────────────────────────────
|
||||||
|
|
||||||
|
def _write_sloppak_manifest(server, name, extra_yaml=""):
|
||||||
|
d = server.DLC_DIR / name
|
||||||
|
d.mkdir(parents=True)
|
||||||
|
(d / "manifest.yaml").write_text(
|
||||||
|
"title: Thunderstruck\nartist: AC/DC\nduration: 292\n"
|
||||||
|
"arrangements: []\nstems: []\n" + extra_yaml,
|
||||||
|
encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def test_manifest_mbid_tier0(server, mb):
|
||||||
|
mbid = "12345678-abcd-4ef0-9876-0123456789ab"
|
||||||
|
_write_sloppak_manifest(server, "a.sloppak", f"mbid: {mbid}\n")
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
mb.recording_lookups[mbid] = mb_doc(rid=mbid)
|
||||||
|
server._background_enrich()
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["match_state"] == "matched"
|
||||||
|
assert row["match_source"] == "mbid"
|
||||||
|
assert row["match_score"] == 1.0
|
||||||
|
assert row["mb_recording_id"] == mbid
|
||||||
|
assert mb.search_calls == [] # trusted key — no text search
|
||||||
|
|
||||||
|
|
||||||
|
def test_manifest_isrc_tier1(server, mb):
|
||||||
|
_write_sloppak_manifest(server, "a.sloppak", "isrc: AUAP09000045\n")
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
mb.isrc_lookups["AUAP09000045"] = {"recordings": [mb_doc()]}
|
||||||
|
server._background_enrich()
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["match_state"] == "matched"
|
||||||
|
assert row["match_source"] == "isrc"
|
||||||
|
assert mb.search_calls == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_bad_manifest_mbid_falls_through_to_text(server, mb):
|
||||||
|
mbid = "12345678-abcd-4ef0-9876-0123456789ab"
|
||||||
|
_write_sloppak_manifest(server, "a.sloppak", f"mbid: {mbid}\n")
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
mb.recording_lookups.clear() # lookup 404s (typo'd manifest)
|
||||||
|
mb.search_response = {"recordings": [mb_doc()]}
|
||||||
|
server._background_enrich()
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["match_state"] == "matched"
|
||||||
|
assert row["match_source"] == "text"
|
||||||
|
|
||||||
|
|
||||||
|
# ── manual is sacred ──────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
def test_manual_never_overwritten_by_matcher(server, mb):
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
server.meta_db.set_enrichment_manual(
|
||||||
|
"a.sloppak", {"recording_id": "user-pick", "title": "Thunderstruck",
|
||||||
|
"artist": "AC/DC"}, source="search")
|
||||||
|
mb.search_response = {"recordings": [mb_doc(rid="machine-pick")]}
|
||||||
|
server._background_enrich()
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["match_state"] == "manual"
|
||||||
|
assert row["mb_recording_id"] == "user-pick"
|
||||||
|
# The writer refuses machine writes onto manual outright.
|
||||||
|
ok = server.meta_db.apply_enrichment_match(
|
||||||
|
"a.sloppak", row["content_hash"], "matched", source="text", score=1.0,
|
||||||
|
cand={"recording_id": "machine-pick", "title": "X"})
|
||||||
|
assert ok is False
|
||||||
|
|
||||||
|
|
||||||
|
# ── review routes ─────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
def _seed_review(server, mb, fn="a.sloppak", title="Thunderstruck (v2)"):
|
||||||
|
# Distinct raw titles give distinct content hashes (else the match cache
|
||||||
|
# legitimately copies an earlier row instead of running the text tiers).
|
||||||
|
_put(server, fn, title=title, artist="AC/DC ft Nobody")
|
||||||
|
mb.search_response = {"recordings": [mb_doc()]}
|
||||||
|
server._background_enrich()
|
||||||
|
assert server.meta_db.get_enrichment(fn)["match_state"] == "review"
|
||||||
|
|
||||||
|
|
||||||
|
def test_review_queue_route(server, mb, client):
|
||||||
|
_seed_review(server, mb)
|
||||||
|
body = client.get("/api/enrichment/review").json()
|
||||||
|
assert body["total_review"] == 1
|
||||||
|
assert body["songs"][0]["filename"] == "a.sloppak"
|
||||||
|
assert body["songs"][0]["candidates"][0]["recording_id"] == "rec-1"
|
||||||
|
|
||||||
|
|
||||||
|
def test_review_accept_route(server, mb, client):
|
||||||
|
_seed_review(server, mb)
|
||||||
|
r = client.post("/api/enrichment/review/a.sloppak/accept",
|
||||||
|
json={"recording_id": "rec-1"})
|
||||||
|
assert r.status_code == 200
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["match_state"] == "manual"
|
||||||
|
assert row["match_source"] == "review"
|
||||||
|
assert row["canon_artist"] == "AC/DC"
|
||||||
|
assert client.get("/api/enrichment/review").json()["total_review"] == 0
|
||||||
|
# Accepting a candidate that isn't in the stored list → 404.
|
||||||
|
_seed_review(server, mb, fn="b.sloppak", title="Thunderstruck (Live)")
|
||||||
|
r = client.post("/api/enrichment/review/b.sloppak/accept",
|
||||||
|
json={"recording_id": "nope"})
|
||||||
|
assert r.status_code == 404
|
||||||
|
|
||||||
|
|
||||||
|
def test_review_reject_route_never_retries(server, mb, client):
|
||||||
|
_seed_review(server, mb)
|
||||||
|
r = client.post("/api/enrichment/review/a.sloppak/reject")
|
||||||
|
assert r.status_code == 200
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["match_state"] == "failed"
|
||||||
|
assert row["match_source"] == "rejected"
|
||||||
|
# Rejected rows are excluded from the retry backoff forever…
|
||||||
|
n = len(mb.calls)
|
||||||
|
server._background_enrich()
|
||||||
|
assert len(mb.calls) == n
|
||||||
|
# …but an identity edit re-queues (the user fixed the metadata).
|
||||||
|
_put(server, "a.sloppak", artist="AC/DC")
|
||||||
|
server._background_enrich()
|
||||||
|
assert server.meta_db.get_enrichment("a.sloppak")["match_state"] == "matched"
|
||||||
|
# Rejecting a manual row is refused.
|
||||||
|
r = client.post("/api/enrichment/review/a.sloppak/reject")
|
||||||
|
assert r.status_code == 200 # matched → rejectable
|
||||||
|
client.post("/api/enrichment/review/a.sloppak/pick",
|
||||||
|
json={"candidate": {"recording_id": "rec-9", "title": "T"}})
|
||||||
|
r = client.post("/api/enrichment/review/a.sloppak/reject")
|
||||||
|
assert r.status_code == 404
|
||||||
|
|
||||||
|
|
||||||
|
def test_pick_route_fix_match(server, mb, client):
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
r = client.post("/api/enrichment/review/a.sloppak/pick", json={"candidate": {
|
||||||
|
"recording_id": "rec-77", "title": "Thunderstruck", "artist": "AC/DC",
|
||||||
|
"album": "The Razors Edge", "year": "1990", "genres": ["hard rock"],
|
||||||
|
"junk_key": "dropped"}})
|
||||||
|
assert r.status_code == 200
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["match_state"] == "manual"
|
||||||
|
assert row["match_source"] == "search"
|
||||||
|
assert row["mb_recording_id"] == "rec-77"
|
||||||
|
# Malformed candidate → 400; unknown song → 404.
|
||||||
|
assert client.post("/api/enrichment/review/a.sloppak/pick",
|
||||||
|
json={"candidate": {"title": "no id"}}).status_code == 400
|
||||||
|
assert client.post("/api/enrichment/review/ghost.sloppak/pick",
|
||||||
|
json={"candidate": {"recording_id": "x", "title": "t"}}
|
||||||
|
).status_code == 404
|
||||||
|
|
||||||
|
|
||||||
|
def test_search_proxy(server, mb, client, monkeypatch):
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
mb.search_response = {"recordings": [mb_doc()]}
|
||||||
|
assert client.get("/api/enrichment/search").status_code == 400
|
||||||
|
body = client.get("/api/enrichment/search",
|
||||||
|
params={"title": "Thunderstruck", "artist": "AC/DC",
|
||||||
|
"filename": "a.sloppak"}).json()
|
||||||
|
assert body["candidates"][0]["recording_id"] == "rec-1"
|
||||||
|
assert body["candidates"][0]["score"] > 0.9
|
||||||
|
# Transport failure surfaces as 503, not a 500.
|
||||||
|
def _down(path, params):
|
||||||
|
raise server.EnrichTransportError("down")
|
||||||
|
monkeypatch.setattr(server, "_mb_http_get", _down)
|
||||||
|
r = client.get("/api/enrichment/search", params={"title": "x"})
|
||||||
|
assert r.status_code == 503
|
||||||
|
|
||||||
|
|
||||||
|
# ── the match facet on /api/library + /api/library/stats ─────────────────────
|
||||||
|
|
||||||
|
def test_match_facet_filters_grid_and_stats(server, mb, client, monkeypatch):
|
||||||
|
_put(server, "auto.sloppak") # → matched
|
||||||
|
_put(server, "rev.sloppak", title="Revsong", artist="AC/DC ft Nobody") # → review
|
||||||
|
_put(server, "fail.sloppak", title="Failsong", artist="Zzz") # → failed
|
||||||
|
_put(server, "pend.sloppak", title="Pendsong", artist="Yyy") # stays unscanned
|
||||||
|
|
||||||
|
def _routed(path, params):
|
||||||
|
q = params.get("query", "")
|
||||||
|
if "thunderstruck" in q:
|
||||||
|
return {"recordings": [mb_doc()]}
|
||||||
|
if "revsong" in q:
|
||||||
|
return {"recordings": [mb_doc(rid="rec-r", title="Revsong")]}
|
||||||
|
return {"recordings": []}
|
||||||
|
monkeypatch.setattr(server, "_mb_http_get", _routed)
|
||||||
|
server._background_enrich()
|
||||||
|
# Pendsong got failed by the pass (no results); reset it to unscanned to
|
||||||
|
# represent the not-yet-scanned band.
|
||||||
|
with server.meta_db._lock:
|
||||||
|
server.meta_db.conn.execute(
|
||||||
|
"UPDATE song_enrichment SET match_state='unscanned', attempts=0, "
|
||||||
|
"last_attempt_at=NULL WHERE filename='pend.sloppak'")
|
||||||
|
server.meta_db.conn.commit()
|
||||||
|
|
||||||
|
def names(match):
|
||||||
|
return sorted(s["filename"] for s in client.get(
|
||||||
|
"/api/library", params={"match": match, "size": 50}).json()["songs"])
|
||||||
|
|
||||||
|
assert names("review") == ["rev.sloppak"]
|
||||||
|
assert names("matched") == ["auto.sloppak"]
|
||||||
|
assert names("unmatched") == ["fail.sloppak"]
|
||||||
|
assert names("pending") == ["pend.sloppak"]
|
||||||
|
assert names("review,matched") == ["auto.sloppak", "rev.sloppak"]
|
||||||
|
# Stats agree with the grid (the rail's lockstep contract).
|
||||||
|
total = client.get("/api/library/stats",
|
||||||
|
params={"match": "review"}).json()["total_songs"]
|
||||||
|
assert total == 1
|
||||||
|
# Unknown values are ignored → unfiltered.
|
||||||
|
assert len(names("bogus")) == 4
|
||||||
|
|
||||||
|
|
||||||
|
def test_status_counts_by_state(server, mb, client):
|
||||||
|
_seed_review(server, mb)
|
||||||
|
states = client.get("/api/enrichment/status").json()["states"]
|
||||||
|
assert states.get("review") == 1
|
||||||
|
|
||||||
|
|
||||||
|
# ── settings: enable toggle + auto-apply confidence ───────────────────────────
|
||||||
|
|
||||||
|
def test_auto_threshold_setting_moves_the_auto_review_boundary(server, mb, client):
|
||||||
|
# artist exact (1.0) + title 4/5 token overlap (0.8) and NO year/duration
|
||||||
|
# corroboration → combined exactly 0.90.
|
||||||
|
mb.search_response = {"recordings": [
|
||||||
|
mb_doc(title="Highway Hell", artist="AC/DC", date="", length_ms=None)]}
|
||||||
|
client.post("/api/settings", json={"enrich_auto_threshold": 0.95})
|
||||||
|
_put(server, "a.sloppak", title="Highway to Hell", artist="AC/DC",
|
||||||
|
year="", duration=0)
|
||||||
|
server._background_enrich()
|
||||||
|
assert server.meta_db.get_enrichment("a.sloppak")["match_state"] == "review"
|
||||||
|
# Lower the bar to the default 0.90 → an identity edit re-queues, and the
|
||||||
|
# same 0.90-scored candidate now auto-applies.
|
||||||
|
client.post("/api/settings", json={"enrich_auto_threshold": 0.9})
|
||||||
|
_put(server, "a.sloppak", title="Highway to Hell", artist="AC/DC",
|
||||||
|
year="", duration=0, album="Different Album")
|
||||||
|
server._background_enrich()
|
||||||
|
row = server.meta_db.get_enrichment("a.sloppak")
|
||||||
|
assert row["match_state"] == "matched"
|
||||||
|
assert abs(row["match_score"] - 0.9) < 1e-6
|
||||||
|
|
||||||
|
|
||||||
|
def test_enrich_enabled_setting_gates_background_matching(server, mb, client):
|
||||||
|
client.post("/api/settings", json={"enrich_enabled": False})
|
||||||
|
_put(server, "a.sloppak")
|
||||||
|
mb.search_response = {"recordings": [mb_doc()]}
|
||||||
|
server._background_enrich()
|
||||||
|
assert mb.calls == []
|
||||||
|
assert server.meta_db.get_enrichment("a.sloppak")["match_state"] == "unscanned"
|
||||||
|
# Manual search/fix stays available while the background matcher is off.
|
||||||
|
r = client.get("/api/enrichment/search", params={"title": "Thunderstruck"})
|
||||||
|
assert r.status_code == 200
|
||||||
|
# Re-enable → the next pass matches.
|
||||||
|
client.post("/api/settings", json={"enrich_enabled": True})
|
||||||
|
server._background_enrich()
|
||||||
|
assert server.meta_db.get_enrichment("a.sloppak")["match_state"] == "matched"
|
||||||
|
|
||||||
|
|
||||||
|
def test_settings_validation(server, client):
|
||||||
|
assert "error" in client.post(
|
||||||
|
"/api/settings", json={"enrich_enabled": "yes"}).json()
|
||||||
|
assert "error" in client.post(
|
||||||
|
"/api/settings", json={"enrich_auto_threshold": "high"}).json()
|
||||||
|
assert "error" in client.post(
|
||||||
|
"/api/settings", json={"enrich_auto_threshold": 2.5}).json()
|
||||||
|
ok = client.post("/api/settings", json={"enrich_auto_threshold": 1.01}).json()
|
||||||
|
assert "error" not in ok
|
||||||
|
assert client.get("/api/settings").json()["enrich_auto_threshold"] == 1.01
|
||||||
|
|
||||||
|
|
||||||
|
def test_kick_route(server, mb, client):
|
||||||
|
import time as _t
|
||||||
|
body = client.post("/api/enrichment/kick").json()
|
||||||
|
assert "started" in body
|
||||||
|
# Let the kicked pass settle so its daemon thread can't bleed into the
|
||||||
|
# fixture teardown (the DB connection closes there).
|
||||||
|
for _ in range(200):
|
||||||
|
if not client.get("/api/enrichment/status").json()["running"]:
|
||||||
|
break
|
||||||
|
_t.sleep(0.02)
|
||||||
|
|
||||||
|
|
||||||
|
def test_review_queue_orders_missing_data_first(server, mb, client):
|
||||||
|
# Complete chart first alphabetically, incomplete second — the queue must
|
||||||
|
# surface the incomplete (missing album + year) one first anyway.
|
||||||
|
_seed_review(server, mb, fn="aa.sloppak", title="Thunderstruck (v2)")
|
||||||
|
_put(server, "zz.sloppak", title="Thunderstruck (Live)",
|
||||||
|
artist="AC/DC ft Nobody", album="", year="")
|
||||||
|
server._background_enrich()
|
||||||
|
assert server.meta_db.get_enrichment("zz.sloppak")["match_state"] == "review"
|
||||||
|
songs = client.get("/api/enrichment/review").json()["songs"]
|
||||||
|
assert [s["filename"] for s in songs] == ["zz.sloppak", "aa.sloppak"]
|
||||||
@@ -0,0 +1,227 @@
|
|||||||
|
"""Unit tests for lib/mb_match.py — the pure text-matching engine (P8).
|
||||||
|
|
||||||
|
No network, no database, no server import: denoise/tokenize, similarity,
|
||||||
|
scoring + tier classification, Lucene query building, and MusicBrainz
|
||||||
|
response parsing are all exercised as plain functions.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import mb_match as m
|
||||||
|
|
||||||
|
|
||||||
|
# ── denoise / tokenize ────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
def test_denoise_lowercases_and_strips_punct_and_diacritics():
|
||||||
|
assert m.denoise("Motörhead") == "motorhead"
|
||||||
|
assert m.denoise("Beyoncé!!") == "beyonce"
|
||||||
|
assert m.denoise("Guns N' Roses") == "guns n roses"
|
||||||
|
assert m.denoise(" Weird spacing ") == "weird spacing"
|
||||||
|
|
||||||
|
|
||||||
|
def test_denoise_strips_noise_parentheticals():
|
||||||
|
# The design's explicit list: author suffixes + (440Hz)/(Live)/(No Lead)/(v2).
|
||||||
|
assert m.denoise("Thunderstruck (440Hz)") == "thunderstruck"
|
||||||
|
assert m.denoise("Thunderstruck (Live)") == "thunderstruck"
|
||||||
|
assert m.denoise("Thunderstruck (No Lead)") == "thunderstruck"
|
||||||
|
assert m.denoise("Thunderstruck (v2)") == "thunderstruck"
|
||||||
|
assert m.denoise("Thunderstruck [Remastered 2012]") == "thunderstruck"
|
||||||
|
assert m.denoise("One (Live at Wembley)") == "one"
|
||||||
|
|
||||||
|
|
||||||
|
def test_denoise_strips_author_credits():
|
||||||
|
assert m.denoise("Back in Black (by SomeCharter)") == "back in black"
|
||||||
|
assert m.denoise("Back in Black (charted by X99)") == "back in black"
|
||||||
|
assert m.denoise("Back in Black - by SomeCharter") == "back in black"
|
||||||
|
|
||||||
|
|
||||||
|
def test_denoise_keeps_meaningful_parentheticals():
|
||||||
|
# A parenthetical with no noise term survives (both sides get the same
|
||||||
|
# treatment, so symmetric content still matches).
|
||||||
|
assert m.denoise("Doin' It (All for My Baby)") == "doin it all for my baby"
|
||||||
|
|
||||||
|
|
||||||
|
def test_denoise_leading_the_is_artist_only():
|
||||||
|
assert m.denoise("The Beatles", strip_leading_the=True) == "beatles"
|
||||||
|
# Titles keep their "The" — never strip it there.
|
||||||
|
assert m.denoise("The Trooper") == "the trooper"
|
||||||
|
|
||||||
|
|
||||||
|
def test_ampersand_folds_to_and():
|
||||||
|
assert m.similarity("Angus & Julia Stone", "Angus and Julia Stone", artist=True) == 1.0
|
||||||
|
|
||||||
|
|
||||||
|
# ── similarity ────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
def test_similarity_exact_and_empty():
|
||||||
|
assert m.similarity("Back in Black", "Back In Black!") == 1.0
|
||||||
|
assert m.similarity("", "Anything") == 0.0
|
||||||
|
assert m.similarity(None, None) == 0.0
|
||||||
|
|
||||||
|
|
||||||
|
def test_similarity_folds_spelling_drift_via_compaction():
|
||||||
|
# The headline case: ACDC / AC DC / AC/DC all name the same artist.
|
||||||
|
assert m.similarity("ACDC", "AC/DC", artist=True) == 1.0
|
||||||
|
assert m.similarity("AC DC", "ACDC", artist=True) == 1.0
|
||||||
|
assert m.similarity("Greenday", "Green Day", artist=True) == 1.0
|
||||||
|
|
||||||
|
|
||||||
|
def test_similarity_partial_overlap():
|
||||||
|
s = m.similarity("Highway to Hell", "Highway Hell")
|
||||||
|
assert 0.7 < s < 1.0
|
||||||
|
assert m.similarity("Back in Black", "Paint It Black") < 0.5
|
||||||
|
|
||||||
|
|
||||||
|
# ── scoring + tiers ───────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
SONG = {"artist": "ACDC", "title": "Thunderstruck (v2)", "album": "The Razors Edge",
|
||||||
|
"year": "1990", "duration": 292}
|
||||||
|
|
||||||
|
|
||||||
|
def test_score_exact_match_is_high():
|
||||||
|
cand = {"artist": "AC/DC", "title": "Thunderstruck", "year": "1990", "duration": 292}
|
||||||
|
s = m.score_candidate(SONG, cand)
|
||||||
|
assert s == 1.0
|
||||||
|
assert m.classify(SONG, cand, s) == "auto"
|
||||||
|
|
||||||
|
|
||||||
|
def test_score_cover_never_auto():
|
||||||
|
# Perfect title, wrong artist (a cover) — must not auto-match.
|
||||||
|
cand = {"artist": "Some Cover Band", "title": "Thunderstruck"}
|
||||||
|
s = m.score_candidate(SONG, cand)
|
||||||
|
assert m.classify(SONG, cand, s) != "auto"
|
||||||
|
|
||||||
|
|
||||||
|
def test_missing_artist_caps_at_review():
|
||||||
|
song = {"artist": "", "title": "Thunderstruck", "duration": 292}
|
||||||
|
cand = {"artist": "AC/DC", "title": "Thunderstruck", "duration": 292}
|
||||||
|
s = m.score_candidate(song, cand)
|
||||||
|
# artist half scores 0 → combined ≤ 0.55 + bonuses → review at best.
|
||||||
|
assert m.classify(song, cand, s) != "auto"
|
||||||
|
|
||||||
|
|
||||||
|
def test_year_and_duration_corroborate():
|
||||||
|
# Fuzzy title so the base sits below the 1.0 cap and bonuses are visible.
|
||||||
|
base = {"artist": "AC/DC", "title": "Thunderstruck Thunder"}
|
||||||
|
plain = m.score_candidate(SONG, base)
|
||||||
|
with_year = m.score_candidate(SONG, dict(base, year="1990"))
|
||||||
|
with_dur = m.score_candidate(SONG, dict(base, duration=290))
|
||||||
|
assert with_year > plain
|
||||||
|
assert with_dur > plain
|
||||||
|
|
||||||
|
|
||||||
|
def test_fuzzy_title_with_corroboration_lands_review_or_auto():
|
||||||
|
cand = {"artist": "AC/DC", "title": "Thunderstruck Thunder"}
|
||||||
|
s = m.score_candidate(SONG, cand)
|
||||||
|
assert m.classify(SONG, cand, s) in ("review", "auto")
|
||||||
|
|
||||||
|
|
||||||
|
def test_unrelated_is_none():
|
||||||
|
cand = {"artist": "Norah Jones", "title": "Sunrise"}
|
||||||
|
s = m.score_candidate(SONG, cand)
|
||||||
|
assert m.classify(SONG, cand, s) == "none"
|
||||||
|
|
||||||
|
|
||||||
|
def test_classify_auto_min_override():
|
||||||
|
# The host's "auto-apply confidence" setting: a perfect match autos at
|
||||||
|
# any real threshold, and "Always review" (>1.0) sends even it to review.
|
||||||
|
cand = {"artist": "AC/DC", "title": "Thunderstruck", "year": "1990", "duration": 292}
|
||||||
|
s = m.score_candidate(SONG, cand)
|
||||||
|
assert s == 1.0
|
||||||
|
assert m.classify(SONG, cand, s, auto_min=0.9) == "auto"
|
||||||
|
assert m.classify(SONG, cand, s, auto_min=1.01) == "review"
|
||||||
|
# The per-field floors are independent of the threshold: a wrong-artist
|
||||||
|
# cover stays non-auto even at a permissive auto_min.
|
||||||
|
cover = {"artist": "Some Cover Band", "title": "Thunderstruck"}
|
||||||
|
cs = m.score_candidate(SONG, cover)
|
||||||
|
assert m.classify(SONG, cover, cs, auto_min=0.5) != "auto"
|
||||||
|
|
||||||
|
|
||||||
|
def test_rank_candidates_orders_by_our_score():
|
||||||
|
cands = [
|
||||||
|
{"recording_id": "b", "artist": "Someone Else", "title": "Thunderstruck", "mb_score": 100},
|
||||||
|
{"recording_id": "a", "artist": "AC/DC", "title": "Thunderstruck", "mb_score": 90},
|
||||||
|
]
|
||||||
|
ranked = m.rank_candidates(SONG, cands)
|
||||||
|
assert [c["recording_id"] for c in ranked] == ["a", "b"]
|
||||||
|
assert all("score" in c for c in ranked)
|
||||||
|
|
||||||
|
|
||||||
|
# ── query building ────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
def test_build_recording_query_denoises_and_quotes():
|
||||||
|
q = m.build_recording_query("ACDC", 'Thunderstruck (v2)')
|
||||||
|
assert q == 'recording:"thunderstruck" AND artist:"acdc"'
|
||||||
|
|
||||||
|
|
||||||
|
def test_build_recording_query_escapes_and_handles_missing_artist():
|
||||||
|
q = m.build_recording_query("", 'Say "Hello"')
|
||||||
|
# Quotes are punct-stripped by denoise, so nothing to escape here — but
|
||||||
|
# the artist clause must be absent entirely.
|
||||||
|
assert q.startswith('recording:"')
|
||||||
|
assert "artist:" not in q
|
||||||
|
|
||||||
|
|
||||||
|
# ── MusicBrainz response parsing ──────────────────────────────────────────────
|
||||||
|
|
||||||
|
MB_DOC = {
|
||||||
|
"id": "rec-123",
|
||||||
|
"score": 98,
|
||||||
|
"title": "Thunderstruck",
|
||||||
|
"length": 292773,
|
||||||
|
"isrcs": ["AUAP09000045"],
|
||||||
|
"artist-credit": [
|
||||||
|
{"name": "AC/DC", "joinphrase": "",
|
||||||
|
"artist": {"id": "art-1", "name": "AC/DC", "sort-name": "AC/DC"}},
|
||||||
|
],
|
||||||
|
"releases": [
|
||||||
|
{"id": "rel-compilation", "title": "Greatest Hits", "status": "Official",
|
||||||
|
"date": "2005-01-01", "release-group": {"primary-type": "Compilation"}},
|
||||||
|
{"id": "rel-album", "title": "The Razors Edge", "status": "Official",
|
||||||
|
"date": "1990-09-24", "release-group": {"primary-type": "Album"}},
|
||||||
|
{"id": "rel-boot", "title": "Bootleg", "status": "Bootleg",
|
||||||
|
"date": "1989-01-01", "release-group": {"primary-type": "Album"}},
|
||||||
|
],
|
||||||
|
"tags": [{"name": "hard rock", "count": 10}, {"name": "rock", "count": 4}],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_recording_doc_normalizes():
|
||||||
|
c = m.parse_recording_doc(MB_DOC)
|
||||||
|
assert c["recording_id"] == "rec-123"
|
||||||
|
assert c["title"] == "Thunderstruck"
|
||||||
|
assert c["artist"] == "AC/DC"
|
||||||
|
assert c["artist_id"] == "art-1"
|
||||||
|
# Official Album beats the compilation and the bootleg.
|
||||||
|
assert c["album"] == "The Razors Edge"
|
||||||
|
assert c["release_id"] == "rel-album"
|
||||||
|
assert c["year"] == "1990"
|
||||||
|
assert c["duration"] == 293
|
||||||
|
assert c["isrc"] == "AUAP09000045"
|
||||||
|
assert c["genres"] == ["hard rock", "rock"]
|
||||||
|
assert c["mb_score"] == 98
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_recording_doc_joined_artist_credit():
|
||||||
|
doc = dict(MB_DOC)
|
||||||
|
doc["artist-credit"] = [
|
||||||
|
{"name": "Queen", "joinphrase": " & ",
|
||||||
|
"artist": {"id": "q", "name": "Queen", "sort-name": "Queen"}},
|
||||||
|
{"name": "David Bowie",
|
||||||
|
"artist": {"id": "b", "name": "David Bowie", "sort-name": "Bowie, David"}},
|
||||||
|
]
|
||||||
|
c = m.parse_recording_doc(doc)
|
||||||
|
assert c["artist"] == "Queen & David Bowie"
|
||||||
|
assert c["artist_id"] == "q"
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_recording_doc_rejects_malformed():
|
||||||
|
assert m.parse_recording_doc({}) is None
|
||||||
|
assert m.parse_recording_doc({"id": "x"}) is None
|
||||||
|
assert m.parse_recording_doc(None) is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_search_response():
|
||||||
|
body = {"recordings": [MB_DOC, {"bogus": True}]}
|
||||||
|
cands = m.parse_search_response(body)
|
||||||
|
assert len(cands) == 1
|
||||||
|
assert m.parse_search_response({}) == []
|
||||||
|
assert m.parse_search_response(None) == []
|
||||||
Reference in New Issue
Block a user