mirror of
https://github.com/got-feedBack/feedBack.git
synced 2026-08-11 03:09:57 +00:00
6ece31b0201736745a8646aace3ab21653b45c4b
1
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
73c5ab149e |
feat(enrichment): AcoustID audio-fingerprint identification (opt-in) (#759)
* feat(enrichment): AcoustID audio-fingerprint identification (opt-in) Text search can only guess the version; the definitive fix is content-based — fingerprint the actual audio with Chromaprint (fpcalc) and look it up on AcoustID, which maps the fingerprint to the EXACT MusicBrainz recording (the approach Lidarr uses). Sidesteps the studio-vs-live ambiguity entirely. - lib/acoustid_match.py: pure response parsing + config gating (unit-tested); normalizes AcoustID hits into the same candidate shape as mb_match so the review UI + editor Match popup render fingerprint and text hits identically. - server.py: _fpcalc (Chromaprint subprocess), _acoustid_lookup (throttled, offline-guarded HTTP), _identify_by_fingerprint (also available to the library-enrichment pipeline), and POST /api/enrichment/identify (upload the master audio → candidates). - Fully OPT-IN and graceful: absent the fpcalc binary or an ACOUSTID_API_KEY the whole path is a no-op / 503 and the text matcher runs unchanged. Requires (both optional): the `fpcalc` (Chromaprint) binary on PATH/$FPCALC, and a free AcoustID application key in $ACOUSTID_API_KEY. Pure parsing/gating is unit-tested; the fpcalc + live-lookup path needs those two to exercise. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(enrichment): make AcoustID self-serve — opt-in toggle + API key in settings Fingerprinting was env-var only (ACOUSTID_API_KEY), so only an operator could enable it. Add two core settings so a user can turn it on themselves: - acoustid_enabled (bool, default OFF — opt-in) - acoustid_api_key (string, ≤128 chars, trimmed; env var stays a fallback) _acoustid_available()/_acoustid_lookup() now resolve (enabled, key) from settings via _acoustid_settings(). /api/enrichment/identify distinguishes "not set up" (412 needs_setup — the UI nudges the user to enable it) from "set up but fpcalc/network missing" (503) so the client never fakes a match. Verified: default off; POST round-trips + trims; 412 vs 503 gating; bad types/over-length rejected. acoustid_match unit tests green (8/8). * fix(enrichment): POST the AcoustID lookup instead of GET A Chromaprint fingerprint is multi-KB (a 3.5-min track ≈ 3.5k chars), so sending it as a GET query param overflows the request URL for longer songs and fails spuriously. AcoustID accepts the same params form-encoded — POST them. * fix(acoustid): space-separate the lookup `meta` (was silently dropping metadata) The meta value was `+`-joined ("recordings+releasegroups+compress"). Sent over the wire the literal `+` percent-encodes to %2B, which AcoustID does NOT split into flags — so every hit came back with an empty `recordings` array and the parser produced zero candidates (a fingerprint match that resolved to nothing). AcoustID wants the flags space-separated. Verified against real fingerprints: `+`-joined → 0 recordings; space-joined → 28, resolving Highway to Hell and Living After Midnight to their canonical studio albums as the top hit. * feat(acoustid): resolve the canonical original album + year from the fingerprint AcoustID hits resolved the right recording but a weak album/blank year: the album picker took the first studio-typed group (a later comp/soundtrack typed "Album" could win) and the year took an arbitrary release (often a reissue). Request the `releases` meta (which carries per-release dates) and use them to (1) pick the EARLIEST original studio album among the groups and (2) fill the year from that album's earliest release. Verified against real fingerprints: Smoke on the Water → Machine Head (1972) not a later comp; Highway to Hell → 1979; Living After Midnight → British Steel (1980). +2 unit tests. * feat(acoustid): per-song "Identify by audio" for the library metadata tooling Add POST /api/enrichment/identify/{filename} — fingerprints an EXISTING library song's own master audio (resolves the sloppak's original_audio or a loose folder's audio), the library counterpart to the upload-based /identify used by the editor. Wire an "Identify by audio" action into the match-review / Fix-match modal: it renders fingerprint hits in the same candidate list and pins the pick via the existing /review/{f}/pick. Shared _acoustid_gate() (412 needs_setup / 503) for both endpoints; 404 when a pack has no full mix. Both identify routes added to the demo-mode block list (they spend fpcalc + the AcoustID budget) — fixes a pre-existing miss on the upload route. * fix(acoustid): regenerate stale tailwind CSS + cap identify upload - static/tailwind.min.css was stale vs a fresh rebuild (ci/tailwind-fresh red); regenerated with the pinned tailwindcss@3.4.19 (byte-stable). - /api/enrichment/identify read the whole multipart upload into memory before writing it; stream it to the temp file with a 256 MB cap (413 over) so an oversized upload can't balloon RAM. fpcalc reads from the temp file anyway. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(acoustid): pre-parse upload guard + settings UI to enable it - /api/enrichment/identify is now async: a pre-parse Content-Length check + request.form(max_part_size=…) reject an oversized body BEFORE Starlette spools the multipart to temp disk (mirrors the song-upload endpoint), and the blocking fpcalc subprocess + AcoustID HTTP run off the event loop via run_in_executor. - The v3 Metadata-matching settings card gains an 'Identify by audio' opt-in toggle (acoustid_enabled, default OFF) + an AcoustID key input (acoustid_api_key), wired in match-review.js — so the advertised feature is reachable from the UI instead of only via a manual settings POST. Reuses existing classes only; committed tailwind.min.css stays fresh. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |