mirror of
https://github.com/got-feedBack/feedBack.git
synced 2026-07-21 20:31:21 +00:00
* feat(enrichment): AcoustID audio-fingerprint identification (opt-in)
Text search can only guess the version; the definitive fix is content-based —
fingerprint the actual audio with Chromaprint (fpcalc) and look it up on
AcoustID, which maps the fingerprint to the EXACT MusicBrainz recording (the
approach Lidarr uses). Sidesteps the studio-vs-live ambiguity entirely.
- lib/acoustid_match.py: pure response parsing + config gating (unit-tested);
normalizes AcoustID hits into the same candidate shape as mb_match so the
review UI + editor Match popup render fingerprint and text hits identically.
- server.py: _fpcalc (Chromaprint subprocess), _acoustid_lookup (throttled,
offline-guarded HTTP), _identify_by_fingerprint (also available to the
library-enrichment pipeline), and POST /api/enrichment/identify (upload the
master audio → candidates).
- Fully OPT-IN and graceful: absent the fpcalc binary or an ACOUSTID_API_KEY
the whole path is a no-op / 503 and the text matcher runs unchanged.
Requires (both optional): the `fpcalc` (Chromaprint) binary on PATH/$FPCALC,
and a free AcoustID application key in $ACOUSTID_API_KEY. Pure parsing/gating
is unit-tested; the fpcalc + live-lookup path needs those two to exercise.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(enrichment): make AcoustID self-serve — opt-in toggle + API key in settings
Fingerprinting was env-var only (ACOUSTID_API_KEY), so only an operator could
enable it. Add two core settings so a user can turn it on themselves:
- acoustid_enabled (bool, default OFF — opt-in)
- acoustid_api_key (string, ≤128 chars, trimmed; env var stays a fallback)
_acoustid_available()/_acoustid_lookup() now resolve (enabled, key) from
settings via _acoustid_settings(). /api/enrichment/identify distinguishes
"not set up" (412 needs_setup — the UI nudges the user to enable it) from
"set up but fpcalc/network missing" (503) so the client never fakes a match.
Verified: default off; POST round-trips + trims; 412 vs 503 gating; bad
types/over-length rejected. acoustid_match unit tests green (8/8).
* fix(enrichment): POST the AcoustID lookup instead of GET
A Chromaprint fingerprint is multi-KB (a 3.5-min track ≈ 3.5k chars), so
sending it as a GET query param overflows the request URL for longer songs and
fails spuriously. AcoustID accepts the same params form-encoded — POST them.
* fix(acoustid): space-separate the lookup `meta` (was silently dropping metadata)
The meta value was `+`-joined ("recordings+releasegroups+compress"). Sent over
the wire the literal `+` percent-encodes to %2B, which AcoustID does NOT split
into flags — so every hit came back with an empty `recordings` array and the
parser produced zero candidates (a fingerprint match that resolved to nothing).
AcoustID wants the flags space-separated. Verified against real fingerprints:
`+`-joined → 0 recordings; space-joined → 28, resolving Highway to Hell and
Living After Midnight to their canonical studio albums as the top hit.
* feat(acoustid): resolve the canonical original album + year from the fingerprint
AcoustID hits resolved the right recording but a weak album/blank year: the
album picker took the first studio-typed group (a later comp/soundtrack typed
"Album" could win) and the year took an arbitrary release (often a reissue).
Request the `releases` meta (which carries per-release dates) and use them to
(1) pick the EARLIEST original studio album among the groups and (2) fill the
year from that album's earliest release. Verified against real fingerprints:
Smoke on the Water → Machine Head (1972) not a later comp; Highway to Hell →
1979; Living After Midnight → British Steel (1980). +2 unit tests.
* feat(acoustid): per-song "Identify by audio" for the library metadata tooling
Add POST /api/enrichment/identify/{filename} — fingerprints an EXISTING library
song's own master audio (resolves the sloppak's original_audio or a loose
folder's audio), the library counterpart to the upload-based /identify used by
the editor. Wire an "Identify by audio" action into the match-review / Fix-match
modal: it renders fingerprint hits in the same candidate list and pins the pick
via the existing /review/{f}/pick. Shared _acoustid_gate() (412 needs_setup /
503) for both endpoints; 404 when a pack has no full mix. Both identify routes
added to the demo-mode block list (they spend fpcalc + the AcoustID budget) —
fixes a pre-existing miss on the upload route.
* fix(acoustid): regenerate stale tailwind CSS + cap identify upload
- static/tailwind.min.css was stale vs a fresh rebuild (ci/tailwind-fresh red);
regenerated with the pinned tailwindcss@3.4.19 (byte-stable).
- /api/enrichment/identify read the whole multipart upload into memory before
writing it; stream it to the temp file with a 256 MB cap (413 over) so an
oversized upload can't balloon RAM. fpcalc reads from the temp file anyway.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(acoustid): pre-parse upload guard + settings UI to enable it
- /api/enrichment/identify is now async: a pre-parse Content-Length check +
request.form(max_part_size=…) reject an oversized body BEFORE Starlette spools
the multipart to temp disk (mirrors the song-upload endpoint), and the blocking
fpcalc subprocess + AcoustID HTTP run off the event loop via run_in_executor.
- The v3 Metadata-matching settings card gains an 'Identify by audio' opt-in
toggle (acoustid_enabled, default OFF) + an AcoustID key input
(acoustid_api_key), wired in match-review.js — so the advertised feature is
reachable from the UI instead of only via a manual settings POST. Reuses
existing classes only; committed tailwind.min.css stays fresh.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
121 lines
4.4 KiB
Python
121 lines
4.4 KiB
Python
"""Pure-function tests for AcoustID fingerprint response parsing + config
|
|
gating. No network, no fpcalc binary — server.py owns those seams."""
|
|
import acoustid_match as a
|
|
|
|
|
|
def _resp(score=0.97, rec_id="rec-1", title="Highway to Hell", artist="AC/DC",
|
|
rg_title="Highway to Hell", rg_type="Album", secondary=None,
|
|
year=1979, duration=208.4):
|
|
return {
|
|
"status": "ok",
|
|
"results": [{
|
|
"id": "acoustid-uuid",
|
|
"score": score,
|
|
"recordings": [{
|
|
"id": rec_id,
|
|
"title": title,
|
|
"duration": duration,
|
|
"artists": [{"id": "a1", "name": artist}],
|
|
"releasegroups": [{
|
|
"id": "rg1", "title": rg_title, "type": rg_type,
|
|
"secondarytypes": secondary or [],
|
|
"releases": [{"date": {"year": year}}],
|
|
}],
|
|
}],
|
|
}],
|
|
}
|
|
|
|
|
|
def test_parse_maps_the_studio_recording():
|
|
out = a.parse_lookup_response(_resp())
|
|
assert len(out) == 1
|
|
c = out[0]
|
|
assert c["recording_id"] == "rec-1"
|
|
assert c["title"] == "Highway to Hell"
|
|
assert c["artist"] == "AC/DC"
|
|
assert c["album"] == "Highway to Hell"
|
|
assert c["year"] == "1979"
|
|
assert c["duration"] == 208
|
|
assert c["studio"] is True
|
|
assert c["source"] == "acoustid"
|
|
assert c["mb_score"] == 97 # 0.97 → 0..100 confidence band
|
|
assert c["score"] == 0.97
|
|
|
|
|
|
def test_live_release_group_is_not_studio():
|
|
out = a.parse_lookup_response(_resp(rg_type="Album", secondary=["Live"]))
|
|
assert out[0]["studio"] is False
|
|
|
|
|
|
def test_compilation_is_not_studio():
|
|
out = a.parse_lookup_response(_resp(secondary=["Compilation"]))
|
|
assert out[0]["studio"] is False
|
|
|
|
|
|
def test_prefers_studio_group_for_album_display():
|
|
resp = _resp()
|
|
# Add a comp release-group first; the studio one must win the album pick.
|
|
resp["results"][0]["recordings"][0]["releasegroups"].insert(0, {
|
|
"id": "rg0", "title": "Greatest Hits", "type": "Album",
|
|
"secondarytypes": ["Compilation"], "releases": [{"date": {"year": 2000}}],
|
|
})
|
|
c = a.parse_lookup_response(resp)[0]
|
|
assert c["album"] == "Highway to Hell"
|
|
assert c["studio"] is True
|
|
|
|
|
|
def test_earliest_studio_album_wins_over_later_one():
|
|
# Two studio "Album" groups (e.g. a later soundtrack typed Album). The
|
|
# ORIGINAL — earliest release year — must win the album pick, not whichever
|
|
# AcoustID happened to list first. (Real case: "Machine Head" over a later
|
|
# comp for "Smoke on the Water".)
|
|
resp = _resp(rg_title="Machine Head", year=1972)
|
|
resp["results"][0]["recordings"][0]["releasegroups"].insert(0, {
|
|
"id": "rg-late", "title": "Later Studio Album", "type": "Album",
|
|
"secondarytypes": [], "releases": [{"date": {"year": 1997}}],
|
|
})
|
|
c = a.parse_lookup_response(resp)[0]
|
|
assert c["album"] == "Machine Head"
|
|
assert c["year"] == "1972"
|
|
|
|
|
|
def test_year_is_earliest_release_not_a_reissue():
|
|
# A group's first-listed release is often a reissue; the year must be the
|
|
# EARLIEST across the group's releases (real case: British Steel's 1980
|
|
# original, not a 2010 reissue listed first).
|
|
resp = _resp(rg_title="British Steel", year=2010)
|
|
resp["results"][0]["recordings"][0]["releasegroups"][0]["releases"].append(
|
|
{"date": {"year": 1980}})
|
|
c = a.parse_lookup_response(resp)[0]
|
|
assert c["year"] == "1980"
|
|
|
|
|
|
def test_dedupes_recording_across_results():
|
|
resp = _resp()
|
|
resp["results"].append(dict(resp["results"][0])) # same recording again
|
|
assert len(a.parse_lookup_response(resp)) == 1
|
|
|
|
|
|
def test_non_ok_status_and_garbage_return_empty():
|
|
assert a.parse_lookup_response({"status": "error"}) == []
|
|
assert a.parse_lookup_response({}) == []
|
|
assert a.parse_lookup_response(None) == []
|
|
assert a.parse_lookup_response({"status": "ok", "results": []}) == []
|
|
|
|
|
|
def test_higher_acoustid_score_ranks_first():
|
|
resp = _resp(score=0.55, rec_id="low")
|
|
resp["results"].append(_resp(score=0.99, rec_id="high")["results"][0])
|
|
out = a.parse_lookup_response(resp)
|
|
assert out[0]["recording_id"] == "high"
|
|
|
|
|
|
def test_config_gating(monkeypatch):
|
|
monkeypatch.delenv("ACOUSTID_API_KEY", raising=False)
|
|
assert a.api_key() == ""
|
|
assert a.is_configured() is False
|
|
assert a.is_configured("explicit-key") is True
|
|
monkeypatch.setenv("ACOUSTID_API_KEY", "envkey")
|
|
assert a.api_key() == "envkey"
|
|
assert a.is_configured() is True
|