Add tests for job cancellation and retry logic in BackendJobs

This commit is contained in:
barlind
2026-06-18 00:41:35 -07:00
committed by Bret Mogilefsky
parent ed4e9fdd67
commit 73dd472f52
5 changed files with 194 additions and 16 deletions
+3 -1
View File
@@ -157,7 +157,9 @@ Diagnostics live under `slopsmith.note_detection_capability.v1` and contain prov
The jobs slice promotes `jobs` as a privileged core provider-coordinator implemented by [static/capabilities/jobs.js](../static/capabilities/jobs.js), with a backend visibility companion exposed through `context["jobs"]`, `/api/jobs`, `/api/jobs/providers`, `/api/jobs/{id}`, and `/ws/jobs`. It coordinates long-running plugin work such as conversion, import, cache-building, update, preview, and studio-style background tasks while providers keep the actual file writes, subprocesses, downloads, and private payloads inside their own code. Backend plugin queues can register/adopt/progress/settle redaction-safe job summaries without requiring their browser screen to be loaded.
The browser command surface is `register-provider`, `unregister-provider`, `list-providers`, `enqueue`, `adopt`, `list`, `inspect`, `cancel`, `pause`, `resume`, `retry`, and `record-bridge-hit`. Provider operations are `job.enqueue`, `job.status`, `job.cancel`, `job.pause`, `job.resume`, `job.retry`, `job.recover`, and `job.adopt`. `list` and `inspect` are prompt-free and side-effect-free: they never trigger provider work, writes, downloads, subprocesses, or external calls. Fresh privileged `enqueue` and `retry` requests require an explicit `authorization: "user-action"` or a matching approved-continuation scope before provider callbacks run. The backend companion is deliberately reporting-first in this slice: provider backends call `register_provider`, `adopt`, `update_progress`, `complete`, `fail`, `cancelled`, `mark_provider_unavailable`, and `record_bridge_hit`; `/api/jobs/{id}/cancel` and `/api/jobs/{id}/retry` return `unsupported-operation` until a backend dispatch/authorization bridge exists.
The browser command surface is `register-provider`, `unregister-provider`, `list-providers`, `enqueue`, `adopt`, `list`, `inspect`, `cancel`, `pause`, `resume`, `retry`, and `record-bridge-hit`. Provider operations are `job.enqueue`, `job.status`, `job.cancel`, `job.pause`, `job.resume`, `job.retry`, `job.recover`, and `job.adopt`. `list` and `inspect` are prompt-free and side-effect-free: they never trigger provider work, writes, downloads, subprocesses, or external calls. Fresh privileged `enqueue` and `retry` requests require an explicit `authorization: "user-action"` or a matching approved-continuation scope before provider callbacks run. Backend provider code can call `register_provider`, `adopt`, `update_progress`, `complete`, `fail`, `cancelled`, `mark_provider_unavailable`, and `record_bridge_hit`; providers may also register private action callbacks for advertised operations such as `job.cancel` and `job.retry`. `/api/jobs/{id}/cancel` and `/api/jobs/{id}/retry` inspect the redaction-safe job owner, enforce the user-action gate for retry, and dispatch to the provider callback without exposing raw payloads to core state, diagnostics, or API responses.
This slice supports dispatch before execution ownership. Core routes authorized backend job actions to provider callbacks (`cancel` for queued jobs, `retry` for failed jobs, and later provider-level actions where the contract is explicit) while providers keep their private queue stores and worker code. Moving scheduling or execution itself into core is a later, separate upgrade for shared worker patterns across multiple providers; it must not erase domain-specific semantics such as queue-level pause/resume, non-interruptible running jobs, remote service retries, or plugin-owned artifact indexing.
Jobs are scheduled by provider capacity and priority (`user-approved-interactive` before `background-maintenance`). State transitions are explicit (`queued`, `running`, `paused`, `cancellation-requested`, terminal cancelled/completed/failed/provider-unavailable/orphaned), and outcomes use the shared canonical vocabulary: `handled`, `queued`, `denied`, `user-action-required`, `unavailable`, `no-owner`, `no-handler`, `no-target`, `unsupported-command`, `unsupported-operation`, `incompatible`, `incompatible-version`, `provider-selection-required`, `validation-failed`, `stale`, `cancelled`, `completed`, `failed`, `timeout`, and `retry-started`.
+2
View File
@@ -66,6 +66,8 @@ The jobs slice promotes `jobs` from a deferred domain to an active privileged pr
Providers keep actual privileged work private. Core stores only safe job summaries, provider metadata, selected/default provider choices, bounded lifecycle history, terminal outcomes, and recovery references. It does not persist raw payloads, active non-recoverable work, DB schemas, paths, filenames, URLs, tokens, command lines, media/artifacts, recordings, live handles, or provider-private values.
Backend action dispatch is the middle ground before core-owned execution: provider route code declares supported actions such as queued-job cancel or failed-job retry; core checks the caller's authorization and job ownership; then it dispatches to provider-owned callbacks while the provider keeps raw payloads, queue semantics, and domain-specific execution. Future jobs upgrades should only promote shared backend scheduling/execution into core if multiple first-party plugins need the same worker behavior for provider-declared recoverable jobs. That larger step must define queue-vs-job actions, restart recovery, cancellation truthfulness, retry attempts, payload privacy, provider failure handling, diagnostics, and migration from plugin-owned queue stores before replacing plugin execution routes.
Jobs bridge removal gates are: bundled and first-party long-running workflows use native `jobs` provider registration/dispatch; normal conversion/import/update/cache smoke runs show no unexpected legacy bridge hits; diagnostics distinguish queued, denied, user-action-required, provider-selection-required, stale, cancelled, completed, failed, timeout, retry-started, orphaned, and provider-unavailable cases; repeated plugin hydration does not duplicate providers or jobs; and reload recovery restores only provider-declared safe references.
## Recommended Next Slices