What the product does
A candidate uploads a resume. Claude extracts a structured profile; pgvector retrieves candidate listings, sourced from public Department of Labor filings that indicate sponsorship history; an LLM reranker orders the shortlist and writes a rationale for each match. The public homepage is open; the matching flow itself needs an account and a resume.
- Ingest listings from source data this case
- Normalize and chunk into child records this case
- Vector retrieval against the candidate profile unchanged
- LLM rerank unchanged
- Written rationale per match unchanged
The application is built like production even though it is a solo product: row-level security, prompt-injection hardening, per-user quotas, model routing with prompt caching, and daily health checks.
The problem
Job data has to be both usable and consistent. Ingestion runs in the background and writes listings in batches of child chunks. When a run is interrupted, it leaves records in an unfinished state. A repair job that simply re-processes the leftovers introduces a second writer, and two writers touching the same rows produce duplicates and partial writes that are worse than the original interruption.
My role
I own JobSignal end to end. For this change I defined the requirement and the acceptance criteria, designed the guard, and reviewed the implementation, which was produced with AI coding tools under the cross-LLM workflow I run. The design, review, and the call on what counted as done are mine; this page does not claim every line was hand-written.
Decisions
- One shared lease for every relevant write path. The regular ingestion path and the repair path both have to hold the same lease before writing, rather than each checking its own flag.
- Bounded work per run. A run is limited in both processing time and row count, so a repair can never turn into an unbounded rewrite.
- Stop writing after losing the lease. The lease is re-checked between batches, not only at the start; losing it halts subsequent writes.
- Keep the guard at the application layer and say so. This does not claim to solve every database-level concurrency issue; it makes the known failure mode bounded and observable.
How it was verified
- A behavior test that fails on the old code. The regression test reproduces the interrupted-run scenario and fails against the previous implementation before passing with the lease in place.
- A mutation check. With the lease check removed, the repair job kept going and wrote three batches of child chunks. With the guard restored, only the first batch executed. That is the evidence that the guard, and not something incidental, is what stops the second writer.
| Test snapshot | 2026-09-14, candidate commit 45696f8: 94 test files, 1,487 tests. Not re-run for this page and not a live production figure; Playwright end-to-end runs are counted separately. |
|---|---|
| CI | Passing on the merged change. |
Status and boundaries
Code merged CI passing Production behavior acceptance pending
The change is merged and CI is green. Acceptance of the behavior in production, against real interrupted runs, has not been completed yet and is tracked separately. The application-level lease is not presented as a substitute for database-level coordination.
Where to see the evidence
- The live product, including the public demo data notes.
- The pull request, comparison report, and test snapshot behind this case are in a private repository and can be walked through on request.