No mock data anywhere in the live flow. The GUI, the seam, and the memory corpus were checked end to end: everything Sam sees is a real posting from a real company, fetched from the open web by the live loop. Mocks exist only where they should — confined to CI fixtures and the dormant PR #94 branch, by design. The one hygiene note: rotate the poc-token at deploy time.
Two leftover test documents were found in the DB and deleted. What remains: 28 real jobs — 7 alerted · 6 watched · the rest dismissed/expired — all from real Zurich AI companies: the DeepJudge cluster, Lakera, Rapidata, LGT, Ahead Health, and peers. The corpus the judge reasons over is the corpus you'd want it to reason over.
| Result | Count | What it means |
|---|---|---|
| Direct 200 good | 25 / 28 | The stored link resolves straight to the live posting. |
| OpenAI bot-403 handled | 1 | The careers page blocks non-browser agents. A handled state — the case notes it; a human click works. |
| Superseded aggregator link handled | 1 | The aggregator rotated its URL; the direct employer link in the case is the canonical one. Handled. |
| Dead posting (Accenture) handled | 1 | The posting closed — and the scout correctly auto-expired the case. This is the system working, not failing. |
Why this framing matters: a link audit that only counts 200s would call three healthy behaviours "failures." The honest metric is 25 direct-200 + 3 correctly-handled states = 28/28 accounted for. Zero silent rot.
Ahead Health duplicate. Their site serves a .md variant of the same posting path, and URL-hash canonicalization treated the two as different jobs — so one posting became two cases. The operator reconciled the pair today; the root cause (suffix noise defeats hash-dedup) feeds directly into shortlist item #1 below. This is exactly the kind of defect a real-data audit exists to catch: no test fixture serves .md variants of itself.
We asked the operating agent to narrate its own pipeline — not the spec's version, the actual version — and to name where it's thin. Here is the case-build as it really runs:
web_fetch the candidate pagesThe 0.64 measured similarity floor is only a domain-noise gate — it keeps cooking blogs out, it does not rank jobs. The judge owns fit. The proof is in the re-judge swings that made sense: Google 90→28 (a generic brand-halo score collapsed once the actual requirements were quoted against the actual CV) and Swissquote 46→58 (real evidence surfaced that the first pass missed). A scorer whose corrections move in the direction the evidence points is a scorer you can trust — that's the whole cited-verdict thesis in two data points.
.md duplicate is the proof.Why we publish the thin list: the product's differentiator is honesty in verdicts — that only stays credible if the system is equally honest about itself. Every thin item above reappears, ranked, in the shortlist (§7).
The burn-in weekend surfaced one outage: an OAuth credential expired mid-weekend and the loop went silently dark until a human noticed. The fix wasn't just re-authing — it was making sure the system can never again fail silently.
*/30 steady state — the loop runs every 30 minutes, unattended.poc-token. One bearer token gates the seam. Rotate at deploy; per-user tokens come with the multi-user layer.To judge what the scout is worth, map the whole human job hunt — every stage, its typical time cost, its documented pain — and mark honestly where the scout helps today. The pattern that falls out is the strategy.
| # | Stage | Typical time | The pain | ollscout today |
|---|---|---|---|---|
| 1 | Define target | 2–5 h once | Most people never write down what they actually want; search drift follows | Partial — standing queries + the structured profile encode it, but building them is manual |
| 2 | Discover postings | 3–10 h / week | Board-refreshing is the single biggest recurring time sink | YES — this is our 30-min open-web loop |
| 3 | Company research | 30–60 min / target | Funding, size, culture, reviews — skipped when tired, regretted at the interview | No — zero today (shortlist #3) |
| 4 | Assess fit | 10–20 min / posting | Humans skim and guess; hope inflates fit | YES — the flagship: cited verdicts, quote vs quote |
| 5 | Tailor CV | 15–20 min / app | 48% call this the hardest step of the entire hunt | No — but our gap-verdicts are the perfect input (shortlist #4) |
| 6 | Cover letter | 30 min+ / app | Dreaded; generic AI letters are detectable and hurt | Partial — grounded drafting is the next milestone |
| 7 | Apply / ATS forms | 45 min each · 100 apps ≈ 75 h | Re-typing the same CV into Workday, forever | Deliberate NO — no submit tool, by construction |
| 8 | Track applications | 1–2 h / week | Spreadsheet entropy; lost threads | Partial — applied/replied states exist; no nudges yet |
| 9 | Follow up | mostly skipped | 60% cite black-box silence after applying as the worst part | No — (shortlist #5) |
| 10 | Interview prep | 4–10 h / interview | Scramble the night before | No — but the case file is ideal raw material (shortlist #6) |
| 11 | Negotiate | 1–2 h | 55% never even try | No |
The takeaway: we're strongest exactly where humans are weakest and volume tools are dishonest (discover + assess), absent where the hours bleed (tailor / apply) and where offers are actually won (referrals / prep). The shortlist in §7 moves down this table in leverage order.
Every notable tool in the category, what it's genuinely good at, and what its own users complain about loudest. The complaints are the map — they tell you where the category's trust is broken.
| App | Price | Genuinely good at | Top complaints / honest read |
|---|---|---|---|
| Teal | $29/mo | Great application tracker | Generic AI content; billing complaints |
| Huntr | $40/mo | Kanban pipeline view | Price gripes at that tier |
| Simplify | Free | Autofill: ~90% on Greenhouse/Lever, ~70% Workday | 0% on long-form questions — the part that actually hurts |
| Jobright | ~$20–40/mo | Closest to us; 4.6 Trustpilot | BUT: black-box match score + ghost jobs + billing = its top three complaints |
| LoopCV | paid tiers | — | Auto-apply that fails its core promise; 20% 1-star |
| Sonara / JobCopilot | paid tiers | Volume auto-appliers | 1–2% interview conversion; duplicate-filter blowback — hard numbers validating our no-auto-apply stance |
| Careerflow | freemium | LinkedIn profile optimizer is good | Autofill corrupts form data |
| HiringCafe | Free | Employer-pages-only corpus — superb data hygiene | No alerts. Our 30-min push is literally its missing feature |
| AIHawk | Open source | Proves real self-hosted demand exists | Auto-applier with no quality layer at all |
Read the column on the right as one dataset: the recurring failure modes are unexplained scores, ghost jobs, and auto-apply that doesn't convert. Nobody fails at UI. The category fails at trust.
Every competitor emits a naked percentage. Unexplained and inflated match scores are the recurring #1 product complaint of the whole category — and structurally so: their subscriptions are fed by volume, so honesty is against their business model.
Our cited posting-quote ↔ CV-quote verdicts attack exactly the trust gap the market documents. A score you can audit — requirement quoted, evidence quoted, met/partial/gap ruled — is the one artifact no competitor can copy without breaking its own volume economics. Sharpening this beats adding anything else.
The operator's thin-list (§2) and the market's complaint-list (§5) pointed at the same items — so the shortlist was locked as a sprint on 2026-07-13, split by owner (Claude Code = seam + GUI code, Claw = standing-prompt behaviours), and execution finished the same day: all 6 items live by the afternoon — the operator built the final one (2b expiry sweep) itself. Only the explicitly parked items remain parked.
| # | Item | Owner | Status | What actually shipped |
|---|---|---|---|---|
| 1 | Dedup hardening | Claude Code | ✅ Live 07-13 | Canonicalization strips .md/.html/index suffix noise; job_store warns "possible duplicate of <id>" on a company+title match at a different URL — the seam judges nothing, the operator rules. 92 tests (+18), live-smoked both incident classes, deployed to the running seam, on PR #95. |
| 2a | Ghost/stale guard (pre-alert) | Claw | ✅ Live | GUARD 6 in the standing prompt — a fresh re-fetch of the posting URL immediately before ANY ≥70 alert; dead/redirected → expired, no alert. Wired into the rate-limit degradation path so a degraded run cannot skip it. |
| 2b | Daily expiry sweep | Claw | ✅ Live 07-13 | Operator-built as a separate silent cron job-scout-expiry-sweep (daily ~07:03): re-fetches every alerted/watched/applied URL, expires the dead, tags survivors with still-live-days. Own lock; no Telegram output — it caught the cron-add announce-by-default footgun preemptively and set no-deliver. Read-merge-write on the store, so it never clobbers the lifecycle block the main tick writes. |
| 3 | Company dossier on ≥70 | Claw | ✅ Live | One extra search pass; an optional company:{stage,size,signal,sources[]} block in the verdict contract + one alert line — never folded into the fit score, which stays CV-grounded. |
| 4 | Grounded CV tailoring | Claude Code | ✅ Live 07-13 | Two seam verbs — ollam_cover_letter + ollam_tailor_cv — evidence-anchored: every claim traces to a CV or posting quote, a GROUNDED-IN footer lists the anchors, and abstained cases get an honest refusal instead of a draft (113 seam tests). The Dossier GUI grew DRAFT COVER LETTER / TAILOR CV PLAN buttons with drafts persisted to the store (regenerate = upsert). Proven with one real letter on the DeepJudge 83 case — only real employers, the law-domain gap addressed honestly. |
| 5 | Lifecycle nudges | Claw | ✅ Live | Self-bootstraps lifecycle.applied_seen_at in verdict_note (no per-record timestamps exist in the tool schema — an honest workaround); nudges once at ≥7 days, max 1/week/case. |
| 6 | Interview-prep packs | — | Parked | Until after #4 — the case file already contains the raw material. |
The live proof of #4 caught a grounding slip — the model named a framework absent from the CV — so the no-invent rules in both seam and GUI now explicitly forbid tools/frameworks not present in the evidence.
Two adjacent bugs were fixed by the operator during the same patch: stale burn-in cadence language in the standing prompt, and the run-lock threshold raised 20→25 min under the */30 cadence. First live tick under the v3 prompt is due within 30 min of 09:07 — verification in practice, not just on paper.
Auto-apply Never (1–2% interview conversion + ToS blowback — market-validated, unambiguous) and volume dashboards ("you applied to 200 jobs this week!"). The competitor field proves these are the losing bet; our economics don't need them, so we get to be honest.
The Dossier direction won the design study; v2 keeps it the case file where actions live on the same sheet as the evidence.
After the sprint, Sam ordered a thorough review — and THREE independent reviewers ran in parallel: adversarial code-attack agents, the operating agent (Claw) reviewing its own runs, and an external QA/UX review (Claude Desktop driving a real browser). Each found a distinct class of real bug the others missed. That is the day's architectural lesson: no single reviewer — not even a good one — sees the whole failure surface. Defense-in-depth review is where correctness came from.
| Reviewer | Vantage point | The bug class only IT could see |
|---|---|---|
| Adversarial code attack | Reads the code hostilely, proves claims by execution | Latent data-destruction and injection paths that never fire in a happy-path demo |
| The operator (Claw) | Lives inside the runs — sees its own inputs, locks, and prompt budget | Operational races and a real prompt-injection sitting in the stored corpus |
| External QA (real browser) | A stranger's hands on the actual GUI, no context | Output-quality and UX failures the builders were blind to |
Seam: NO-SHIP, two HIGH bugs proven by execution. (1) The store persisted the canonical URL, so .md-stripping could hand out dead links — github-blob postings 404'd when the stripped URL was served back. (2) Every status transition silently wiped score + verdict_note — the grounding was destroyed the moment a job was marked applied. Plus a CONSTRUCTED prompt-injection: a malicious posting could forge a byte-identical CV-EVIDENCE block above the real one.
Operator verification of v4 (post-publish same hour): the operating agent verified all 17 tools live, found and repaired a live data regression (a dedup-merge had overwritten one verdict with prose — honest re-judge landed Ahead Health at 68 on a real healthcare-vs-fintech domain gap), exposed the legacy-id ambiguity (pre-v4 records resolve by document_id, not URL — seam guidance patched), retired its cache workaround, and migrated both crons to document_id-keyed read-merge-write.
GUI / watchdog: two more HIGHs, both reproduced. The GUI's server actions were invocable from the LAN (bound 0.0.0.0 — actions POSTed from another machine executed). And the watchdog's heal path, run overnight with a locked Keychain, would TRUNCATE the container's credentials to zero bytes and report HEALED — causing the very outage it guards against, then lying about it.
ollam_job_get — the read half of read-modify-write the operator never had; document_id on list rows; contract: v4 markers so contract drift is detectable instead of chat-archaeology.AGENTS.md silently truncating ~25% at session load — a failure class worth remembering platform-wide: config exceeding an injection budget with zero error.suppressHydrationWarning.[Your Name] placeholder, and weak-model prose (rubric/model/preflight round queued, needs Sam's exemplar letter); and pending-approval governance for agent-learned sources — a memory-poisoning surface (queued with operator coordination).a/p/l/d · word-boundary truncation with expand · provenance folded into a case-history collapse · aria/headings · "grounded on N exhibits" · pager.Sam pivoted the evening's mandate: "we dont want new features instead better quality and ui/ux… we might launch sooner than planned." So the round built nothing new. It did three things instead: walked every feature in a real browser, audited the prompts and retrieval that make the verdicts, and polished the rough edges the walk found — all same evening.
Every feature row below was clicked through in the actual GUI against the live corpus — so this matrix doubles as the product's feature documentation, verified-as-working rather than as-intended. Core all PASS; the ROUGH rows became the polish list (§4) and were fixed the same evening.
| Feature | Verdict | One-line note |
|---|---|---|
| Case rail | PASS | Groups, counts, independent scroll — all correct on 58 cases |
| Dual-mode filter | PASS | Instant local narrowing + semantic search at a measured 82–90 ms |
| Exhibits | PASS | Honest truncation, every quote cited to posting or CV |
| Sticky decision panel | PASS | Stays pinned through long case sheets |
| Sub-scores | PASS | Per-dimension breakdown renders and sums honestly |
| Stamps | PASS | Tier/status stamps match the stored verdicts |
| Closest-miss | PASS | The nearest-gap callout is real, not decorative |
| ≥70 letterhead gate | PASS | Draft actions correctly gated to alert-tier cases only |
| Lifecycle chips | PASS | Full path clickable; "tracked since" stays honest about what's known |
| Rule-yourself persistence | PASS | Human rulings survive reloads and re-ticks |
| Draft versioning | PASS | Regenerate = upsert; history preserved |
| Tick band | PASS | Renders REAL sweep data — is-it-alive at a glance, no synthetic ticks |
| Corpus page | PASS | The full stored corpus, browsable and truthful |
| Keyboard verbs | PASS | a/p/l/d rule from the keyboard |
| Honest not-found | PASS | Bad case id → honest state, no crash, no fake sheet |
| Semantic no-match honesty | ROUGH | Gibberish queries returned confident hits → honesty floor added (§4) |
| Ruling/draft pending states | ROUGH | No feedback while filing → "filing…" pending states added (§4) |
| Failed-draft presentation | ROUGH | Failed drafts rendered like successes → now LOOK failed (§4) |
| Search-hit hygiene | ROUGH | Raw entities + near-duplicate families in hits → decode + dedup (§4) |
| Draft signer | ROUGH | A letter signed "[Your Name]" → real signer, always (§4) |
| Cross-section context | ROUGH | Which role am I looking at? → role subtitles across sections (§4) |
| Closest-miss fallback | ROUGH | No full match → blank; now falls back to PARTIAL honestly (§4) |
| 404 / robots | ROUGH | Framework default 404, no robots → house-voice 404 + robots (§4) |
| Case-history & letterhead detail | ROUGH | Label + ellipsis nits → "ARCHIVE & RULED", word-safe tick-band ellipsis (§4) |
The audit (which converged with the operator's independent self-audit — two readers, same findings) put the weakest link exactly where quality is made: the judge.
abstain_reason the whole system reads downstream could never actually be stored under the prompt's own instructions.Cron prompts rebuilt as v6/v4 — drafted by the operator itself, applied under four-eyes review. What went in: an injection guard at the judging moment; anchor bands (88–95 = all-MET → <45 = majority-GAP) plus tier-edge anchors, plus a worked example that teaches the German-clip lesson from the real incident; cv_read-before-NO-EVIDENCE (the judge must read full CV sections before ruling a gap); abstention now resolves to dismissed with a stored reason; and ~35% prompt compression.
Code side: prompt contract p1 now byte-identical seam↔GUI (hard rules first+last, fencing, 9k budgets, LANGUAGE line) · floor baked to the measured 0.64 · interest clips 400 chars · list caps raised · ef_search pinned 80 · gateway token rotated — the day-one committed default is finally dead. Seam: 185 tests.
cv_read + the anchor bands; the residual risk is on record, not waved away.The close: zero new features shipped in this round — deliberately. The round only made what already exists honest, consistent, and demonstrable. That is what "we might launch sooner than planned" needs.