Role-Liveness Investigator
A tool that takes a job link and checks the evidence (is the posting still live, how fresh is it, has it been reposted, what is the company doing) before answering apply now, quick apply, wait or skip, with each reason citing a dated source.
Problem
Some job postings were filled months ago and never taken down. Some are reposted again and again, and some belong to companies that have just frozen hiring. Tailoring an application for one of them is time spent on a role that was never really open.
The Role-Liveness Investigator takes a job link and checks the evidence first: whether the posting is still live on the company's own hiring system, how recently it was published or refreshed, whether it has been reposted or open for a very long time, and whether the company is hiring on the same team or in the news for layoffs or a freeze. It then gives one of four answers, each with short reasons that link to the dated source they rest on. It never calls a job "fake"; it reports what the evidence shows and how strong that evidence is.
| Answer | Meaning |
|---|---|
apply_now |
Strong, fresh evidence the role is real and wanted. Worth real effort. |
quick_apply |
Probably open, but the evidence is thin. Apply with little tailoring. |
wait |
Something is unresolved, for example news of a hiring freeze. Check again in a set number of days. |
skip |
Closed, or a long-running repost with signs the company is not really hiring. |
The research question is narrower than the product: can an LLM agent that chooses which checks to run reach the same answers as running every check, while running fewer of the expensive checks than a fixed rulebook does?
How it works
- Collect. Every day a systemd timer records the public job boards of 351 tech companies that use Greenhouse, Ashby or Lever (it runs at 00:05 and 12:05 and catches up after sleep). Older copies of those boards come from the Wayback Machine, and dated company events (funding, layoffs, freezes) from news headlines. This history is what makes reposts and long-open roles visible. A posting's exact closing time is rarely observed, so a closure is stored as an interval between the last time it was seen open and the first time it was seen gone.
- Investigate. For a given link, a resolver and a board snapshot always run. Small checks called probes can then add evidence: repost history (low cost), requirements drift and company events (medium), and a team signal built from the company's own board history (high). Each probe returns timestamped facts, not verdicts.
- Decide. A fixed, rule-based policy turns the evidence into one of the four answers. The decision is always made by these rules, never by the LLM.
- Explain. The answer comes with reasons that cite evidence ids, and each cited id is checked against the evidence.
- Three ways to pick probes. A (full) runs every probe and is the reference. B (rules) follows a fixed checklist. C (agent) lets an LLM propose which probes are worth running, while a deterministic controller enforces budgets, step limits and the allowed tools. A question counts as unresolved while one of the policy's inputs is still empty, and a probe is eligible only if it could fill one, so the controller stops C as soon as nothing left could change the answer.
- Delivery. A command-line tool, and a FastAPI service with a small web page.
POST /investigatereturns the answer and falls back to the rules, markeddegraded: true, when the LLM is unreachable.POST /outcomeslogs what happened after applying, and/watchrechecks a role when it is due.
Architecture
In words: for a job link, the resolver and a board snapshot always run and build a case file. The LLM investigator reads it and proposes probes, or proposes to stop. The deterministic controller drops probes that are ineligible, off the allowlist or over budget, and either runs the best remaining one or stops. Each probe's evidence goes back to the investigator (the right-hand rail), and the loop repeats until the controller stops it because no remaining probe could change the answer, a budget or step cap is reached, or a call would repeat (the left-hand rail). The fixed action policy then picks the answer from the evidence, and a second LLM call writes reasons that must cite existing evidence ids.
Stack
- Language and tooling: Python 3.12, uv, a Typer command-line interface
- Storage: SQLite, which also holds the run traces (every probe, controller decision, model call, cost and latency)
- Collection: httpx behind a per-host allowlist (exact host match, HTTPS only, private hosts rejected, redirects checked per hop), the Greenhouse and Ashby public APIs, Lever's public postings, JSON-LD parsing with BeautifulSoup, and the Wayback Machine
- Validation: Pydantic for probe arguments and for the JSON schemas the model's output must match
- LLM: any OpenAI-compatible endpoint; the default is Mistral's
ministral-8bon the free tier, and a local Ollama model keeps everything on the machine - Analysis: lifelines for interval-censored closure curves
- API and UI: FastAPI and a vanilla-JS page, on localhost unless an API token is set
- Testing: 1,446 tests with pytest, plus ruff
Results and evaluation
| Companies tracked | 351 |
|---|---|
| Job postings seen | 39,696 |
| Posting closures observed | about 21,000 |
| Days of my own daily collection | 23 (since 2026-09-07) |
| Wayback board captures | 1,452 |
| Dated company events | 773 |
The evaluation is a point-in-time replay. Each case is a posting at a past date T, on a 7-day grid, and the system may only see evidence that was available on or before T. Replay runs offline: a live tool call raises an error and is written into the trace, and every run is audited for future-data leaks. There are two development splits, one by time and one by company.
| Measure | Time split | Company split |
|---|---|---|
| Postings / companies | 300 / 246 | 200 / 138 |
| Cases built / scored | 7,905 / 7,618 | 2,683 / 2,534 |
| Expensive probes per case, agent (C) vs rules (B) | 0.98 vs 1.43 | 0.98 vs 1.55 |
| Ratio C / B (gate: at most 0.70) | 0.68 | 0.63 |
| Agent's answers matching the full system (A) | 100% | 100% |
| Rules' answers matching the full system (A) | 100% | 100% |
| Future-data leaks | 0 | 0 |
| Agent gate | pass | pass |
"Expensive" means the medium and high cost tiers. The gate requires C to use at most 70% of B's expensive probes while staying within 2 points of B's agreement with A, both overall and averaged per answer. Agreement is 100% on both measures for both systems. Scored cases leave out those whose posting could not be assigned to a split (287 and 149), since an unassigned posting could belong to the held-out set. C ran on Mistral's ministral-8b on the free tier, with four workers: all 10,588 cases took about 3 h 15 min.
Agreement has to be read next to the answer distribution, because a policy that nearly always gives the same answer agrees with itself easily. Across the 7,618 scored time-split cases, all three systems answered quick_apply 7,342 times, skip 227, apply_now 45 and wait 4. 7,115 of those cases come from before my own daily snapshots began, so their only board evidence is Wayback captures, which count as weak, and every one of them came out quick_apply or skip. The cases from after daily collection began are the ones that look like real use:
| Split | Cases | apply_now | quick_apply | wait | skip |
|---|---|---|---|---|---|
| Time split | 503 | 45 | 313 | 4 | 141 |
| Company split | 465 | 60 | 327 | 25 | 53 |
Even here most cases are rated weak (439 of 503 in the time split), so quick_apply dominates. Tracing that, I found the main cause on 2026-10-03: the daily snapshot fetched the Greenhouse and Ashby publish dates and then threw them away, which accounted for 435 of those 439 weak cases. Publish dates are now stored with every capture, and the datasets will be rebuilt and rerun once dated captures have built up.
What I don't claim
- That the advice is right. Matching the full system shows the agent is consistent with it, not that its answers are correct. That needs real application outcomes, none are logged yet, and the product value is unproven.
- That the agent is more accurate than the rules. The rules also match the full system on every case; the agent's advantage is running fewer expensive checks.
- Final numbers. The policy is not frozen yet, and the held-out test split has not been touched.
- A full-size company split. It has 200 postings, below the 300 the spec sets as the target for a headline evaluation.
Decisions and trade-offs
- The LLM never decides. It only proposes probes and writes the explanation. Its output must match a schema, it can reach only allow-listed sites, it has hard budget and step limits, and it cannot repeat a call. The answer itself always comes from the fixed policy.
- No looking ahead. Every piece of evidence carries the time it became available, and a replay at time T sees only evidence available on or before T. Archived evidence is dated by its capture, never backdated to when the underlying event happened.
- No guessing. A failed or missing capture is recorded as a gap, never as "job closed", and a closure is kept as a date range rather than an invented date.
- Reasons are checked against evidence. In the time-split report, all 10,988 reasons cite evidence ids that exist.
- Job and news text is data. Posting and news text is delimited in prompts and treated as data, never as instructions to the model.
- Probe savings are measured against the rules, not the full system. A runs every eligible probe whether or not it could change the answer, so any probe count compared with A flatters a leaner system by construction. That is why the gate compares C with B.
- A learned probe ranker was tried and not kept. The spec allows replacing the deterministic, cost-aware probe ranking with a small learned model. On the time split its training label had only one class, so it could not be compared with the deterministic ranking, and nothing was kept.
- Deliberately out of scope. No "ghost job" verdicts, no LinkedIn scraping, no crowd scores and no made-up 0–100 scores.
- What leaves the machine. The posting title, URL, company name, evidence snippets and news headlines are sent to the LLM provider, and free tiers may train on them. Pointing the config at a local Ollama model keeps everything local.