Methodology

How the corpus is built, how every published figure is computed, and where the limits are.

Figures on this page last computed · published by FindRole

Live postings
23,289
Employers
294
Disclose pay
78%

1. Where the data comes from

Every posting in the corpus is collected from an employer's own career page or applicant tracking system. FindRole does not license or resell a third-party job feed, and does not accept paid placement — an employer cannot pay to appear, to rank higher, or to be excluded from the statistics.

Salary figures are advertised pay: the range the employer published in the posting. They are not self-reported by employees, not survey responses, and not modelled estimates. Where a posting discloses no range, it contributes to the corpus but not to any pay figure.

2. How often it updates

The corpus is refreshed nightly. Each cycle adds newly discovered postings, re-checks existing ones against their source page, and removes roles that no longer appear on the employer's site. Salary pages and market statistics are recomputed from the refreshed corpus on the same nightly cycle, so a figure is at most one day old — and every page carries the date its numbers were computed.

A posting's date is the date the employer published it where that is stated, and the date FindRole first saw it otherwise.

3. AI-extracted fields, and their limits

Job postings are unstructured prose. To make them comparable, a language model reads each posting and extracts structured fields — a normalized job title, seniority level, required skills, and the locations the role can be performed from. These power search filters, the salary groupings and the skill rankings.

This is the part of the pipeline most worth knowing the limits of:

  • Skills are grounded, not invented. Extracted skills are checked back against the posting's own text; an entry the source doesn't support is dropped rather than kept.
  • Seniority is inferred. Where a posting states a level or a years-of-experience requirement, that is used. Where it states neither, the level is inferred from the title, and inference is imperfect — titles are not standardised across employers, and the same words mean different things at different companies.
  • Titles are normalized for grouping. Salary and skill aggregates group by a cleaned title, so "Sr. Software Engineer II" and "Senior Software Engineer" land in the same bucket. The original title is always shown on the listing itself.
  • Extraction is re-run when the source changes. Fields are keyed to the posting text they were derived from, so an edited posting is re-parsed rather than left with stale metadata.

Found a field that's wrong? Email sales@findrole.ai with the listing URL and it will be corrected.

4. How salary figures are computed

A posting's pay is the midpoint of its advertised range. Percentiles on a salary page are computed over those midpoints for every live posting matching that title (and region, on a regional page). Ranges that are obviously not annual base pay — hourly rates, placeholder zeroes, and figures outside a plausible annual window — are excluded, so a contract hourly rate cannot masquerade as a salary.

A small sample produces a confident-looking number that means nothing. So a page is only published once it clears explicit gates, and is automatically unpublished if it later falls below them:

Gate Threshold
Postings with disclosed pay, national page at least 30
Distinct employers, national page at least 5
Postings with disclosed pay, regional page at least 30
Distinct employers, regional page at least 3
Postings in a seniority or remote/onsite band at least 10, from 3+ employers
Maximum share of a band from one employer 50%

Every salary page states its own sample size and employer count, so the strength of any individual figure is visible on the page it appears on.

5. How listing trust is scored

Job boards carry stale reposts and low-substance listings. Each posting gets a trust score, shown as a badge on the listing, derived from signals about the posting itself rather than about the employer: how distinctive its description is against the rest of the corpus (near-duplicate reposts score lower), how substantive the description is, how recently it was published, and whether the same role has been reposted repeatedly.

The score is calibrated against the live corpus, so a typical posting reads as mid-scale rather than alarming. It is a signal about listing quality, not a judgement about the employer, and not a claim that a role is or isn't still open. The insights page publishes the full distribution across the corpus.

6. How employer posting behavior is measured

Company pages carry a short “how this employer posts” summary. Unlike everything above, most of it is derived from listings that are no longer live: when a posting closes, we keep a permanent record of when it first appeared, when it was last confirmed open, and whether the same employer had another listing under the same title at that moment. Nothing about the posting’s content is kept.

  • Reposted is the share of the listings we watched close where the same employer had another live listing under the same title on the day the first one closed. It is a signal, not a verdict: two genuinely separate openings can share one title, and large employers hiring the same role in several places will register here without doing anything questionable.
  • Listings we saw close is a count of observed closures, not of the employer’s total hiring. We only see listings we were tracking.
  • Salary disclosure is the share of that employer’s currently open roles publishing any pay range. This one needs no closure history and is shown for every employer.

Repost figures are published only for employers with at least 20 closed listings on record. Below that a handful of listings moves the share far enough to misrepresent an employer, so the section is hidden rather than shown with a caveat. The record began in September 2026, so these figures describe recent behavior and will get steadier over time.

How long listings stay up is reported as a survival estimate, not an average. The obvious approach — averaging the listings that have already closed — is badly wrong over a short record: the ones that have closed are disproportionately the short-lived ones, while every long-running listing is still up and so contributes nothing. We published such a figure briefly on 2026-09-04 and withdrew it the same day.

The replacement is a Kaplan-Meier estimate, the standard method for this problem. A listing that is still up is not left out; it counts as having lasted at least as long as we have watched it. Listings already running when the window opened enter at the age they had then, so we never credit ourselves with time we did not observe. We report the share still up after 30 days, and a halfway point only where the data actually reaches one — for most employers it does not, and we say so rather than extrapolating. Employers with fewer than 20 closed listings get no survival figures at all.

7. Known limitations

  • Coverage is technology roles, mostly US. The corpus is not a census of the whole labor market, and figures should not be read as one.
  • Advertised pay is not paid pay. Published ranges are wide, sometimes aspirational, and exclude equity and bonus. They describe what employers advertise, not what people earn.
  • Pay-transparency coverage is uneven. 78% of postings disclose a range; disclosure is far more common in states that mandate it, so national medians lean toward those markets.
  • A live listing is not a guaranteed-open role. A posting is removed once it leaves the employer's site, but an employer that fills a role without taking the posting down will keep appearing until they do.
  • The same requisition can appear twice. An employer publishing one role at two URLs can produce two listings; de-duplication is conservative, because collapsing two genuinely distinct roles is the worse error.

8. Using these figures

FindRole's published statistics are free to quote and reuse, including commercially, with attribution and a link to the page the figure came from. Quote the date shown alongside the figure — the numbers move nightly. For a cut of the data that isn't published, or a question about how something was computed, email sales@findrole.ai .