Paper Detail

JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes

Iman Johary, Guillaume Bied, Alexandru C. Mara, Tijl De Bie

arxiv Score 15.2

Published 2026-07-13 · First seen 2026-07-14

General AI

Abstract

Large-scale, richly annotated career trajectory data underpins workforce planning, job recommendation, and labour market analysis, yet publicly available datasets are either small, closed to independent use, or built from pre-standardized occupational codes with LLM-synthesized rather than authentic free text. We present JobHop~v2, an improved version of the publicly available JobHop dataset, constructed through end-to-end large language model (LLM) extraction from a corpus of ${\sim}440{,}000$ pseudonymized, multilingual resumes provided by VDAB, the Flemish Public Employment Service. The released dataset comprises $355{,}315$ career trajectories annotated with ESCO occupational codes, quarter-level temporal information, and normalized five-level education attainment, broadening both the coverage and the annotation richness of the original release. Relative to v1, JobHop~v2 introduces a redesigned extraction pipeline based on reasoning-controlled LLM inference with a retry mechanism (achieving a 100% JSON parse rate), a richer extraction schema, and a revised evaluation protocol scored against three complementary annotation baselines. Evaluated against these baselines, our best extractor comes closest to the inter-annotator agreement ceiling among all compared models, trailing it by only 1.1-2.7 percentage points. The dataset and code are publicly released to support reproducible career-trajectory research.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{johary2026jobhop,
  title = {JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes},
  author = {Iman Johary and Guillaume Bied and Alexandru C. Mara and Tijl De Bie},
  year = {2026},
  abstract = {Large-scale, richly annotated career trajectory data underpins workforce planning, job recommendation, and labour market analysis, yet publicly available datasets are either small, closed to independent use, or built from pre-standardized occupational codes with LLM-synthesized rather than authentic free text. We present JobHop\textasciitilde{}v2, an improved version of the publicly available JobHop dataset, constructed through end-to-end large language model (LLM) extraction from a corpus of \$\{\textbackslash{}sim\}440\{,\}000\$ },
  url = {https://arxiv.org/abs/2607.11715},
  keywords = {cs.CL},
  eprint = {2607.11715},
  archiveprefix = {arXiv},
}

Metadata

{}