Science1 distinct publisher3 min readUpdated
No official O*NET-SOC to ANZSCO correspondence exists. The one just published chains three many-to-many joins, and its author says the European Commission's method is probably better.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The problem is arithmetic before it is taxonomy. The best SOC-to-ISCO correspondence the author could find maps 3,349 categories onto 958 O*NET occupations, and none of the matches are one-to-one [8]. That works out at roughly 3.5 source categories per O*NET occupation [9], and it happens at the first of three joins. ISCO-08 correspondences are published only at unit group level, so OSCA occupations arrive bundled into large groups on the way through [10].
The bundling does not run in one direction, which is what stops the chain being fixed by more careful joining. Engineering Technologist is assigned to several ISCO-08 groupings at once [7]. The ISCO/ESCO category for sports, recreation and cultural centre managers picks up two O*NET jobs, while O*NET's legislators splits across more than one ISCO/ESCO category [11]. On the Australian side, Production Manager (Manufacturing) carries more than one ANZSCO grouping [6]. Many-to-many at every step means the person building the table is deciding, not deriving. The author says so directly: analyst judgement is required to make the match fit the use case, and anyone working on trucking occupations, for instance, should confirm by hand that the data has been sensibly assigned [12].
That instruction is the part with teeth for anyone citing O*NET-derived numbers in an Australian paper. If the mapping requires judgement, the mapping is a result, with as much claim on the methods section as the model that consumes it.
The pointer to the European Commission's approach is unusual enough to sit with [4]. A crosswalk arrives carrying a note from its own author that a better method probably exists elsewhere and has not been reproduced here. Whoever picks up the table inherits that note, whether or not they pass it on.
Then there is decay. OSCA is the new standard and the successor to ANZSCO [13], which the author expects will give this crosswalk a short useful life [14]. The underlying correspondence tables were sourced from O*NET and the ABS on 21 August 2026 [15], so the artefact is pinned to a date and will drift at both ends. The author's own guess at why no official table exists is that OSCA is new and that O*NET-SOC does not match cleanly to ANZSCO or to the intermediate tables anyway [17]. The same mismatch turns up whenever datasets use different definitions of industries, administrative boundaries or products [18].
Worth noting how the work was made: the code was written largely with Claude, while the write-up was left more or less alone [16]. For a data-cleaning and joining exercise, that is the reverse of the usual concern. The thing a reader needs to review is the joins, and the joins came out of a machine. The whole exercise began as groundwork for a 2025 advisory project mapping occupational transition pathways in India [1], where the viability of a move between jobs was judged partly on how similar the two occupations were, conditional on geography, wage differentials and education [2].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
In 2025 the author served as an adviser for a project to map occupational transition pathways in India.
The India project's methodology leaned heavily on studies using O*NET occupational profile data, with the viability of a transition pathway determined in part by how similar two jobs are, conditional on geography, wage rate differentials and education.
There is no official crosswalk between the O*NET occupational taxonomy and ANZSCO, despite Australian researchers frequently using the O*NET database.
The post tells readers to be suspicious of relying on the crosswalk it builds, and suggests a better methodology is probably the one applied by the European Commission.
The crosswalk is assembled in steps: ANZSCO to OSCA, ISCO-08 (or ESCO) to OSCA, and SOC to ISCO-08 (or ESCO), with problems arising at every step from SOC to ANZSCO.
In most cases ANZSCO occupations have been assigned to one or more OSCA group, but the opposite also occurs, such as Production Manager (Manufacturing), which has been assigned more than one ANZSCO grouping.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Concrete and self-documented, but single-source and unvalidated
The post exposes its inputs, code, threshold assumptions and named failure cases, and its central factual claims are internally checkable (3,349 to 958 is arithmetically consistent with the ~3.5 collapse ratio). Against that, everything comes from one self-published item with no independent verification, no accuracy evaluation of the resulting mapping, and no coverage statistics for unmatched occupations.
One published artifact, no observed downstream use
Adoption evidence stops at publication: a single blog release of the crosswalk plus code, with the author's own planned follow-up post as the only stated use. No third-party researcher, agency or product is reported to have taken it up, and the author expects OSCA succession to shorten its life.
Understated by its own author
The framing runs below what the artifact delivers: the post calls itself 'the boring part', instructs readers to be suspicious of the crosswalk, nominates the European Commission's method as probably better, and dates its own obsolescence, while still shipping code, inputs and documented failure modes. Negative rather than strongly negative because the underlying evidence base is a single unvalidated source, so restraint is partly warranted.
Disclosed personal research pipeline, no commercial stake shown
The author's interests are stated plainly: a 2025 advisory role on an Indian occupational transition project motivated the method, the crosswalk exists to enable a planned follow-up post, and AI code authorship is disclosed. No vendor, funder, client deliverable or revenue relationship appears in the source, so the incentive to overstate is low and largely reputational.
Reproducible mechanics, unverified accuracy
Confidence is moderate: the descriptive facts about missing official crosswalks, the join chain and the collapse ratio are well specified and consistent within the source, and the author's own caveats align with the evidence. It is held down by single-publisher sourcing, absent validation of the mapping, and no observed external adoption to corroborate usefulness.
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
product
A five-hour script beats Claude's watermark, so stop treating it as provenance3 distinct publishers
science
3.2 million replies, 88 candidates, and the gender gap that prevalence metrics erase1 distinct publisher
leadership
Daycare Does Not Break Children's Brains, And It Does Not Fix Economies Either1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 20, 2026