Analyzing O*NET Occupational Data
Given a row like:
11-1011.00 Chief Executives 4.A.2.b.1 Making Decisions and Solving Problems IM Importance 4.85 30 0.066 4.7182 4.9882 N 08/2023 Incumbent
Read it as: Chief Executives rate "Making Decisions and Solving Problems" at Importance = 4.85/5 (very important), based on 30 survey respondents, with a tight confidence interval (4.72–4.99) indicating high consensus.
| Column | Meaning |
|---|---|
| O*NET-SOC Code | Occupation identifier (e.g., 11-1011.00) |
| Title | Occupation name |
| Element ID | Hierarchical code for the work activity/descriptor (e.g., 4.A.2.b.1 = Generalized Work Activities) |
| Element Name | Human-readable descriptor (skill, activity, or knowledge area) |
| Scale ID / Scale Name | IM=Importance (1-5 scale), LV=Level (0-7 scale, how much/how advanced) |
| Data Value | The rating itself |
| N | Sample size (respondents) — low N (<20) means less reliable |
| Standard Error, Lower/Upper CI Bound | Statistical reliability of Data Value |
| Recommend Suppress | Y = data unreliable, exclude from analysis |
| Not Relevant | Y = this element doesn't apply to this occupation (only meaningful for LV) |
| Date | Data collection/update date |
| Domain Source | Incumbent (job holders) vs Occupational Expert |
Key distinction: Importance answers "how important is this?" (1=Not Important, 5=Extremely Important). Level answers "how much of this is needed?" (0=none, 7=highest degree) — scale differs by element type (e.g., education level vs. activity frequency).
When analyzing an O*NET data extract:
- Filter out rows where
Recommend Suppress = Yunless explicitly asked to include unreliable data - Group rows by Element ID pairs (each element has an IM row and LV row) to get a complete picture per activity
- Check
Not Relevant = Y— flag these as excluded from that occupation's profile rather than scored 0 - Sort by Data Value (descending) within Scale ID to rank top activities/skills
- Note N size — treat ratings with N < 15 as lower confidence
- If comparing occupations, align on identical Element IDs, not just Element Names (names are stable but confirm code match)
- Report both IM and LV together per element when summarizing a job profile — a high-Importance/low-Level activity is a different signal than high/high
Example 1: Ranking top work activities for an occupation Input: Full O*NET row set for Chief Executives (11-1011.00), Element Name family = Generalized Work Activities (4.A.x) Output:
Top 5 activities by Importance:
1. Establishing and Maintaining Interpersonal Relationships — IM 4.80, LV 6.54
2. Making Decisions and Solving Problems — IM 4.85, LV 6.00
3. Communicating with Supervisors, Peers, or Subordinates — IM 4.75, LV 5.82
4. Guiding, Directing, and Motivating Subordinates — IM 4.74, LV 6.07
5. Coaching and Developing Others — IM 4.71, LV 5.85
Notably de-emphasized: Repairing and Maintaining Mechanical Equipment
(IM 1.21, LV 0.27, flagged Recommend Suppress=Y, Not Relevant=Y) — excluded from profile.
Example 2: Interpreting a single row
Input: 4.A.3.b.1 Working with Computers IM 4.17 LV 3.03
Output: "Working with Computers is moderately-to-highly important (4.17/5) for this role, but the required proficiency level is only moderate (3.03, likely on a 0–7 scale) — suggesting the role requires computer use but not deep technical expertise."
Example 3: Data quality check
Input: Row with Recommend Suppress = Y, e.g., Repairing and Maintaining Mechanical Equipment LV = 0.27
Output: "Exclude this data point — flagged for suppression due to low reliability (likely small N or wide CI relative to value). Also marked Not Relevant=Y, confirming this activity doesn't apply to Chief Executives."
- Always cross-reference Element ID, not just name, when merging/joining datasets across files — O*NET reuses similar names across different taxonomies (Work Activities vs. Skills vs. Knowledge use different ID prefixes: 4.A = Work Activities, 2.B = Skills, 2.C = Knowledge, etc.)
- When N is small (<20) and CI range is wide (>1.0 point), caveat any conclusion drawn from that value
- Distinguish Importance from Level explicitly in any summary — never report one number without specifying which scale
- When ranking across occupations, normalize by using percentile rank within each occupation's own distribution if scales differ in practical use
- Watch for malformed SOC codes (e.g., "11-1011.002" instead of "11-1011.00") — likely data entry artifacts; verify against title match before treating as a distinct occupation
- Don't average IM and LV together — they measure different things and are not comparable/summable
- Don't treat "Not Relevant = Y" rows as a Level score of 0 in aggregate calculations — exclude them instead
- Don't ignore Recommend Suppress flags when computing rankings or summaries
- Don't assume Level scales are always 0–7; some Level scales (e.g., education) have different ranges — check context/Scale Name if unsure
- Don't confuse the hierarchical Element ID structure — e.g., "4.A.4.b" is a parent category grouping multiple child elements like "4.A.4.b.1", "4.A.4.b.2"; avoid double-counting parent and child rows if a dataset includes both