Purpose

This article introduces the Integrated Sociotechnical Exposure Framework (ISEF) to assess generative AI exposure in civil and environmental engineering by separating technical AI suitability (TA) from institutional resistance (IR). The contribution is an operational measurement framework, not a claim that institutional constraint is a new concept.

Design/methodology/approach

The framework is applied to 674 O*NET tasks across 28 occupations. Task-variable pairs are scored by three large language models under a fixed rubric and then aggregated to occupation level. Exposure is modelled with a multiplicative form, and sensitivity checks compare additive, min-rule, and geometric alternatives. A small external practitioner survey is used as a calibration check rather than as task-level validation.

Findings

A large group of occupations score well on technical AI suitability but stay heavily constrained by liability, compliance, and sign-off requirements. The cleaned practitioner survey points in the same direction: across 15 task vignettes, higher perceived institutional constraint is strongly associated with lower adjusted AI exposure.

Research limitations/implications

The scoring still leans on LLMs, and the practitioner sample is small. Because no same-task expert panel is available in the present revision, the results should be read as a transparent first-stage index. The nine-response practitioner exercise is preliminary calibration, not validation of the occupation-level index or its constructed adjusted-score analogue. The natural next step is a larger same-task expert panel paired with adoption-linked field evidence.

Practical implications

For engineering organisations, adoption strategy is better anchored in accountable workflows, auditability, and professional review processes than in model capability on its own.

Originality/value

The paper pulls governance constraints into the exposure measure itself, and uses that lens to explain why technical-only indices tend to overstate near-term generative AI uptake in regulated engineering work. Its measurement logic is potentially transferable to other regulated professions, but any such application requires domain-specific specification and calibration of the institutional-resistance construct.

Civil and environmental engineering is a good place to take the phrase “AI exposure” seriously. Generative systems can now draft reports, produce code, and assist analytical work at a level that would have looked implausible only a few years ago (OpenAI, 2023; Bubeck et al., 2023). Engineering decisions still pass through licensure, liability, and review chains that were built long before these tools existed. Whether AI can produce a particular output, and whether that output can enter an accountable workflow, are two different questions; much of the current confusion around generative AI in engineering comes from treating them as one.

The exposure literature has built strong instruments for the first question. Task-content measures identified codifiable work, routine procedures, and pattern-based activity as the natural unit of analysis (Autor, 2015; Frey & Osborne, 2017). That foundation remains useful for regulated engineering, but it is incomplete. A task may read as technically suitable for AI assistance and still sit inside a sign-off, audit, or liability structure that keeps AI-generated outputs out of the accountable record (Jasanoff, 2016).

This paper develops the Integrated Sociotechnical Exposure Framework (ISEF) in response. ISEF is a task-level framework for civil and environmental engineering that scores technical AI suitability and institutional resistance through the same rubric-based protocol, so neither side gets reduced to background. Its novelty is practical and methodological: it turns a familiar sociotechnical insight into a reproducible occupation-level measurement design.

Research questions. Three questions guide the analysis:

  1. How can technical AI suitability and institutional resistance be operationalised separately for civil and environmental engineering tasks?

  2. How do occupation-level exposure patterns change when institutional resistance is incorporated into technical exposure measurement?

  3. How sensitive are the resulting rankings and groupings to modelling choices, including functional form and clustering assumptions?

The multiplicative form developed below treats mandatory sign-off and similar requirements as binding limits on workable exposure, not as small adjustments at the margin. The analysis covers 28 engineering occupations and 674 O*NET task statements (Peterson et al., 2001). Throughout, exposure is used as a descriptive indicator of where technical suitability meets institutional permission, and it is kept separate from claims about adoption timing or labour displacement.

This section positions the framework in four bodies of work: task-based exposure measurement, professional regulation, sociotechnical governance, and digital transformation in engineering workflows. The review is not used to claim that institutions have been ignored in prior scholarship. The point is narrower: existing literatures explain why adoption is institutionally mediated, while the present paper turns that insight into an occupation-level measurement protocol.

Task-based exposure measures have a long lineage in labour economics. The early generation linked routine task content to substitution risk and set the analytical template later work would extend (Autor, 2015; Frey & Osborne, 2017). Labour-market studies then connected task displacement to job polarisation and wage inequality, while also showing why occupation-level averages can hide task-level variation (Goos, Manning, & Salomons, 2014; Arntz, Gregory, & Zierahn, 2017; Acemoglu & Restrepo, 2022). Subsequent studies narrowed the frame further, moving from whole occupations to the specific capabilities a machine-learning system can perform (Brynjolfsson, Mitchell, & Rock, 2018; Webb, 2020). Reviews of AI, labour, and productivity also stress that exposure measures need to be separated from realised firm-level adoption (Raj & Seamans, 2019). Generative AI has revived that agenda: recent indices ask whether language models can perform or assist with occupational tasks (Eloundou, Manning, Mishkin, & Rock, 2023; Felten, Raj, & Seamans, 2021; Tolan et al., 2021), while experimental and field evidence shows that gains are task-contingent rather than uniform (Noy & Zhang, 2023; Dell’Acqua et al., 2023). These measures handle technical fit competently, but they remain thinner on the institutional side of adoption, which is the part of the problem this paper focuses on.

ISEF keeps task-level measurement but separates two dimensions that exposure scores typically fold together. Technical AI suitability (TA) asks whether generative AI can plausibly support the task. Institutional resistance (IR) asks whether responsibility, compliance duties, data governance, or organisational accountability make practical use harder. In engineering, the second question often matters as much as the first: a useful draft or calculation still has to clear review before it becomes actionable evidence.

Institutional accounts help explain why technical fit alone tends to undersell adoption frictions. Abbott (1988) argues that professional jurisdictions are sustained through expertise, credentialing, and control over what counts as legitimate practice; Freidson (2001) reads that control as a durable institutional logic, not a transitional friction. The implication for digital tools is sharper in Susskind and Susskind (2015): technology can redistribute the tasks professionals carry out without redistributing authority, sign-off, or liability. STS and organisational scholarship add the safety-critical angle. Jasanoff (2004) treats science and technology as co-produced with social order, and Jasanoff (2016) extends the argument to how risk technologies become acceptable—through standards, review procedures, and accountability arrangements, alongside any efficiency case. Sociomaterial and algorithmic-work research makes the same point at organisational scale: digital systems alter work through routines, documentation, and control relations rather than through technical capability alone (Leonardi, 2012; Kellogg, Valentine, & Christin, 2020).

Civil and environmental engineering sits at the sharp end of this argument because outputs are tied to public safety, environmental compliance, procurement rules, and professional sign-off. AI can help with drafting, checking, simulation support, and documentation, but the institutional layer still decides how those outputs are reviewed, recorded, and accepted. AI governance scholarship is useful here because it treats accountability, transparency, risk classification, and auditability as design conditions rather than optional afterthoughts (Calo, 2017; Floridi & Cowls, 2019; Raji et al., 2020; Veale & Borgesius, 2021; Weidinger et al., 2022). ISEF scores this institutional layer at task level using the same rubric-based logic applied to technical suitability, so the two sides of the problem stay comparable.

Engineering research has tracked the spread of BIM, digital twins, machine learning, reservoir modelling, and other data-driven tools in infrastructure and subsurface work (Bock, 2015; Sacks, Eastman, Lee, & Teicholz, 2018; Mohaghegh, 2017; Azimi, Eslamlou, & Pekcan, 2020). Generative AI adds a particular complication on top of this trajectory: outputs can read as competent while remaining hard to verify, attribute, or certify. Foundation-model and generative-AI reviews make this verification problem especially salient because general-purpose models can be repurposed across domains faster than institutional controls can be standardised (Bommasani et al., 2021; Dwivedi et al., 2023). That gap carries weight when a flawed design note, permit draft, or environmental assessment has legal and safety consequences.

This distinction matters for Digital Transformation and Society because generative AI is not simply another productivity tool inserted into an existing workflow. In regulated engineering, model outputs become useful only when they can be documented, checked, attributed, insured, and signed off. A measurement framework that scores technical suitability without also scoring those governance conditions will tend to overstate the near-term exposure of safety-critical professional work.

Exposure is therefore treated here as a sociotechnical construct: a description of where task suitability and institutional constraint meet, rather than a forecast of job loss or a snapshot of current firm-level uptake. The originality of ISEF lies in operationalisation. It combines a task-level rubric, a multi-model scoring protocol, and a gating specification that can be recalculated by other researchers. The empirical contribution remains descriptive, and the LLM-based scores should be read as a first-stage measurement exercise—one that can be inspected, challenged, and extended with stronger practitioner validation.

The study examines 28 civil and environmental engineering occupations from the O*NET database (version 28.0) (Peterson et al., 2001). The final sample consists of 674 unique tasks across managers, engineers, technologists/technicians, and related professionals. Table 1 provides an overview of the sample structure.

Table 1

Sample structure by occupation category

CategoryOccupationsTasks
Managers484
Engineers10248
Technologists/Technicians7169
Scientists & Others7173
Total28674
Source(s): O*NET 28.0

Technical and institutional factors are evaluated using the Generative AI Impact Index (GAII), which consists of 12 variables (seven TA; five IR). Table 2 provides a quick reference, while the full definitions and rubric are available in Supplementary Material B. The rubric is written out in full so that scoring decisions can be checked line by line.

Table 2

ISEF variable quick reference

VariableTypeIntuition
Cognitive/Data AnalysisTAStructured reasoning, statistical modelling, data interpretation
Technological IntegrationTADigital workflows, IoT/BIM coordination
AI ReadinessTARule-based, procedural tasks suitable for generative AI assistance
Automation PotentialTAGeneral suitability for hybrid human-AI execution
CollaborationTATeamwork, stakeholder coordination
Innovation/SynergyTAPattern discovery, creative synthesis
Domain ExpertiseTASpecialised engineering knowledge and judgement
Compliance/ResponsibilityIRLegal liability, regulatory sign-off requirements
Industry BarriersIRProfessional gatekeeping, institutional inertia
Algorithm LimitationIRTasks requiring emotional intelligence, negotiation
Data Security/PrivacyIRSensitive data, confidentiality constraints
Sociotechnical ChallengesIROrganisational politics, public trust issues

The variable list was built through an iterative review of three sources: task-based AI exposure studies, sociotechnical and professional-regulation concepts, and civil engineering workflow characteristics. The TA variables capture recurring forms of task suitability in prior exposure measures, including data analysis, structured procedures, digital integration, and augmentation potential. The IR variables capture barriers repeatedly identified in professional and safety-critical domains, including legal responsibility, privacy, organisational trust, and sectoral gatekeeping. The list is not claimed to be exhaustive; it is a parsimonious operationalisation designed for transparent scoring and future recalibration. Domain Expertise is treated consistently as a TA variable because specialised knowledge can make a task more amenable to AI-assisted retrieval, drafting, or checking when appropriate expert oversight remains in place. Institutional accountability for that expert judgement is captured separately by Compliance/Responsibility and related IR variables.

The analysis measures exposure: the joint degree to which tasks are technically suitable and institutionally permissible. It does not forecast replacement or job loss.

The task–variable mapping protocol is set out in Supplementary Material A. The complete measurement protocol, defined in Supplementary Material B, documents every scoring rule used in the analysis. For each task, the initial mapping was conducted with a focus on the specific context of civil engineering. This conservative protocol involved reading task descriptions and identifying relevant variables supported by the text. The mapping was checked for internal consistency before model scoring, but it should not be confused with independent task-level expert validation. Supporting diagnostic tables and figures are provided in Supplementary Material C.

Task-variable pairs were evaluated using three large language model systems under fixed settings and a prompt template: ChatGPT-5, DeepSeek-R1, and Qwen-3. The same prompt, rubric, and aggregation code can be rerun by other researchers without proprietary tuning.

Design considerations for the scoring process included:

  1. Three independent evaluation runs per model to mitigate session effects.

  2. A single prompt template requiring a 0–10 score with a brief justification tied to the task text.

  3. No post-hoc cleaning of outliers to maintain a realistic assessment of variance.

  4. Clearing of session history between runs to ensure independent measurements.

Consistency was monitored across runs and stayed within the limits acceptable for a descriptive measurement exercise. Inter-model agreement is useful but cannot stand in for external validity. The three systems likely share training data, alignment pressures, and shared assumptions about professional work, so their consensus is treated here as a scalable first-stage signal rather than a substitute for domain-expert scoring. In response to reviewer concerns about circularity, a small practitioner survey was added during the revision period to check whether the TA/IR split survives outside the model-scoring loop; respondents rated 15 task vignettes from the questionnaire on AI assistance potential and institutional constraints. This survey is reported as calibration evidence only. It does not estimate expert–LLM agreement on the same tasks, and the manuscript therefore avoids presenting LLM consensus as independent validation.

The aggregation of task-level evaluations into occupation-level indices follows a structured procedure:

  • Stage 1: Consensus Scores. A mean is computed across all systems and runs for each task-variable pair (Equation 1):

  • where stv(m,r) is the score for task t and variable v from system m and run r⁠, with M=3 systems and R=3 runs per system.

  • Stage 2: Variable Aggregation. Evaluations are aggregated to the variable level using O*NET importance weights (Equation 2):

  • where wto is the normalised importance weight for task t in occupation o⁠, with ∑t∈owto=1⁠.

  • Stage 3: Index Construction. TA and IR are computed as unweighted means of their constituent variables and rescaled to a 0–100 range (Equation 3):

  • Exposure is then calculated through the multiplicative specification (Equation 4), incorporating the institutional context:

The multiplicative specification is used as the primary model because it represents institutional resistance as a partial gate on technical capability. The overall ISEF procedure is visualised in Figure 1. Exploratory K-means diagnostics are then applied to the standardised occupation-level scores to identify broad patterns within the data; the resulting labels are not treated as validated occupational strata.

Figure 1
A flowchart illustrating the methodology overview of the Integrated Sociotechnical Exposure Framework (ISEF).A flowchart titled 'Integrated Sociotechnical Exposure Framework (ISEF) Methodology Overview' depicts the process of ISEF. The flowchart starts with '1. Task Universe' which includes 674 tasks across 28 occupations. This leads to '2a. TA Variables' with seven forward indicators of technical suitability, and '2b. IR Variables' with five indicators of liability, privacy, and governance friction. These variables feed into '3. Scoring Protocol' using ChatGPT-5, DeepSeek-R1, and Qwen-3 with three runs each. The next step is '4. Aggregation' where importance-weighted task scores are aggregated to the occupation level. This leads to '5. ISEF Exposure' calculated as Exposure = TA x (1 - IR/100) reported as a descriptive indicator. The final step is '6. Robustness and Calibration' involving alternative functional forms plus preliminary practitioner calibration.

Complete ISEF workflow for the 674-task, 28-occupation sample: task universe; TA and IR indicators; three-model rubric scoring; occupation-level aggregation; ISEF exposure calculation; and functional-form sensitivity plus preliminary practitioner calibration. Domain Expertise is treated as a TA variable, while accountability and sign-off barriers are captured through IR variables

Figure 1
A flowchart illustrating the methodology overview of the Integrated Sociotechnical Exposure Framework (ISEF).A flowchart titled 'Integrated Sociotechnical Exposure Framework (ISEF) Methodology Overview' depicts the process of ISEF. The flowchart starts with '1. Task Universe' which includes 674 tasks across 28 occupations. This leads to '2a. TA Variables' with seven forward indicators of technical suitability, and '2b. IR Variables' with five indicators of liability, privacy, and governance friction. These variables feed into '3. Scoring Protocol' using ChatGPT-5, DeepSeek-R1, and Qwen-3 with three runs each. The next step is '4. Aggregation' where importance-weighted task scores are aggregated to the occupation level. This leads to '5. ISEF Exposure' calculated as Exposure = TA x (1 - IR/100) reported as a descriptive indicator. The final step is '6. Robustness and Calibration' involving alternative functional forms plus preliminary practitioner calibration.

Complete ISEF workflow for the 674-task, 28-occupation sample: task universe; TA and IR indicators; three-model rubric scoring; occupation-level aggregation; ISEF exposure calculation; and functional-form sensitivity plus preliminary practitioner calibration. Domain Expertise is treated as a TA variable, while accountability and sign-off barriers are captured through IR variables

Close Figure 1

K-means clustering is applied to standardised occupation-level scores as a diagnostic exercise. Solutions for k=2 through k=10 are evaluated using the elbow method, silhouette scores, and gap statistics (Tibshirani, Walther, & Hastie, 2001). With only 28 occupations, clustering carries clear risks: small shifts in input weighting can move borderline cases between groups. The main presentation reports the original k=5 diagnostic for transparency, but interpretation is deliberately collapsed to broad lower-, middle-, and higher-exposure patterns rather than discrete occupational strata. Supporting diagnostic tables and figures are provided in Supplementary Material C.

One feature of the k=5 diagnostic needs a direct word. The singleton diagnostic category contains only one occupation, Engineering Teachers, Postsecondary. The most plausible reading is structural: O*NET task wording for teaching roles is heavy on instruction, curriculum design, and student assessment, which sits awkwardly against the design-and-permit vocabulary that dominates the rest of the sample. The singleton is therefore retained only as a transparent diagnostic artefact. It is not interpreted as a cluster and is not used to support substantive claims. Bootstrap stability is also only moderate in the wider solution; the discussion therefore stays at the level of broad exposure bands and avoids fine-grained cluster identities.

Table 3 reports agreement across the three systems. On occupation-level averages, Cronbach's α and the ICC both sit around 0.93, which is enough for the descriptive purposes of the paper. Kendall's W tells a less tidy story. Its mean is 0.87, but the range stretches from 0.09 to 0.99, and that spread deserves a serious look before any operational use of the index.

Table 3

Cross-model reliability statistics (occupation-level averages)

StatisticMean valueRange
Cronbach's α0.930.67–0.99
ICC(2,1)0.930.67–0.99
Kendall's W0.870.09–0.99

Most occupations sit near the top of the W distribution, where the three systems agree on both the average scores and the within-occupation rank order of tasks. The very low W values are concentrated in a small group of compliance-heavy roles whose task statements lean on words such as “inspect,” “certify,” “ensure,” and “approve.” For these occupations the systems converged on the overall exposure level but disagreed on which sign-off and licensing tasks were the most binding inside the bundle. Two implications follow. First, fine-grained rank ordering is the weakest part of the analysis, and the affected occupations should be flagged for targeted human re-rating before the index is used to inform decisions on individuals or specific workflows. Second, the high mean W cannot be read as independent validation. Where three LLMs trained on overlapping corpora interpret legalistic phrasing in similar ways, the agreement may sit inside the models' shared priors as much as inside the task content. We treat the W extremes as a flag for human review rather than as background noise.

Figure 2 reports ordered ISEF exposure scores for the 28 occupations under the multiplicative specification. Table 4 summarises the exploratory diagnostic groups, and Table 5 reports the occupation-level scores. Scores range from 8.1 to 15.7 on the 0–100 realised exposure index. That observed range is narrow on purpose: once institutional resistance is multiplied in, the safety-critical character of the whole domain compresses exposure even where technical AI suitability is high. Exact ordered scores are retained for transparency and reproducibility, but the analysis does not treat adjacent ranks as substantively distinct.

Figure 2
A bar graph showing ordered ISEF exposure scores across 28 engineering occupations.The bar graph presents ordered ISEF exposure scores for 28 engineering occupations. The x-axis lists the occupations, while the y-axis represents the ISEF exposure scores ranging from 0 to 20. The graph features five data series, each represented by different colors indicating exploratory diagnostic groupings: Diagnostic group A, Diagnostic group B, Diagnostic group C, Diagnostic group D, and Singleton diagnostic. Each bar represents the exposure score for a specific occupation, with the highest score being 15.7 for Petroleum Engineers and the lowest score being 8.1 for Mining and Geological Engineers. The bars are grouped vertically for each occupation, showing the scores for each diagnostic group. The graph highlights that the scores range from 8.1 to 15.7 on the 0 to 100 realized exposure index. The observed range is narrow due to the safety-critical character of the engineering domain, which compresses exposure even where technical AI suitability is high.

Ordered ISEF exposure scores across 28 engineering occupations. Colours indicate exploratory diagnostic groupings only, and labels report one-decimal exposure scores; adjacent ranks should not be interpreted as substantively distinct

Figure 2
A bar graph showing ordered ISEF exposure scores across 28 engineering occupations.The bar graph presents ordered ISEF exposure scores for 28 engineering occupations. The x-axis lists the occupations, while the y-axis represents the ISEF exposure scores ranging from 0 to 20. The graph features five data series, each represented by different colors indicating exploratory diagnostic groupings: Diagnostic group A, Diagnostic group B, Diagnostic group C, Diagnostic group D, and Singleton diagnostic. Each bar represents the exposure score for a specific occupation, with the highest score being 15.7 for Petroleum Engineers and the lowest score being 8.1 for Mining and Geological Engineers. The bars are grouped vertically for each occupation, showing the scores for each diagnostic group. The graph highlights that the scores range from 8.1 to 15.7 on the 0 to 100 realized exposure index. The observed range is narrow due to the safety-critical character of the engineering domain, which compresses exposure even where technical AI suitability is high.

Ordered ISEF exposure scores across 28 engineering occupations. Colours indicate exploratory diagnostic groupings only, and labels report one-decimal exposure scores; adjacent ranks should not be interpreted as substantively distinct

Close Figure 2
Table 4

Exploratory diagnostic group summary statistics (exposure)

Diagnostic groupNMeanMedianSDRepresentative occupations
Diagnostic group A48.858.850.29Geoscientists; Environmental Scientists and Specialists; Construction and Building Inspectors; Electrical and Electronic Engineering Technologists and Technicians
Diagnostic group B1111.6011.332.08Petroleum Engineers; Environmental Engineering Technologists and Technicians; Marine Engineers and Naval Architects; Wind Energy Operations Managers; Environmental Economist
Diagnostic group C812.7912.361.29Mechanical Engineering Technologists; Urban and Regional Planners; Industrial Engineers; Environmental Engineers; Water Wastewater Engineers
Diagnostic group D413.3913.041.06Surveyors; Environmental Science Teachers, Postsecondary; Construction Managers; Civil Engineering Technologists and Technicians
Singleton diagnostic112.8212.820.00Engineering Teachers, Postsecondary

Note(s): Summary statistics are calculated from the unrounded occupation-level exposure scores. Table 5 rounds occupation exposures to one decimal place, so manual recomputation from the displayed values may differ slightly in the second decimal. These are exploratory diagnostics from the k=5 solution, not validated occupational strata. The singleton diagnostic is retained only to make the instability of the solution visible

Table 5

ISEF occupation scores (multiplicative exposure)

OccupationSOCTAIRRealisedExposureDiagnostic
Petroleum Engineers17-2171.0073.278.521.515.7Diagnostic group B
Surveyors17-1022.0065.477.222.814.9Diagnostic group D
Mechanical Engineering Technologists17-3027.0060.375.324.714.9Diagnostic group C
Urban and Regional Planners19-3051.0058.675.124.914.6Diagnostic group C
Environmental Engineering Technologists and Technicians17-3025.0071.380.919.113.6Diagnostic group B
Environmental Science Teachers, Postsecondary25-1053.0063.579.220.813.2Diagnostic group D
Marine Engineers and Naval Architects17-2121.0070.281.618.412.9Diagnostic group B
Construction Managers11-9021.0069.081.318.712.9Diagnostic group D
Engineering Teachers, Postsecondary25-1032.0059.678.521.512.8Singleton diagnostic
Civil Engineering Technologists and Technicians17-3022.0066.681.218.812.5Diagnostic group D
Industrial Engineers17-2112.0056.377.822.212.5Diagnostic group C
Environmental Engineers17-2081.0055.177.322.712.5Diagnostic group C
Water Wastewater Engineers17-2051.0255.778.022.012.2Diagnostic group C
Environmental Compliance Inspectors13-1041.0156.378.321.712.2Diagnostic group C
Non-Destructive Testing Specialists17-3029.0154.577.722.312.2Diagnostic group C
Wind Energy Operations Managers11-9199.0965.581.518.512.1Diagnostic group B
Environmental Economist19-3011.0169.082.817.211.9Diagnostic group B
Emergency Management Director11-9161.0069.383.716.311.3Diagnostic group B
Civil Engineers17-2051.0062.882.018.011.3Diagnostic group B
Transportation Engineers17-2051.0168.183.616.411.2Diagnostic group B
Architecture and Engineering Management11-9041.0054.379.420.611.2Diagnostic group C
Mechanical Engineer17-2141.0069.285.914.19.7Diagnostic group B
Industrial Engineering Technologists and Technicians17-3026.0072.386.513.59.7Diagnostic group B
Geoscientists19-2042.0070.887.112.99.1Diagnostic group A
Environmental Scientists and Specialists19-2041.0064.185.814.29.1Diagnostic group A
Construction and Building Inspectors47-4011.0071.287.912.18.6Diagnostic group A
Electrical and Electronic Engineering Technologists and Technicians17-3023.0067.587.312.78.6Diagnostic group A
Mining Geological Engineers17-2151.0072.188.811.28.1Diagnostic group B

Note(s): TA, IR, Realised, and Exposure are shown to one decimal place. Exposure is calculated from the unrounded TA and IR scores using TA×(1−IR/100) before final rounding, so recomputation from the displayed one-decimal TA and IR values may differ by 0.1 in a small number of rows. The final column reports exploratory diagnostic labels only; adjacent ranks and group labels should not be read as stable occupational strata

Civil Engineers (11.3) and Transportation Engineers (11.2), for example, sit too close to treat as meaningfully different given the measurement noise documented above. The interpretive weight sits in the TA/IR decomposition and in broad bands: a small group of occupations combine high TA with especially high IR, others sit on lower IR floors, and another set is moderate on both axes. The diagnostic summaries in Table 4 and Figure 3 carry the same caveat—the singleton diagnostic group and only moderate bootstrap stability mean the labels are visual aids, not discrete categories. Most visible variation comes from differences in institutional resistance, with technical capability moving less.

Figure 3
A line graph showing occupation exposure by diagnostic group with group-level 95% confidence intervals.A line graph titled Occupation Exposure by Diagnostic Group with Group-Level 95% CI. The horizontal axis represents different occupations, while the vertical axis represents the BSF exposure values ranging from 0.0 to 20.0. The graph includes multiple diagnostic groups represented by different symbols: circles for Diagnostic group A, squares for Diagnostic group B, triangles for Diagnostic group C, diamonds for Diagnostic group D, and pentagons for Singleton diagnostics. Each data point is accompanied by error bars indicating the 95% confidence intervals. Notable trends include higher exposure values for certain occupations such as Environmental Scientist, Environmental Engineer, and Civil Engineer, with values around 15.0 to 16.0. Other occupations like Chemist, Construction & Building Inspector, and Electrical & Electronic Engineer show lower exposure values around 8.0 to 9.0.

Occupation exposure by exploratory diagnostic group with group-level 95% confidence intervals. Points and intervals are descriptive diagnostics, not validated occupational strata

Figure 3
A line graph showing occupation exposure by diagnostic group with group-level 95% confidence intervals.A line graph titled Occupation Exposure by Diagnostic Group with Group-Level 95% CI. The horizontal axis represents different occupations, while the vertical axis represents the BSF exposure values ranging from 0.0 to 20.0. The graph includes multiple diagnostic groups represented by different symbols: circles for Diagnostic group A, squares for Diagnostic group B, triangles for Diagnostic group C, diamonds for Diagnostic group D, and pentagons for Singleton diagnostics. Each data point is accompanied by error bars indicating the 95% confidence intervals. Notable trends include higher exposure values for certain occupations such as Environmental Scientist, Environmental Engineer, and Civil Engineer, with values around 15.0 to 16.0. Other occupations like Chemist, Construction & Building Inspector, and Electrical & Electronic Engineer show lower exposure values around 8.0 to 9.0.

Occupation exposure by exploratory diagnostic group with group-level 95% confidence intervals. Points and intervals are descriptive diagnostics, not validated occupational strata

Close Figure 3

Splitting TA and IR brings out a pattern that gets blurred when both are folded into one score. Among occupations with similarly high technical potential, the ones with stronger institutional resistance end up with noticeably lower exposure. The whole sample is regulated work, which keeps the gap muted, but the pattern is still readable in the numbers.

Mining & Geological Engineers and Petroleum Engineers make the point concrete. Both sit above 72 on TA. The Mining role pairs that with very high IR (88.8) and ends up at the bottom of the exposure ladder (8.1); the Petroleum role sits on lower resistance (78.5) and lands at 15.7. The comparison is not meant to suggest that one occupation is “more exposed” in any absolute sense. The point is that two technically similar roles get pulled in opposite directions by their institutional surroundings.

This is the same gap that technical-only exposure indices tend to read over. A task can look well suited to AI assistance and still sit inside a workflow where a licensed engineer's signature is the final binding step.

To see how the rankings diverge from technical-only measures, the relationships between exposure, TA, and IR are examined across the 28 occupations. Exposure correlates strongly and negatively with institutional resistance (⁠r≈−0.88⁠); its correlation with technical potential is small and slightly negative (⁠r≈−0.26⁠). These are descriptive diagnostics, not causal estimates, but they point to a clear pattern within this regulated domain: institutional resistance does most of the work in shaping the final exposure scores.

The same pattern helps explain why technical-only indices can mis-rank occupations whenever IR is both high and unevenly distributed. In civil engineering, governance and liability constraints look more consequential for near-term exposure than marginal differences in model capability. Occupation-level TA, IR, and exposure profiles are reported in Figure 4; variable-level model means are reported in Table 6.

Figure 4
A bar graph showing occupation-level profiles of technical AI suitability, institutional resistance, and realized ISEF exposure.The bar graph compares various occupations based on their technical AI suitability, institutional resistance, and realized ISEF exposure. The x-axis lists different occupations, while the y-axis measures ISEF exposure. There are multiple bars for each occupation, representing different data points. The bars are vertical and grouped. Key labels on the x-axis include occupations such as Petroleum Engineers, Surveyors, Mechanical Engineering Technicians, and others. The y-axis is labeled with ISEF exposure values ranging from 0 to 20. The graph uses a color scheme where dark blue bars represent ISEF exposure, and lighter markers represent technical AI suitability and institutional resistance. Notable trends include higher ISEF exposure for Petroleum Engineers and lower exposure for occupations like Environmental Scientists and Geological Technicians. The graph highlights the variation in exposure levels across different occupations. All values are approximated.

Occupation-level profiles of technical AI suitability (TA), institutional resistance (IR), and realised ISEF exposure

Figure 4
A bar graph showing occupation-level profiles of technical AI suitability, institutional resistance, and realized ISEF exposure.The bar graph compares various occupations based on their technical AI suitability, institutional resistance, and realized ISEF exposure. The x-axis lists different occupations, while the y-axis measures ISEF exposure. There are multiple bars for each occupation, representing different data points. The bars are vertical and grouped. Key labels on the x-axis include occupations such as Petroleum Engineers, Surveyors, Mechanical Engineering Technicians, and others. The y-axis is labeled with ISEF exposure values ranging from 0 to 20. The graph uses a color scheme where dark blue bars represent ISEF exposure, and lighter markers represent technical AI suitability and institutional resistance. Notable trends include higher ISEF exposure for Petroleum Engineers and lower exposure for occupations like Environmental Scientists and Geological Technicians. The graph highlights the variation in exposure levels across different occupations. All values are approximated.

Occupation-level profiles of technical AI suitability (TA), institutional resistance (IR), and realised ISEF exposure

Close Figure 4
Table 6

ISEF variable mean scores by model

VariableDir.ChatGPT-5DeepSeekQwen-3Consensus
Cognitive/Data AnalysisFwd6.877.398.027.38
Technological IntegrationFwd6.817.497.857.40
AI ReadinessFwd7.678.138.388.05
Automation PotentialFwd7.017.517.757.46
CollaborationFwd3.064.585.804.37
Innovation/SynergyFwd5.305.686.635.83
Domain ExpertiseFwd3.925.046.004.96
Compliance/ResponsibilityBarrier8.328.218.768.38
Industry BarriersBarrier7.457.597.537.57
Algorithm LimitationBarrier8.508.378.798.54
Data Security/PrivacyBarrier7.247.077.527.27
Sociotechnical ChallengesBarrier7.937.698.187.91

Figures 5 and 6 show technical-side and institutional-side heatmaps across the 28 occupations. The heatmaps use a consistent DTS-style blue-to-deep-blue scale and place occupation names on the bottom axis for readability. Reading the two side by side makes one feature obvious: institutional resistance is rarely spread evenly across a role. In most occupations it concentrates in a small cluster of legally sensitive tasks—permits, signoffs, safety checks, and audit-facing documentation—and those few tasks tend to drive the IR score for the whole occupation.

Figure 5
A heat map displaying technical suitability and exposure measures across various occupations.A heat map titled 'Technical Suitability and Exposure Measures by Occupation' displays data across 28 occupations. The heat map uses a blue-to-deep-blue color scale to represent values, with darker shades indicating higher values. The x-axis lists the occupations, while the y-axis includes four metrics: TSA, Additive Mean, Geometric Mean, and Multiplicative. Each cell within the grid shows a specific value corresponding to the intersection of an occupation and a metric. Notable trends include higher values for occupations like Petroleum Engineers and lower values for roles such as Mining and Geological Engineers. The heat map reveals variations in technical suitability and exposure across different occupations, highlighting areas with higher or lower measures.

Technical-side heatmap across the 28 occupations

Figure 5
A heat map displaying technical suitability and exposure measures across various occupations.A heat map titled 'Technical Suitability and Exposure Measures by Occupation' displays data across 28 occupations. The heat map uses a blue-to-deep-blue color scale to represent values, with darker shades indicating higher values. The x-axis lists the occupations, while the y-axis includes four metrics: TSA, Additive Mean, Geometric Mean, and Multiplicative. Each cell within the grid shows a specific value corresponding to the intersection of an occupation and a metric. Notable trends include higher values for occupations like Petroleum Engineers and lower values for roles such as Mining and Geological Engineers. The heat map reveals variations in technical suitability and exposure across different occupations, highlighting areas with higher or lower measures.

Technical-side heatmap across the 28 occupations

Close Figure 5
Figure 6
A bar graph showing institutional resistance and exposure measures by occupation.The bar graph compares institutional resistance, min-rule exposure, and multiplicative exposure across 28 different occupations. It features horizontal bars for each occupation, with three data lines representing institutional resistance, min-rule exposure, and multiplicative exposure. The x-axis lists the occupations, while the y-axis measures the values of institutional resistance, min-rule exposure, and multiplicative exposure. The graph uses a blue-to-deep-blue color scale. Institutional resistance values range from approximately 78.5 to 81.8, min-rule exposure values range from approximately 11.2 to 24.9, and multiplicative exposure values range from approximately 8.1 to 15.7. The graph highlights variations in institutional resistance and exposure measures across different occupations, with some occupations showing higher resistance and exposure than others. All values are approximated.

Institutional-side heatmap of institutional resistance (IR), min-rule exposure, and multiplicative exposure across the 28 occupations. The redundant realised-permission row is omitted because it is a direct linear transformation of IR

Figure 6
A bar graph showing institutional resistance and exposure measures by occupation.The bar graph compares institutional resistance, min-rule exposure, and multiplicative exposure across 28 different occupations. It features horizontal bars for each occupation, with three data lines representing institutional resistance, min-rule exposure, and multiplicative exposure. The x-axis lists the occupations, while the y-axis measures the values of institutional resistance, min-rule exposure, and multiplicative exposure. The graph uses a blue-to-deep-blue color scale. Institutional resistance values range from approximately 78.5 to 81.8, min-rule exposure values range from approximately 11.2 to 24.9, and multiplicative exposure values range from approximately 8.1 to 15.7. The graph highlights variations in institutional resistance and exposure measures across different occupations, with some occupations showing higher resistance and exposure than others. All values are approximated.

Institutional-side heatmap of institutional resistance (IR), min-rule exposure, and multiplicative exposure across the 28 occupations. The redundant realised-permission row is omitted because it is a direct linear transformation of IR

Close Figure 6

The comparison also has implications for how exposure is interpreted. Where institutional resistance is concentrated in a limited set of accountability-related tasks, changes in governance arrangements and signing responsibilities may have a greater effect on occupational exposure than incremental improvements in the underlying AI tools. This suggests that exposure cannot be assessed from technical capability alone, particularly for tasks involving formal approval, safety assurance, and regulatory accountability.

Validation checks

Several checks were run on the scoring protocol, the functional form, and the plausibility of the IR rubric. These checks are deliberately separated from validation claims. The practitioner survey helps calibrate the TA/IR split across 15 concrete task vignettes, the FAR-clause mapping illustrates why institutional resistance belongs in the construct, and the benchmarking exercise checks the technical component against prior exposure measures. None of these exercises is treated as a substitute for a same-task expert panel.

Functional Form Justification. The multiplicative specification (Equation 4) treats institutional resistance as a gate, not a marginal offset. In high-stakes engineering work, a licensed engineer's signature requirement can suppress the practical exposure of a design task that otherwise looks highly suitable for AI assistance. Exposure therefore falls towards zero as resistance approaches its upper bound. The logic is theoretical, not statistically identified, which is why alternative specifications are tested below.

Sensitivity Analysis. Occupational scores were compared across four functional forms: additive mean exposure (⁠(TA+[100−IR])/2⁠), multiplicative exposure (⁠TA×[1−IR/100]⁠), a min-rule specification (⁠min(TA,100−IR)⁠), and an equal-weight geometric mean of technical potential and realised institutional permission. Table 7 and Figure 7 summarise the diagnostics. The additive mean produces a noticeably different fine-grained ordering from the multiplicative model (Spearman ρ=0.44⁠), while the min-rule is more closely aligned but still changes several ranks (Spearman ρ=0.80⁠). The equal-weight geometric mean is rank-equivalent to the multiplicative model (Spearman ρ=1.00⁠), because both are monotonic functions of TA×(100−IR)⁠. These results strengthen the main caution of the paper: broad exposure bands are more defensible than exact occupational ranks. The occupations that move most across specifications include Industrial Engineering Technologists and Technicians, Water Wastewater Engineers, Environmental Engineers, and Non-Destructive Testing Specialists. Weighted and unweighted O*NET aggregation produce nearly identical multiplicative rankings (Spearman ρ=0.998⁠).

Table 7

Functional-form sensitivity checks

SpecificationRank correlationInterpretation
MultiplicativePrimary modelTA×(1-IR/100); treats institutional resistance as a partial gate on technical suitability
Additive meanSpearman ρ = 0.44Equivalent in ranking to TA+(100-IR) after rescaling. It gives less force to high institutional resistance than the multiplicative model
Min-ruleSpearman ρ = 0.80Uses the lower of TA and realised permission. It compresses high-resistance occupations and supports broad-band rather than fine-rank interpretation
Geometric meanSpearman ρ = 1.00Equal-weight geometric mean is rank-equivalent to the multiplicative model because both are monotonic functions of TA×(100-IR)
O*NET weighting checkSpearman ρ = 0.998Weighted and unweighted aggregation produce nearly identical rankings

Note(s): Additive mean, min-rule, and geometric mean are computed from the same occupation-level TA and IR scores as the primary multiplicative model. Full occupation-level scores are provided in the replication repository

Figure 7
A line graph showing functional-form sensitivity across different occupation-level exposure specifications.A line graph titled 'Functional-Form Sensitivity Across Occupations' displays four different functional forms: Multiplicative ISEF, Additive mean, Min-rule, and Geometric mean. The horizontal axis represents various occupations, while the vertical axis shows exposure scores under each specification. The Multiplicative ISEF line, represented by a solid line, shows a decreasing trend. The Additive mean line, represented by a dashed line, fluctuates but generally stays above the Multiplicative ISEF line. The Min-rule line, represented by a dotted line, shows a relatively stable trend. The Geometric mean line, represented by a dash-dotted line, follows a similar pattern to the Multiplicative ISEF line. Notable occupations with significant movement across specifications include Industrial Engineering Technologists and Technicians, Water Wastewater Engineers, Environmental Engineers, and Non-Destructive Testing Specialists.

Functional-form sensitivity across multiplicative, additive, min-rule, and equal-weight geometric-mean occupation-level exposure specifications

Figure 7
A line graph showing functional-form sensitivity across different occupation-level exposure specifications.A line graph titled 'Functional-Form Sensitivity Across Occupations' displays four different functional forms: Multiplicative ISEF, Additive mean, Min-rule, and Geometric mean. The horizontal axis represents various occupations, while the vertical axis shows exposure scores under each specification. The Multiplicative ISEF line, represented by a solid line, shows a decreasing trend. The Additive mean line, represented by a dashed line, fluctuates but generally stays above the Multiplicative ISEF line. The Min-rule line, represented by a dotted line, shows a relatively stable trend. The Geometric mean line, represented by a dash-dotted line, follows a similar pattern to the Multiplicative ISEF line. Notable occupations with significant movement across specifications include Industrial Engineering Technologists and Technicians, Water Wastewater Engineers, Environmental Engineers, and Non-Destructive Testing Specialists.

Functional-form sensitivity across multiplicative, additive, min-rule, and equal-weight geometric-mean occupation-level exposure specifications

Close Figure 7

Illustrative Anchoring via Contract Clauses. Supplementary Material D maps three representative US federal procurement workflows to the institutional-resistance rubric. Engineering work in these settings is governed by inspection, audit, records, and acceptance clauses such as FAR 52.246-4 (Inspection of Services) and FAR 52.215-2 (Audit and Records). Three contracts cannot validate the index. They are useful for showing why liability, auditability, and review requirements belong inside the IR construct in the first place.

Cross-Index Benchmarking. TA scores were benchmarked against established technical exposure indices (Eloundou et al., 2023; Felten et al., 2021). TA here means technical suitability for generative AI assistance, with no claim of full task automation. The correlation with prior “Exposure” scores is strong (⁠r≈0.72⁠), so the technical side of ISEF tracks existing capability measures closely. Final ISEF exposure diverges from those rankings once IR is applied; the divergence is part of the contribution, since governance acts as a gate on what counts as workable exposure. Table 8 summarises the comparison.

Table 8

Cross-index benchmarking of the technical AI suitability component

MeasureBenchmark evidenceInterpretation
Technical AI Suitability (TA)Pearson correlation with existing technical AI exposure indices: r ≈ 0.72The TA component aligns with prior technical-only exposure measures. This supports the technical side of the rubric without claiming realised adoption
ISEF ExposureDescriptive divergence from technical-only rankingsDivergence is expected because ISEF applies institutional resistance as a gating factor. The difference is part of the sociotechnical argument, not a failure of the TA measure

Note(s): This benchmark is used as a concurrent-validity diagnostic for technical AI suitability. It does not validate realised adoption. Final ISEF exposure intentionally diverges from technical-only indices once institutional resistance is included

Preliminary Practitioner Calibration. A short practitioner survey was fielded during the revision period as preliminary calibration, not validation, because a fair criticism of the main scoring exercise is that LLMs are being asked to grade work LLMs may one day do. The questionnaire in question.png presented 15 concrete civil, environmental, and adjacent engineering task vignettes, including drilling-plan design, environmental permit preparation, site-plan compliance review, legal boundary descriptions, notices of violation, FEA stress analysis, and building-inspection sign-off. For each vignette, respondents scored two independent 0–10 dimensions. D1 asked how much generative AI could technically assist the task assuming no legal, liability, or sign-off restrictions. D2 asked how much actual sign-off requirements, professional liability, regulatory compliance, and audit traceability limit the use of AI-generated outputs.

The raw Google Forms export contained 11 anonymous records. Two were removed because they did not contain the full set of D1/D2 task ratings and professional-judgement confirmations. The cleaned file therefore retains 9 complete responses across all 15 vignettes. The retained panel includes academic/researcher, mid-level engineer, senior engineer, and government/regulator roles; 6 of the 9 respondents reported holding a PE licence. Table 9 reports the vignette-level means after these exclusions. This preliminary calibration does not validate the 674-task scoring item by item, the occupation-level ISEF index, or the constructed calibration-only adjusted score; no expert–LLM task-level correlation is reported. What it checks is narrower: whether practitioners draw a similar broad distinction between technical usefulness and institutional permissibility when judging realistic engineering tasks.

Table 9

Preliminary practitioner calibration survey after removing incomplete responses

TaskSurvey vignetteNAI helpBarriersCalibration-only adjusted score
T01Petroleum Engineer – Design drilling plans95.25.32.5
T02Petroleum Engineer – Analyze geological data to determine drilling locations96.04.93.1
T03Environmental Engineer – Prepare environmental permit applications97.24.63.7
T04Environmental Engineer – Inspect facilities for environmental compliance94.28.01.1
T05Mining Engineer – Prepare technical reports for regulatory agencies96.05.42.6
T06Geoscientist – Analyze soil, rock, and groundwater samples95.25.22.9
T07Civil Engineer – Review site plans for building code compliance95.07.21.0
T08Surveyor – Prepare legal property boundary descriptions92.77.70.9
T09Construction Manager – Review and approve contractor change orders96.35.82.9
T10Water/Wastewater Engineer – Design water treatment processes95.35.82.5
T11Environmental Inspector – Issue notices of violation92.68.30.4
T12NDT Specialist – Interpret radiographic images to identify defects97.04.73.8
T13Engineering Teacher – Develop course syllabi and materials98.72.76.4
T14Mechanical Engineer – Perform stress analysis using FEA software94.66.71.7
T15Construction Inspector – Sign off on completed building inspections91.48.30.2

Note(s): Values are means on a 0–10 scale from 9 complete responses after excluding 2 incomplete records from the raw Google Forms CSV. The 15 vignettes and dimension wording are taken from question.pdf. D1 asks how much generative AI can technically assist the task assuming no legal, liability, or sign-off restrictions. D2 asks how much actual sign-off, liability, compliance, and audit-traceability requirements limit use of AI-generated outputs. A response was retained only when all 15 D1/D2 paired ratings were present and each task-level professional-judgement confirmation was marked yes. The calibration-only adjusted score is computed at respondent-task level as AI Help ×(1−Barriers/10) and then averaged. It is a descriptive calibration summary, not a validated exposure measure

On that limited question, the pattern is consistent with the framework. Technical assistance is rated highest for course-material development (Engineering Teacher, mean D1 = 8.7), permit application preparation (Environmental Engineer, D1 = 7.2), and radiographic-image interpretation (NDT Specialist, D1 = 7.0). Institutional constraints are rated highest for notices of violation (Environmental Inspector, mean D2 = 8.3), completed building-inspection sign-off (Construction Inspector, D2 = 8.3), facility compliance inspection (Environmental Engineer, D2 = 8.0), and legal boundary descriptions (Surveyor, D2 = 7.7). The constructed calibration-only adjusted score (AI Help ×(1−Barriers/10)⁠) is highest for course-material development (6.4) and lowest for building-inspection sign-off (0.2) and notices of violation (0.4); it is a descriptive summary of these nine responses, not a validated exposure measure. Across the 15 task means, AI assistance potential and institutional constraints are strongly negatively associated (⁠r=−0.93⁠; p<0.001⁠), while AI assistance potential and the calibration-only adjusted score are strongly positively associated (⁠r=0.93⁠; p<0.001⁠). The survey therefore supports the core TA/IR distinction, while still falling short of the larger same-task expert panel needed for full validation.

Reliability and Dispersion. Within-system variation averaged 0.73 on a 0–10 scale, and 89% of task-variable pairs showed low dispersion. Even so, a small set of occupations showed weaker cross-model rank agreement, which is why the cluster results are treated as exploratory. The silhouette analysis in Figure 8 supports the use of groupings for description, but not as stable occupational classes.

Figure 8
A line graph titled Mean Silhouette Diagnostic showing the mean silhouette coefficient on the y axis and the number of groups on the x axis.The line graph titled Mean Silhouette Diagnostic presents the mean silhouette coefficient on the y axis, ranging from 0.0 to 0.7, and the number of groups on the x axis, ranging from 2 to 10. The graph shows data points for each number of groups, connected by a line. The mean silhouette coefficient starts at approximately 0.45 for 2 groups, increases to around 0.5 for 3 groups, and remains relatively stable around 0.5 for 4 and 5 groups. It then decreases to about 0.45 for 6 groups, drops further to approximately 0.4 for 7 groups, rises again to around 0.55 for 8 groups, and finally decreases to about 0.45 for 9 and 10 groups. All values are approximated.

Silhouette analysis across candidate K-means solutions. The reported groupings are exploratory diagnostics, not validated occupational strata

Figure 8
A line graph titled Mean Silhouette Diagnostic showing the mean silhouette coefficient on the y axis and the number of groups on the x axis.The line graph titled Mean Silhouette Diagnostic presents the mean silhouette coefficient on the y axis, ranging from 0.0 to 0.7, and the number of groups on the x axis, ranging from 2 to 10. The graph shows data points for each number of groups, connected by a line. The mean silhouette coefficient starts at approximately 0.45 for 2 groups, increases to around 0.5 for 3 groups, and remains relatively stable around 0.5 for 4 and 5 groups. It then decreases to about 0.45 for 6 groups, drops further to approximately 0.4 for 7 groups, rises again to around 0.55 for 8 groups, and finally decreases to about 0.45 for 9 and 10 groups. All values are approximated.

Silhouette analysis across candidate K-means solutions. The reported groupings are exploratory diagnostics, not validated occupational strata

Close Figure 8

Civil and environmental engineering is governed at least as much by review chains as by technical skill. Licensure, liability, and regulated sign-off decide what counts as acceptable evidence, who is allowed to issue it, and who carries the risk when something goes wrong. ISEF starts from that working reality and keeps it inside the measurement instead of treating it as background.

Several occupations in the sample combine high technical potential with substantial institutional resistance. The resulting exposure scores stay inside a moderate band, so “high exposure” here is a relative position within a regulated profession; it is not a prediction of displacement. The plausible entry points for generative AI in this domain sit at the edges of the workflow—drafting, checking, summarising, and supporting documentation—where outputs can be folded into review routines that already exist.

ISEF does not claim to capture every driver of adoption. Firm strategy, procurement culture, informal professional norms, and individual project leadership all matter and sit largely outside the index. What the framework adds is a way of keeping governance conditions in the same analytical frame as technical suitability, which task-only exposure measures struggle to do.

The roles that look most open to AI assistance—structured design calculations, simulation support, technical documentation—are also the ones tightly coupled to public safety and infrastructure reliability. The same coupling keeps institutional resistance high in precisely the places where technical scores are high.

There is a forward-looking edge to this observation. As generative models become more capable, the consequences of an undetected error scale with them. If AI-assisted outputs feed into safety-critical work, oversight requirements may tighten before they relax. Institutional resistance can therefore be a durable feature of the profession, not just a temporary lag behind the technology.

The picture is also less uniform than the occupation-level numbers suggest. Project governance, signing authority, and how individual firms allocate work between licensed and unlicensed staff almost certainly matter below the occupation level. Differences across licensure and liability regimes will also pull the same occupation in different directions across jurisdictions.

ISEF is designed around a general regulated-work principle: technical capability becomes practically relevant only after it passes domain-specific accountability, documentation, privacy, and authority requirements. That principle may be useful in professions such as healthcare, law, accountancy, architecture, and regulated utilities, where AI-assisted outputs may be technically plausible but cannot be adopted without professional oversight. Transfer is not automatic. Each domain would need its own task universe, regulatory map, IR indicators, weights, and practitioner calibration; for example, clinical governance and patient confidentiality are not interchangeable with engineering sign-off and infrastructure liability. ISEF should therefore be treated as a portable measurement architecture rather than a ready-made cross-profession score.

Both sides of the framework are time dependent. Technical suitability can change as models gain tool use, multimodal reasoning, reliability controls, or domain-specific evaluation evidence. Institutional resistance can also change as regulators issue guidance, insurers revise coverage, clients amend procurement requirements, and professional bodies clarify acceptable review and sign-off practices. A future longitudinal implementation should version the TA and IR rubrics, retain dated regulatory evidence, re-score a fixed task panel at planned intervals, and reconvene a larger practitioner panel when material changes occur. Such updates would allow observed changes in exposure to be separated from changes in the measurement rule itself.

For workforce planning, exposure is most defensible as a descriptive baseline; it should not be read as a direct estimate of job-loss risk. In regulated roles, AI use is likely to settle first on documentation and analytical support, where outputs can be slotted into audit trails and existing oversight routines.

For governance, the framework surfaces a tension between AI capability and existing liability arrangements that will not dissolve quickly. The concrete levers sit exactly where that tension lives: AI-specific clauses in engineering procurement contracts, professional indemnity guidance for AI-assisted deliverables, audit-trail requirements for generated outputs, and clearer rules on whether a professional engineer's sign-off can rely on AI-supported calculations or documentation. These are not marginal adjustments to an exposure score; they decide which parts of a workflow are institutionally permissible in the first place.

For training and professional development, the heterogeneous exposure profiles imply differentiated needs. Some roles benefit from stronger emphasis on integrating and supervising generative tools; others need more weight on competencies tied to accountability, public safety, and document defensibility.

A few limitations deserve to be named openly rather than left to the framing. LLM scoring showed strong agreement across the three systems, and that agreement may partly reflect shared training priors, especially in how legalistic phrasing is read. The practitioner survey adds an external check, but it is not the same as asking engineers to re-score the same O*NET tasks. Nine complete respondents provide preliminary calibration only; they cannot support population-level claims, task-level validation, or validation of the constructed calibration-only adjusted score. A larger, purposively stratified expert panel should re-score the same O*NET tasks sampled from high-, medium-, and low-exposure occupations, report inter-rater agreement and uncertainty, and test whether calibration patterns replicate across professional roles and jurisdictions. This design would allow direct expert–LLM correlations to be estimated.

Institutional resistance is also, openly, a constructed measure. Different rubric items, weightings, or functional forms would shift the resulting exposure profile. The current reading is a cross-section of present-day conditions; it does not track how procurement standards, professional indemnity, or sectoral regulation may move under generative AI. Following exposure through time, and linking it to observed adoption inside real projects, is work for the next study.

ISEF was built to look at generative AI exposure in civil and environmental engineering through two lenses at once: what a task is technically suited to, and what institutional rules will let through. The aim is descriptive. The framework reports where suitability and governance constraints meet; it does not try to time adoption or forecast job loss.

Applied to 674 tasks across 28 occupations, the analysis surfaces a pattern that is easy to miss when only the technical side is measured. Many roles read as open to AI assistance, yet the exposure score stays inside a moderate band once liability, compliance, and sign-off requirements are added to the picture. Indices that leave the institutional side outside the measurement will tend to misread how generative AI is likely to enter regulated engineering work in the near term.

The more useful question that follows is set at the workflow level, not the occupation level. Which parts of a typical project become institutionally permissible for AI support, under what review arrangements, and with what liability conditions? The same question can be tested in other regulated professions only after their institutional conditions have been specified and calibrated. The released data, scoring rubric, and code are intended as a starting point for that kind of follow-up: larger same-task expert panels, adoption-linked case studies, longitudinal rubric updates, and project-level tracing of where review and responsibility actually shift.

The main analysis uses publicly available occupational task data. The revision also reports an anonymous practitioner calibration survey that collected professional judgements without personal identifiers or sensitive personal data. Ethical approval was not required under the author's institutional understanding for this minimal-risk anonymous consultation.

I sincerely thank my supervisor, Dr Stephen Suryasentana, for his guidance and support throughout the research. I am also grateful to all participants who contributed to the calibration. Finally, I thank the reviewers for their detailed and constructive feedback across several rounds of view.

The supplementary material for this article can be found online:

Abbott
,
A.
(
1988
).
The system of professions: An essay on the division of expert labor
.
University of Chicago Press
.
Acemoglu
,
D.
, &
Restrepo
,
P.
(
2022
).
Tasks, automation, and the rise in US wage inequality
.
Econometrica
,
90
(
5
),
1973
–
2016
. doi: .
Arntz
,
M.
,
Gregory
,
T.
, &
Zierahn
,
U.
(
2017
).
Revisiting the risk of automation
.
Economics Letters
,
159
,
157
–
160
. doi: .
Autor
,
D. H.
(
2015
).
Why are there still so many jobs? The history and future of workplace automation
.
Journal of Economic Perspectives
,
29
(
3
),
3
–
30
. doi: .
Azimi
,
M.
,
Eslamlou
,
A. D.
, &
Pekcan
,
G.
(
2020
).
Data-driven structural health monitoring and damage detection through deep learning: State-of-the-art review
.
Sensors
,
20
(
10
),
2778
. doi: .
Bock
,
T.
(
2015
).
The future of construction automation: Technological disruption and the upcoming ubiquity of robotics
.
Automation in Construction
,
59
,
113
–
121
. doi: .
Bommasani
,
R.
,
Hudson
,
D. A.
,
Adeli
,
E.
,
Altman
,
R.
,
Arora
,
S.
,
von Arx
,
S.
, ... and
Liang
,
P.
(
2021
).
On the opportunities and risks of foundation models
. arXiv:.
Brynjolfsson
,
E.
,
Mitchell
,
T.
, &
Rock
,
D.
(
2018
).
What can machines learn?
. In
AEA Papers and Proceedings
(Vol. 
108
, pp. 
43
–
47
).
Bubeck
,
S.
,
Chadrasekaran
,
V.
,
Eldan
,
R.
,
Gehrke
,
J.
,
Horvitz
,
E.
,
Kamar
,
E.
, ... and
Zhang
,
Y.
(
2023
).
Sparks of artificial general intelligence: Early experiments with GPT-4
. arXiv:.
Calo
,
R.
(
2017
).
Artificial intelligence policy: A primer and roadmap
.
UC Davis Law Review
,
51
(
2
),
399
–
435
.
Dell’Acqua
,
F.
,
McFowland III
,
E.
,
Mollick
,
E.
,
Lifshitz-Assaf
,
H.
,
Kellogg
,
K. C.
,
Rajendran
,
S.
, …
Lakhani
,
K. R.
(
2023
).
Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality
.
Harvard Business School Working Paper
,
24
-
013
.
Dwivedi
,
Y. K.
,
Kshetri
,
N.
,
Hughes
,
L.
,
Slade
,
E. L.
,
Jeyaraj
,
A.
,
Kar
,
A. K.
, …
Wright
,
R.
(
2023
).
‘So what if ChatGPT wrote it?’ Multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy
.
International Journal of Information Management
,
71
, 102642. doi: .
Eloundou
,
T.
,
Manning
,
S.
,
Mishkin
,
P.
, &
Rock
,
D.
(
2023
).
GPTs are GPTs: An early look at the labor market impact potential of large language models
. arXiv:.
Felten
,
E.
,
Raj
,
M.
, &
Seamans
,
R.
(
2021
).
Occupational exposure to artificial intelligence
.
Strategic Management Journal
,
42
(
12
),
2195
–
2217
.
Floridi
,
L.
, &
Cowls
,
J.
(
2019
).
A unified framework of five principles for AI in society
.
Harvard Data Science Review
,
1
(
1
). doi: .
Freidson
,
E.
(
2001
).
Professionalism: The third logic
.
University of Chicago Press
.
Frey
,
C. B.
, &
Osborne
,
M. A.
(
2017
).
The future of employment: How susceptible are jobs to computerisation?
.
Technological Forecasting and Social Change
,
114
,
254
–
280
. doi: .
Goos
,
M.
,
Manning
,
A.
, &
Salomons
,
A.
(
2014
).
Explaining job polarization
.
American Economic Review
,
104
(
8
),
2509
–
2526
.
Jasanoff
,
S.
(
2004
).
States of knowledge: The Co-production of science and social order
.
Routledge
.
Jasanoff
,
S.
(
2016
).
The ethics of invention: Technology and the human future
.
W. W. Norton
.
Kellogg
,
K. C.
,
Valentine
,
M. A.
, &
Christin
,
A.
(
2020
).
Algorithms at work: The new contested terrain of control
.
Academy of Management Annals
,
14
(
1
),
366
–
410
. doi: .
Leonardi
,
P. M.
(
2012
). Materiality, sociomateriality, and socio-technical systems: What do these terms mean? How are they related? Do we need them?. In
P. M.
 
Leonardi
,
B. A.
 
Nardi
, &
J.
 
Kallinikos
(Eds.),
Materiality and Organizing: Social Interaction in a Technological World
(pp. 
25
–
48
).
Oxford University Press
.
Mohaghegh
,
S. D.
(
2017
).
Data-driven reservoir modeling
.
Society of Petroleum Engineers
.
Noy
,
S.
, &
Zhang
,
W.
(
2023
).
Experimental evidence on the productivity effects of generative artificial intelligence
.
Science
,
381
(
6654
),
187
–
192
. doi: .
OpenAI
(
2023
).
GPT-4 technical report
. arXiv:.
Peterson
,
N. G.
,
Mumford
,
M. D.
,
Borman
,
W. C.
,
Jeanneret
,
P. R.
,
Fleishman
,
E. A.
,
Levin
,
K. Y.
, ... and
Dye
,
D. M.
(
2001
).
Understanding work using the occupational information network (O*NET): Implications for practice and research
.
Personnel Psychology
,
54
(
2
),
451
–
492
. doi:.
Raj
,
M.
, &
Seamans
,
R.
(
2019
). Artificial intelligence, labor, productivity, and the need for firm-level data. In
A.
 
Agrawal
,
J.
 
Gans
, &
A.
 
Goldfarb
(Eds.),
The Economics of Artificial Intelligence: An Agenda
(pp. 
553
–
566
).
University of Chicago Press
.
Raji
,
I. D.
,
Smart
,
A.
,
White
,
R. N.
,
Mitchell
,
M.
,
Gebru
,
T.
,
Hutchinson
,
B.
, …
Barnes
,
P.
(
2020
).
Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing
. In
Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency
(pp. 
33
–
44
).
ACM
.
Sacks
,
R.
,
Eastman
,
C.
,
Lee
,
G.
, &
Teicholz
,
P.
(
2018
).
BIM handbook: A guide to building information modeling for owners, designers, engineers, contractors, and facility managers
( (3rd edition) ).
Wiley
.
Susskind
,
R.
, &
Susskind
,
D.
(
2015
).
The future of the professions: How technology will transform the work of human experts
.
Oxford University Press
.
Tibshirani
,
R.
,
Walther
,
G.
, &
Hastie
,
T.
(
2001
).
Estimating the number of clusters via the gap statistic
.
JRSS-B
,
63
(
2
),
411
–
423
. doi: .
Tolan
,
S.
,
Pesole
,
A.
,
Martinez-Plumed
,
F.
,
Fernandez-Macias
,
E.
,
Hernandez-Orallo
,
J.
, &
Gomez
,
E.
(
2021
).
Measuring the occupational impact of AI: Tasks, cognitive abilities and AI benchmarks
.
Journal of Artificial Intelligence Research
,
71
,
191
–
236
. doi: .
Veale
,
M.
, &
Borgesius
,
F. Z.
(
2021
).
Demystifying the draft EU AI Act
.
Computer Law Review International
,
22
(
4
),
97
–
112
.
Webb
,
M.
(
2020
).
The impact of artificial intelligence on the labor market
.
SSRN Working Paper No. 3482150
.
Weidinger
,
L.
,
Uesato
,
J.
,
Rauh
,
M.
,
Griffin
,
C.
,
Huang
,
P.-S.
,
Mellor
,
J.
, …
Gabriel
,
I.
(
2022
).
Taxonomy of risks posed by language models
. In
Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency
(pp. 
214
–
229
).
ACM
.
Published in Digital Transformation and Society. Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at Link to the terms of the CC BY 4.0 licence.

Supplementary data

or Create an Account

Close subscription notice
Close access options