Executive takeaway
Supervised support, accountable decisions
Generative AI should be treated as supervised decision support in safety-critical work. Its outputs can reinforce user assumptions, fabricate technical information, vary across prompts, and encourage over-reliance. Qualified professionals must retain authority for verification, approval, escalation, and final decisions.
Why It Matters
Fluent, confident language can make a weak or incorrect answer feel settled before a safety professional has challenged the evidence. In hazard identification, risk assessment, incident investigation, procedure development, and control selection, that influence can redirect attention or normalize an unsupported premise.
In safety-critical work, even an isolated technical error can contribute to high-consequence decisions. Trust therefore has to be calibrated through traceable sources, transparent uncertainty, scenario testing, and accountable human review. Effective adoption controls must address model failure and the human-performance risks of automation bias, confirmation bias, and cognitive offloading.
What Was Examined
The study used simulated, scenario-based evaluation of generative AI for workplace health and safety, not live industrial deployment. Across six safety-related task categories, four subject-matter experts reviewed 124 AI responses for factual accuracy, completeness, regulatory alignment, explainability, required human correction, and trustworthiness. Recommended controls were also mapped to the hierarchy of controls.
Key Findings
Sycophancy appeared in 53 of 124 responses, or 43%.
Hallucinations appeared in 36 of 124 responses, or 29%.
Eight hallucinated cases were judged to have high or critical risk potential.
Approximately 22% of scenario sessions involved acceptance of AI-generated content without further scrutiny or modification.
Only 14% of responses provided citation or justification for recommended controls.
Performance was materially weaker in complex and specialized scenarios than in routine safety questions, with reported accuracy falling from approximately 88% in common tasks to 54% in more complex tasks.
Evidence boundary: These findings are tied to the study design, tested scenarios, platform conditions, prompts, and expert-assessment framework. They should not be treated as universal performance estimates for every model, task, industry, or deployment.
What This Means in Practice
For Engineers
- Treat AI-generated hazard and control outputs as hypotheses, not verified engineering conclusions.
- Verify outputs independently against drawings, specifications, standards, operating conditions, and physical constraints.
- Test performance across edge cases, degraded conditions, and prompt variations before approving a use case.
- Maintain version-controlled prompts, evaluation cases, and traceability from source evidence to analysis and approved decisions.
- Do not permit generative AI to approve safety-critical designs or controls autonomously.
For Practitioners
- Require source checking, structured human review, explicit challenge prompts, and verification checklists.
- Use review gates for incident investigation, risk assessment, procedure development, and control selection.
- Protect against confirmation bias by requiring reviewers to identify contradictory evidence, uncertainty, and unresolved issues.
- Retain manual competence and escalate outputs that are unsupported, inconsistent, or outside the approved use case.
For Leaders
- Assign governance ownership and define approved, restricted, and prohibited uses.
- Set explicit accountability, sign-off, auditability, monitoring, and incident-reporting requirements.
- Fund workforce AI literacy, validation resources, vendor scrutiny, and preservation of independent professional capability.
- Avoid productivity targets or incentives that encourage unchecked reliance on AI-generated content.
Recommended Organizational Actions
- 01
Classify AI use cases by safety criticality.
- 02
Define approved, restricted, and prohibited uses.
- 03
Require authoritative source verification for critical claims.
- 04
Establish qualified human review and approval.
- 05
Test for sycophancy, hallucination, inconsistency, and omission.
- 06
Record the model, version, prompt, sources, reviewer, corrections, and final decision.
- 07
Monitor errors, drift, overrides, and user reliance.
- 08
Preserve independent professional competence through training and periodic AI-free exercises.
These recommendations are decision-support guidance, not legal advice or a universal compliance standard. Applicable requirements must be established for the organization, jurisdiction, technology, and use context.
Evaluation Framework
-
01
Define the safety use case
-
02
Prepare realistic evaluation scenarios
-
03
Capture AI responses without modification
-
04
Conduct qualified expert review
-
05
Score against performance and safety criteria
-
06
Map recommendations to the hierarchy of controls
-
07
Refine the use case, restrictions, and safeguards
The evaluation sequence begins by defining a bounded safety use case and realistic scenarios. Unmodified AI responses are then reviewed by qualified experts, scored against performance and safety criteria, and compared with the hierarchy of controls. The results inform whether the use case should proceed, be restricted, receive additional safeguards, or be rejected.
Methods and Limitations
- The work used simulated scenarios rather than live industrial deployment.
- Performance was evaluated by experts rather than measured through long-term operational outcomes.
- Findings reflect the tested model, prompts, task types, and evaluation design.
- Transferability to other AI systems and industries requires validation.
- Long-term effects on professional judgment and organizational performance require further study.
- Field-based and comparative studies are needed.
Funding, Conflicts, and Institutional Independence
Funding
Not applicable, as reported in the original publication.
Conflicts of interest
The original publication reports no conflicts of interest.
References to employers, universities, laboratories, publishers, professional organizations, credentials, and institutional roles are included solely for factual biographical and publication identification. They do not imply affiliation, sponsorship, endorsement, partnership, authorization, or responsibility for Risk Analytics AI or its services. RAAI content is prepared and presented in an independent professional capacity.