How to Review an AI-Generated COSHH Assessment

Safe Foundry Team16 Sep 20267 min read
How to Review an AI-Generated COSHH Assessment
Key takeaways
  • An AI-generated COSHH assessment is a draft until a competent person verifies the exact source documents and workplace facts.
  • Reviewers should separate SDS facts from inferred task details and reject unsupported quantities, exposure ratings, controls or citations.
  • The review must add local context about method, duration, people, ventilation, maintenance and emergency arrangements.
  • Approval should preserve source, version, reviewer, changes, actions and sign-off rather than accepting a polished output at face value.

Review an AI-generated COSHH assessment as a structured draft, not a safety decision. Verify every chemical fact against the exact SDS, then compare every task and control statement with the workplace. Reject invented details, add missing exposure context, confirm control performance and require competent approval before use.

What should you verify before reading the risk rating?

Confirm the source and product identity first. Check product name, supplier, code, concentration, physical form, SDS version and revision date. If the AI used the wrong document, detailed review of its conclusions is wasted.

HSE states that an SDS helps employers assess risk but is not itself the assessment (HSE SDS guidance). AI cannot fill that gap without accurate workplace input.

How do you check source fidelity?

Trace each factual claim to a visible source section. Review:

Draft contentPrimary check
Product and supplierSDS Section 1 and label
Classification and H-statementsSDS Section 2
CompositionSDS Section 3, respecting disclosure limits
First aid and spill responseSDS Sections 4 and 6
Handling and storageSDS Section 7
Exposure limits and PPE adviceSDS Section 8
Flash point and physical propertiesSDS Section 9
Reactivity and incompatibilitySDS Section 10

Flag a claim when it is absent, contradicted or stronger than the source. A model may turn “no data available” into a reassuring conclusion or combine facts from different products.

The SDS-to-assessment workflow should retain the exact source so reviewers can trace the draft.

Which workplace details is AI most likely to miss?

AI cannot observe local work unless those facts are deliberately supplied and verified. Check:

  • quantity and concentration used;
  • frequency and duration;
  • pouring, spraying, heating, mixing or machining method;
  • normal and non-routine releases;
  • room, enclosure and ventilation;
  • operators, nearby people, cleaners and contractors;
  • existing engineering controls and maintenance;
  • storage, waste and emergency arrangements;
  • changes, incidents and worker feedback.

HSE requires the assessment to account for the work and work practices, not only substance hazards (HSE COSHH assessment guidance).

How should exposure routes and people be reviewed?

Check inhalation, skin, eye and ingestion routes for every task stage. AI may copy a route from classification and miss exposure created by cleaning or maintenance. It may also list only the operator.

Observe the task where practical. Include workers nearby, maintenance, cleaners, waste handlers, contractors and emergency responders. Handle individual health considerations through appropriate confidential processes.

How do you review the controls?

Test whether each control addresses a specific source and can work in practice. Reject vague controls such as “use ventilation” or “wear PPE”. Require type, location, operating standard, user checks, maintenance and failure response.

Use the control order: eliminate, substitute, reduce, enclose, capture, organise work and then add PPE for residual exposure. HSE warns not to default to PPE because it is less reliable than other measures (HSE good control practice).

Ask whether the control exists now or is merely recommended. An assessment must not describe a proposed extractor as an existing safeguard.

How should risk scores be challenged?

A neat score does not compensate for uncertain inputs. Check the method, scale and definitions. Verify that severity comes from supported health information and likelihood reflects actual exposure and controls.

Reject scores derived from hazard pictograms alone, arbitrary multiplication or assumed control effectiveness. Keep the narrative evidence so another reviewer can understand the decision even without the number.

What hallucinations should reviewers expect?

Look for plausible but unsupported specificity. Examples include:

  • invented workplace exposure limits;
  • wrong glove materials or breakthrough times;
  • fabricated legal citations;
  • assumed LEV flow rates;
  • universal review periods;
  • made-up first-aid instructions;
  • claims that a product is safe because no data were supplied;
  • controls copied from another substance or jurisdiction.

Verify current legal and exposure claims through HSE or legislation.gov.uk. Do not ask the same AI system to certify its own answer without checking sources.

What should competent approval record?

Approval should show what the reviewer checked and changed. Record source versions, workplace evidence, observations, reviewer competence, amendments, unresolved actions, interim controls, worker consultation and approval date.

The assessment features can preserve human sign-off and source history. Prevent approval while critical identity, exposure or control questions remain open.

How can AI be used responsibly in the workflow?

Use AI for extraction and drafting where it reduces transcription, then keep people responsible for context and control. Define mandatory inputs, source citations, confidence or missing-data flags, review checklists and audit history.

Monitor recurring errors and improve the workflow. If reviewers repeatedly correct the same field, change the prompt, input requirement or product control rather than relying on memory.

Use the help centre for review workflow guidance. A good AI-assisted assessment is faster to prepare because the evidence is organised, not because judgement was removed.

How should missing information be displayed?

Missing inputs should remain visible and block false certainty. A draft should distinguish “not provided”, “not found in source” and “not applicable with reviewer reason”. These states mean different things and should not become blank fields or reassuring defaults.

Critical gaps include uncertain product identity, absent task description, unknown quantity, unverified ventilation, missing exposure-limit source and unspecified PPE. Define which gaps prevent sign-off and which can become owned actions with interim controls.

If the AI inferred a value, label it as an inference and replace it with evidence before approval. A fluent sentence without provenance should be easier to challenge, not harder.

How can you review a risk rating independently?

Reconstruct the rating from verified inputs instead of checking whether the final number looks plausible. Identify the model or matrix, severity basis, exposure likelihood, control assumptions and residual-risk threshold. Change one input at a time and confirm the result follows the documented method.

Where the method uses qualitative bands, make the definitions visible. “Occasional” should map to a stated frequency; “effective LEV” should require evidence. Do not let the system silently reduce risk because a proposed control appears in the text.

What should a reviewer observe at the workplace?

Compare the draft with one representative task cycle and credible non-routine work. Watch setup, handling, cleaning and waste. Check worker position, splash or emission points, nearby people, ventilation use, glove changes and what happens when equipment jams or supply runs out.

Ask the operator which parts of the draft are impractical or missing. Check maintenance and cleaners because their exposure may differ substantially. Record the date and conditions of observation so a future reviewer knows what was seen.

How should changes to AI output be preserved?

Keep the generated draft, reviewer changes and approved version linked in an audit history. Record why material controls, facts or ratings changed. This helps distinguish model error, poor source data and missing user input.

Do not overwrite the evidence of an unsafe suggestion once it has influenced workflow. Use recurring corrections to improve the prompt, extraction rules and required fields. Sample approved assessments periodically to ensure reviewers are not normalising the same defect.

What should workers receive?

Workers need the approved assessment findings and usable instructions, not a raw AI transcript. Communicate hazards, exposure routes, controls, checks, PPE, emergencies and stop conditions in the format used at the task.

Explain that unusual conditions or product changes require escalation even when the assessment screen appears approved. Human review remains active after sign-off through feedback, incident reporting and scheduled or triggered review.

Frequently asked questions

Can AI approve a COSHH assessment?

AI can assist preparation and checks, but accountable approval requires a competent person who can verify workplace facts and control decisions.

Should every AI statement have a citation?

Chemical, legal and technical facts should be traceable to reliable sources. Workplace statements should be traceable to observations, records or named input.

What if the AI draft looks complete?

Completeness of format is not evidence of accuracy. Use the same source, context, exposure and control checklist for every draft.

Can AI choose PPE from an SDS?

It can extract supplier advice, but a competent reviewer must select PPE for the actual task, residual exposure, wearer and other controls.

coshhartificial intelligencerisk assessmentcompliance

Put this into practice with Safe Foundry

Build your chemical register, connect locations and create structured COSHH or DSEAR assessment records.