Manual vs AI COSHH Assessments: Honest Comparison

- Manual assessment offers direct expert control but can waste time retyping SDS facts and produce inconsistent records across assessors.
- AI can accelerate extraction and drafting, but it cannot observe the workplace, confirm exposure or accept legal responsibility for the assessment.
- The strongest workflow automates traceable document work while requiring a competent person to verify source facts, add site context and approve controls.
- Evaluate AI tools with representative SDS, known failure cases, audit trails and correction workflows rather than headline speed alone.
AI-assisted COSHH assessment is faster at reading and structuring SDS information, while manual assessment gives a person direct control over every entry. Neither method is sufficient without competent workplace judgement, so the most reliable model is usually AI for traceable drafting and humans for verification, exposure assessment and approval.
Speed matters, but the decision should compare corrected output, not the time to produce a first draft.
What does a manual COSHH assessment involve?
A manual process gathers the label and SDS, identifies hazardous properties, observes the task, evaluates routes and level of exposure, selects controls and records actions and review triggers.
Its strengths include:
- Direct engagement with the workplace.
- Flexible treatment of unusual processes.
- Clear professional judgement when the assessor documents reasoning.
- No dependence on model availability or file processing.
- Easier control for very sensitive or restricted documents where local systems are required.
Its weaknesses are repeated transcription, inconsistent wording, slow comparison of revisions and reliance on individual memory. Copy-and-paste can make a manual assessment look specific while carrying old product information forward.
What can AI actually do?
AI tools can support several bounded tasks:
- Extract product identity, hazard statements and ingredients from an SDS.
- Structure handling, storage, exposure and emergency information.
- Compare document versions.
- Prefill a draft form.
- Flag missing fields or inconsistent answers.
- Suggest questions for the assessor to resolve.
- Create a readable summary from confirmed data.
AI cannot see the task unless reliable workplace context is supplied. It does not know whether extraction is switched on, the worker bypasses a lid, the room is smaller than stated or the glove is changed too late.
A fluent paragraph is not evidence that the system observed the work.
How do speed and cost compare?
| Factor | Manual | AI-assisted |
|---|---|---|
| First draft | Slower, especially for long SDS | Often faster |
| Site observation | Human time required | Human time still required |
| Repetitive data entry | High | Reduced where extraction is accurate |
| Correction | Familiar but manual | Can be quick or extensive depending on errors |
| Consistency | Depends on assessor and template | Stronger structure, but consistent errors are possible |
| Direct cost | Staff or consultant time | Licence, credits and review time |
| Scale | Constrained by competent hours | Document throughput can scale faster |
| Accountability | Employer and competent assessor | Employer and competent assessor remain accountable |
Calculate cost per approved, implemented assessment, not per generated report. Include document preparation, correction, observation, consultation, sign-off, worker communication and review.
Which method is more accurate?
Accuracy has several dimensions:
- Source accuracy: were SDS facts transcribed correctly?
- Applicability: do they match the product and market?
- Exposure accuracy: does the assessment describe real work?
- Control accuracy: are measures suitable, available and maintained?
- Document accuracy: are links, versions and statuses clear?
AI can outperform hurried transcription while still fail on a poor scan or complex table. A human can notice implausible output while still overlook a copied concentration.
The system should show source page or section for extracted facts, mark uncertainty and preserve corrections. Avoid tools that turn an unsupported inference into an authoritative-looking value.
What are the main AI failure modes?
Common risks include:
- Reading the wrong product or supplier from a multi-product document.
- Missing a qualifier in a table or footnote.
- Confusing ingredient concentration ranges.
- Treating supplier PPE advice as the final workplace control.
- Inferring a workplace exposure limit that is absent or from another jurisdiction.
- Generating a risk rating without sufficient exposure evidence.
- Repeating confident but unsupported compliance language.
- Losing the connection between an answer and its source.
Prompt instructions reduce some behaviour but do not guarantee accuracy. Use technical controls: validation, provenance, required review, permissions and audit history.
What are the main manual failure modes?
Manual work has familiar risks:
- Using an old SDS because the filename looks right.
- Copying an assessment for a similar product without checking differences.
- Focusing on inhalation and missing skin exposure.
- Selecting PPE before considering elimination or engineering controls.
- Treating one product as one task.
- Leaving action ownership blank.
- Approving a document without speaking to the worker.
“Human-made” is not a quality mark. A strong manual system needs templates, competence, peer review and document control.
What should always remain a human decision?
A competent person should confirm:
- The product and SDS match physical stock.
- The activity description reflects normal and non-routine work.
- People and routes of exposure are complete.
- Existing controls are present and effective.
- Additional controls follow the hierarchy and are practicable.
- Monitoring or health surveillance needs have been considered.
- Actions have owners and timescales.
- Workers receive understandable information.
- Sign-off and review triggers are appropriate.
AI can organise evidence for these decisions but should not conceal where evidence is missing.
What does a good hybrid workflow look like?
A review-first process is:
Upload current SDS → extract with provenance → verify product and hazard facts → add task and site context → assess exposure → select controls → resolve actions → competent sign-off → communicate → review after change.
Lock supplier facts to the confirmed SDS version and distinguish them from assessor judgements. If a person corrects extracted data, record the correction rather than silently regenerating it.
Safe Foundry's workflow uses AI as optional assistance while keeping confirmation and workplace decisions explicit. Review features, pricing and the FAQ when assessing fit.
How should an AI tool be evaluated?
Build a representative test set, with permission to use every file. Include digital PDFs, scans, mixtures, multi-page tables, unusual suppliers and revised versions.
Measure:
- Correct product and revision identification.
- Field-level accuracy and source traceability.
- Unsupported statements.
- Time to correct, not only time to generate.
- Handling of uncertainty and missing data.
- Permission, retention and deletion controls.
- Export quality and vendor lock-in.
- Ability to stop sign-off when review is incomplete.
Have competent assessors compare final outputs against the same acceptance criteria.
What is the honest conclusion?
Use manual assessment where volume is low, context is unusual or data cannot enter the proposed system. Use AI assistance where repetitive SDS work is the bottleneck and the platform keeps sources and review visible.
Do not choose between humans and AI as if they perform the same job. Automate the clerical layer, strengthen the evidence trail and reserve judgement for people who can see the workplace and act on the result.
Frequently asked questions
Can AI write a COSHH assessment?
AI can draft or prefill parts of an assessment, but a competent person must verify the SDS facts, workplace exposure, controls and final decision.
Are manual COSHH assessments more accurate?
Not automatically. Manual work can contain transcription and consistency errors, while AI can misread or infer data. Accuracy depends on evidence, workflow and review.
Does using AI transfer COSHH responsibility to the software provider?
No. The employer remains responsible for ensuring the assessment is suitable and controls are implemented and maintained.
How should I test AI COSHH software?
Use a varied, anonymised test set including scans, tables, mixtures and revised SDS, then measure extraction accuracy, traceability, correction effort and final assessment quality.
Learn how to assess process-generated dust and fumes without an SDS by defining the process, exposure routes, evidence and practical controls.
COSHH Assessment Review Checklist for UK WorkplacesUse this COSHH assessment review checklist to compare records with current tasks, products, exposure routes, controls and workplace evidence.
COSHH Control Measures: How to Go Beyond PPELearn how to choose COSHH control measures beyond PPE by controlling exposure at source, using a worked activity example and review checklist.
