Industry11 min read

AI Document Search for Chemical Manufacturers: What It Can Answer, and What It Must Not

Compliance questionnaires get answered by hand, from safety data sheets and REACH files. Here is what citation-first AI can and cannot do for a chemical firm.

By Nic Chin|

A customer compliance questionnaire lands on a Tuesday. It is in a format nobody has seen before, it asks about substances across nine products, and the only person who can answer it is the technical manager — who will spend the next two days opening safety data sheets one at a time.

That is the job an AI document search tool should take on for a chemical manufacturer. It is also the job where getting it wrong has a name, a customer and a consequence. So this article is written in two halves: what this kind of tool genuinely does for a chemicals business, and — the half most vendors skip — the specific things it must never be allowed to do.

I build one of these tools, so treat the second half as the more useful one.

Why chemical document sets suit retrieval unusually well#

Most industrial document sets are hostile to text search because they are scans of third-party paperwork. Chemical document sets largely are not.

Safety data sheets follow a fixed 16-section structure in a fixed order. That format comes from Annex 4 of the UN Globally Harmonized System and is written into law on both sides of the Atlantic — into OSHA's Hazard Communication Standard at 29 CFR 1910.1200 Appendix D, and into REACH Annex II as amended by Regulation (EU) 2020/878. Section 2 is always classification. Section 8 is always exposure controls. Section 14 is always transport information.

A fixed section order is worth more to a retrieval system than people expect. It means "what is the exposure limit for this substance" has a predictable home in every document in the set, and a tool that finds the right document has very nearly found the right passage.

The rest of the corpus is similarly structured and similarly born-digital: standard operating procedures and work instructions, quality manuals under ISO 9001, 14001 and 45001, technical data sheets, certificates of analysis, COSHH assessments, and customer specifications. These are written in Word and published as PDF. They are not scanned, which means the extraction path is the straightforward one rather than the OCR one.

Start with the procedures, not the safety data sheets#

This is counterintuitive, because safety data sheets are where the pain is loudest. Start somewhere else anyway.

Standard operating procedures, work instructions, quality manuals and customer specifications change slowly. A procedure written in 2023 is usually still the procedure. That stability matters enormously for a first deployment, for a reason covered in detail further down: retrieval tools do not inherently know which revision of a document is current.

Procedures are also where a QHSE manager's time quietly goes. "What does our procedure say about decanting into secondary containers?" and "which work instruction covers the changeover cleaning between batches?" are asked constantly, answered from memory, and occasionally answered from memory of the previous revision. Those are good first questions for a tool, because the answer is checkable in seconds and the cost of a wrong one is low.

Once the team trusts the citations on procedures, move to the substance documents. Not before.

The three things that decide whether it is usable#

For a chemicals business the evaluation criteria are not the ones in a generic AI buying guide. They reduce to three.

1. Does every answer name the document it came from?#

Not "based on your files". The filename and the page. A QHSE manager answering a customer questionnaire is not looking for a summary — they are looking for the sentence they will paste into the questionnaire, plus the document they will name if anyone asks where it came from.

An answer without a citation is unusable here, regardless of whether it happens to be correct, because it cannot be checked before it goes out. This is why citation behaviour, not speed, is the first thing to test. We wrote a fuller version of this argument in why AI document search needs citations.

2. Does it refuse when the answer is not there?#

This is the one that separates a usable tool from a liability.

Ask it something your documents genuinely do not cover — a substance you do not handle, an exposure limit that is not in any of your sheets. A tool worth deploying says it does not have that. A tool that generates a plausible-looking exposure limit has just written a number into a customer questionnaire.

Refusal is not a weakness in this setting; it is the primary safety feature. Abstention, and why over-refusal is the safe failure, goes through how to measure this properly.

3. Where do the documents actually sit?#

Formulations, customer specifications and supplier declarations are commercially sensitive in a way that a generic productivity tool is not built for. The question to ask any vendor is not "is it secure" but "is my document set isolated from every other customer's, and at what layer". Application-level separation and database-level separation are different answers.

What it must not do, and this is not a disclaimer#

Three hard limits. A vendor who is vague about any of them is selling you a problem.

It must not make the regulatory call. A retrieval tool finds and cites what your own documents say. It does not classify a substance, determine whether you are compliant, or interpret a regulation. If a tool offers to tell you whether a product is affected by a classification change, it has crossed from retrieval into advice, and the liability for that sits with you, not the vendor. The correct behaviour is to show you what your dossiers and sheets say and let a competent person decide.

This matters practically, not just legally. When ECHA updated the REACH Candidate List in February 2026 — adding two substances and taking the list to 253 entries — the work that follows is determining which of your products are affected above the 0.1% weight-by-weight threshold that triggers the duty to inform customers. A document search tool can find every place a substance is mentioned across your sheets and specifications in minutes, which is genuinely most of the labour. It cannot and should not tell you what the answer means.

It must not be trusted on revisions. Covered in its own section below, because it is the thing most likely to bite.

It must not be assumed to read drawings. If a tool extracts text, it extracts the text printed on a drawing — a title block, a note, a label. It does not understand what connects to what. Nothing about "which valve sits upstream of P-101" is answerable from text extraction, and a tool that answers it anyway is guessing. When a vendor says "your technical files" or "all your engineering documents", ask them to list the document types by name. The umbrella nouns are where the disappointment lives.

The revision problem, stated plainly#

Retrieval tools find documents. They do not, by default, know which document supersedes which.

If your workspace contains revision 4 and revision 6 of the same safety data sheet, a query can be answered from revision 4 — correctly cited, verifiably real, and out of date. The citation check passes because the citation is genuine. The answer is still wrong.

For most industries this is an annoyance. In chemicals it is a safety gap, because the thing that changed between revisions is often a classification or an exposure limit.

There is one operational fix and you should insist on it before the first upload: keep exactly one current revision of each document in the workspace, and archive superseded revisions somewhere else. Not in a subfolder — outside the searchable set entirely. It is a real constraint, it is work, and any vendor who does not mention it is either unaware of the problem or hoping you do not notice it.

Our own product does not yet track document versions. That is why the instruction above is in this article rather than in a footnote.

The benchmark question you should ask, and our honest answer#

If a vendor quotes an accuracy figure, ask what corpus it was measured on. "Internal evaluation" means they wrote the test and marked their own homework.

For our part: SureCiteAI publishes reproducible benchmark runs across public corpora — CUAD v1 legal contracts, openFDA drug labels and SEC EDGAR filings — measuring retrieval hit rate, abstention accuracy, citation hallucination rate, and calibration, with deterministic samplers and a fixed seed so anyone can reproduce them.

None of those corpora is a safety data sheet. We have not benchmarked chemical documents.

The closest structural analogue in the set is the openFDA drug label suite: regulated product safety documentation, standardised sections, written to be read by someone who needs the exact wording rather than a paraphrase. That is a genuine analogue and it is also not the same thing as testing on your documents.

The honest position is that the published evidence is strong and it is about other document types. If you are evaluating for a chemicals business, the only evidence that counts is what the tool does on a folder of your own procedures, which takes about five minutes to set up and an afternoon to judge. How to tell if an AI document tool is actually accurate sets out the tests to run.

A realistic first month#

Imagine a 30-person blender with a technical manager, a quality manager and no internal IT.

Week one. Upload the quality manual, the SOPs and work instructions, and the current customer specifications. One revision each. Ask twenty questions the team already knows the answers to, and check every citation. The point is not to be impressed; it is to find out whether the tool refuses cleanly when you ask about something absent.

Week two. Add technical data sheets and certificates of analysis. Spreadsheets are indexed with their rows and columns intact, so a substance register in Excel is searchable — though formulas, cross-sheet references and pivot tables are not evaluated, only the cell text is read.

Week three. Add safety data sheets, having first confirmed one current revision per product. Run the last compliance questionnaire you answered by hand, and compare.

Week four. Decide. The measure is not whether the tool is clever. It is whether the technical manager would let an answer go to a customer on the strength of the citation.

Frequently asked questions#

Can it answer a customer compliance questionnaire for me? It can find and cite the passages in your own documents that answer each question, which is the majority of the labour. It should not produce the submitted answer unreviewed, and it cannot make a determination your documents do not already contain. Treat it as a very fast way to assemble the evidence, with a competent person still signing it off.

Does it work on scanned supplier safety data sheets? Text on a scanned page is extracted via OCR, so a scanned sheet is searchable. Quality depends on the scan. Since most modern sheets arrive as born-digital PDFs, this is usually a smaller problem in chemicals than in trades that deal mainly in scanned third-party certificates.

What about REACH dossiers — are they too large to search usefully? Length is not the obstacle; structure is, and dossiers have plenty. The practical limit is that very large spreadsheets get truncated rather than partially read, and the truncation is announced in the indexed text rather than happening silently. For documents, retrieval finds the relevant section regardless of overall length.

Is our formulation data safe in a hosted tool? Each company gets an isolated workspace on its own subdomain, with tenant separation enforced at the database row level rather than only in application code, and documents are not used to train models. If your procurement has a hard SOC 2 Type II requirement, say so at the first conversation — we do not hold that certification today, and finding out late wastes everyone's time.

We are in Malaysia and some documents are in Bahasa Malaysia. Does that work? The system is optimised for English document sets. A mixed corpus will give noticeably weaker results on the non-English portion, and that is worth testing on your own files before committing rather than taking anyone's word for it.


If your compliance questionnaires are answered by one person opening one sheet at a time, the fastest way to find out whether this helps is to put a folder of your procedures in front of it and ask the questions you already know the answers to. Try SureCiteAI or see how it handles SOP and policy retrieval first.

Stop Searching. Start Finding.

Upload your documents and get AI-powered answers in minutes. No coding, no IT department, no complex setup.

No credit card required. Setup takes less than 5 minutes.