Should Your Firm Use an AI Agent for Company Documents?
Agents like Hermes and OpenClaw act on your files; ChatGPT and Claude read them. Neither verifies its citations — why that matters for client work.
Probably not — at least not as the place your client documents live. An AI agent is built to act on your behalf; a document workspace is built to answer from a fixed set of files and prove where the answer came from. Those are different jobs, and the second one is what professional-services work actually requires. This guide explains the four options your team is likely weighing — ChatGPT, Claude, Hermes, and OpenClaw — what each was designed for, and how to decide between them without learning the difference the expensive way.
What's the difference between an AI assistant and an AI agent?#
Both terms get used loosely, so here is the plain-English split.
An assistant is reactive. You ask, it answers, the conversation ends. ChatGPT and Claude are assistants — extremely capable ones, with file upload and long memory, but the shape of the interaction is a conversation.
An agent is autonomous. It runs continuously, keeps a memory of you between sessions, and takes actions: browsing, running commands, sending messages, triggering scheduled tasks. Hermes Agent from Nous Research describes itself as "The Agent That Grows With You," with persistent memory that "learns your projects, auto-generates skills, and never forgets how it solved a problem." OpenClaw describes itself as "Your AI assistant, on your own hardware, in every chat app you already use" — a self-hosted gateway that wires AI agents into Discord, Slack, WhatsApp, iMessage, Signal, Microsoft Teams and Telegram.
A document workspace is neither. It is a narrow tool with one job: given your firm's files, answer questions about them and show the source. It doesn't act, doesn't roam, and doesn't reach outside the documents you gave it. That narrowness is the product.
The failure mode firms hit is picking from the wrong category. An agent that can do anything is impressive in a demo and unbounded in a compliance review.
What do ChatGPT and Claude actually do with your documents?#
They read them, well, and then answer from everything they know.
That last clause is the whole problem. When you upload a contract and ask about an indemnity cap, a general assistant will answer from the document if the answer is there — and from its training data if it isn't. Both answers arrive in the same confident tone. There is no visible seam between "I found this in your file" and "this is how indemnity caps usually work."
The citation problem follows from the same design. When an assistant names a source, that name is generated text, not a lookup. Nothing in the system tests the cited document against the set of documents actually retrieved before the answer reaches you. Independent comparisons of the major assistants for enterprise use tend to separate them on how often they fabricate references rather than on whether the architecture prevents it (IntuitionLabs, 2026 enterprise guide). Claude is generally the better-behaved of the two here. "Better-behaved" is a probability, not a guarantee, and a probability is a difficult thing to put in front of a client.
The third issue is separation. A shared assistant workspace has no concept of a matter, an engagement, or an ethical wall. Everyone searching the same corpus sees the same corpus. For a firm holding documents for two clients with adverse interests, that isn't a preference — it's a structural problem. We covered that specific question in can law firms use ChatGPT for client-confidential work.
What are Hermes and OpenClaw, and who are they for?#
Both are open-source, MIT-licensed and self-hosted, both are genuinely impressive engineering, and neither was built for cited document retrieval.
Hermes Agent is a persistent daemon you run on your own infrastructure. It accumulates memory across sessions, generates its own reusable skills from experience, and reaches 200+ models through Nous Portal. Nous has announced business and enterprise tiers, which reporting describes as running on customer-controlled infrastructure with single sign-on and service-level agreements (tao.media). For an organisation whose hard requirement is sovereignty — the AI must run inside our walls, full stop — that is a serious answer, and a better one than any hosted product can give.
What its own product page does not mention anywhere: document search, citations, retrieval accuracy, or multi-tenant isolation. That isn't a criticism. It's a different product.
OpenClaw is a gateway between your messaging apps and AI agents, running as a single process on your own hardware with no telemetry by default. It is deservedly popular with developers. For a firm holding privileged material, though, look carefully at the shape of it: an agent with tool access and persistent memory, bridged into WhatsApp and iMessage. Every one of those bridges is a path for content to move, and an agent that reads untrusted inbound messages while also holding your client files is an injection surface, not a confidentiality posture.
Both also carry an operating cost that the free licence hides. Someone has to patch it, monitor it, hold the model API keys, and answer for it when a client asks where their file went. For a 12-lawyer firm with no IT department, that someone is a partner's evenings. We wrote about this trade-off in why not to build your own AI document search — self-hosting somebody else's agent lands in much the same place.
How do the four approaches compare?#
| | ChatGPT / Claude | Hermes Agent | OpenClaw | SureCiteAI | |---|---|---|---|---| | Primary job | General assistant | Autonomous agent with memory | Messaging-to-agent gateway | Cited answers from your files | | Where it runs | Vendor cloud | Self-hosted or Nous cloud | Self-hosted, your hardware | Private tenant on your own subdomain | | Answers only from your documents | No — falls back to training data | Not its purpose | Not its purpose | Yes — refuses when the files don't cover it | | Citations checked against retrieved sources | No — generated, unverified | Not applicable | Not applicable | Yes — unverified citations are blocked | | Published retrieval accuracy | Not published | Not published | Not published | Yes — public benchmark runs | | Separation between clients or matters | Shared workspace | Task-level sandboxing | Per-instance | Row-level tenant isolation | | Who operates it | Vendor | You | You | Vendor | | Takes actions on your behalf | Limited | Yes, by design | Yes, by design | No — answers only |
Read the "published retrieval accuracy" row twice. It is the one thing a buyer can check before signing, and it is blank in three of the four columns.
Why does a verified citation change the calculation?#
Because in professional services you are not buying answers. You are buying the ability to defend one.
The distinction is architectural, not attitudinal. In a system with a citation verifier, the model proposes a source and a separate check tests that source against the documents actually retrieved for that question. A citation that fails the check is rewritten, or the answer is held back. The model's honesty stops being load-bearing.
On SureCiteAI's most recent published run (25 September 2026, 297 cases across legal, accounting, healthcare, real-estate and consulting suites plus the external PatronusAI FinanceBench suite), that produced zero citation hallucinations — 0 out of 297 — at a 93% retrieval hit rate. On FinanceBench specifically, 150 questions over 84 real 10-K and 10-Q filings, retrieval surfaced the correct source filing 140 times out of 150, with an expected-calibration error of 0.064. That run's pass rate did not compare answers to the reference answers, so read it as a ceiling until the stricter check ships. The methodology and the run artefacts are public, and benchmarks explained walks a non-technical reader through how to read them.
The number we're less flattered by is just as instructive. The legal CUAD suite passed 17 of 35 — 49%. CUAD is adversarial by construction: many of its questions ask about clause categories simply absent from the contract being queried, and the correct behaviour is to say so. Most of that gap is the system declining to commit, not the system inventing. We publish it anyway, because a perfect score on a corpus we wrote ourselves would prove nothing, and because refusing to answer is the safe failure for work that carries professional liability. A tool that is occasionally too cautious costs you a re-read. A tool that is occasionally confident and wrong costs you a client.
What does the difference cost in billable hours?#
Work it through with hypothetical but ordinary numbers. Imagine a 12-lawyer firm billing an average of $280 an hour.
Scenario one: an unverified answer gets used. A cited clause turns out not to exist in the document the assistant named, and someone catches it at review — the good case. Re-checking the answer, tracing the real source and re-drafting the affected section costs three hours across two people. Roughly $840, plus the partner's confidence in the tool.
Scenario two: nobody catches it until the other side does. There is no clean way to price that hour.
Scenario three: the verified system abstains on a question the documents actually answered — the over-cautious failure. A lawyer spends twelve minutes opening the file and reading it. Call it $56.
Three near-misses a month in the first scenario is about $2,500. Twenty over-abstentions a month is about $1,100 — and produces no exposure at all. That asymmetry, not raw answer quality, is the real argument. If you want a structured way to pressure-test it with any vendor, the questions to ask an AI document search tool checklist is built for exactly that conversation.
When should you not choose SureCiteAI?#
Three honest cases, and they are real ones.
If you need on-premise or air-gapped deployment. SureCiteAI is a hosted product with per-tenant isolation; it is not something you run inside your own datacentre today. If that requirement is absolute, Hermes is the more honest fit, and we would rather say so now than in month three.
If you need the AI to do things rather than answer things. Filing, scheduling, sending, triggering workflows — that is agent territory. SureCiteAI answers from documents and stops there, deliberately.
If you need broad multi-language support today. The system is English-optimised.
And if you are a solo operator with no confidentiality constraints, a ChatGPT or Claude subscription is genuinely the right call. Not every problem is a platform problem.
The bottom line#
Agents and assistants are optimised for breadth: do more things, for more people, from more sources. A document workspace for professional services is optimised for the opposite — a closed corpus, a verified citation, and a refusal when the files come up short.
If your documents are ordinary business files and nobody is harmed by a confidently wrong answer, use the assistant you already pay for. If a wrong citation in front of a client, a regulator or a court is a professional-standards problem, then the architecture has to carry that weight, not the model's good manners.
SureCiteAI runs from $39/month, and every product claim in this article maps to a public benchmark run or to source code you can read. Compare plans, or book a 30-minute walkthrough and bring your hardest document.
Frequently asked questions#
Can an AI agent like Hermes or OpenClaw search my company documents?#
It can be made to, but neither publishes retrieval accuracy or citation-verification results, and neither product presents document search as its purpose. You would be building a document-QA system on top of an agent framework and owning the result.
Is a self-hosted AI agent more secure than a hosted document workspace?#
Self-hosting moves control to you; it does not automatically move risk down. You inherit patching, key management, monitoring and incident response. For a firm with no IT function, an isolated hosted tenant with a documented boundary is usually the lower-risk option. The honest question is who is better placed to operate it.
Why can't ChatGPT or Claude just be told to cite sources?#
They can be, and they will comply. The issue is that the citation is generated text rather than a verified lookup — nothing downstream tests the named source against the documents actually retrieved. Prompting improves the odds; it does not change the architecture. Why citations matter covers the distinction in detail.
What is a citation hallucination?#
An answer that names a source which does not support it: a document that was never retrieved, a page that does not contain the quoted text, or a file that does not exist. It is more dangerous than an ordinary wrong answer, because the citation is the thing that makes it look checkable.
Should a vendor publish a benchmark score below 100%?#
Yes — and be suspicious when one doesn't. A perfect score on a corpus the vendor authored proves nothing. A mid-range score on a public benchmark that anyone can re-run, with the run artefacts published, is the only accuracy claim a buyer can independently verify.
Where this applies
SureCiteAI for law firms
Contract, matter and due-diligence search with a citation on every answer.
SureCiteAI for accounting firms
Workpaper, return and engagement-file search that names its source.
SureCiteAI for real estate teams
Lease, listing and transaction-file search across a whole brokerage.
SureCiteAI for consultancies
Proposal, deliverable and research-library search for client work.
Accuracy benchmarks
297 test cases across six suites on public corpora, with the methodology, run artifacts and failure cases published in full.
Stop Searching. Start Finding.
Upload your documents and get AI-powered answers in minutes. No coding, no IT department, no complex setup.
No credit card required. Setup takes less than 5 minutes.