Comparison10 min read

ChatGPT Projects for Company Documents: What It Does Well, and Where It Stops

ChatGPT Projects keeps files and chats together and answers well from them. Here is exactly where it stops being enough for client-confidential firm work.

By Nic Chin|

The verdict first, because you probably came here for it: ChatGPT Projects is a genuinely good place to keep a working document set, and a poor place to keep a firm's system of record. It is excellent for a piece of work with a beginning and an end. It is not built to be the thing your team queries every day about client files, and the reasons are structural rather than a matter of the model being good or bad.

I build a competing tool, which you should factor in. I have tried to make the criticisms specific enough that you can check each one yourself in about twenty minutes.

What Projects actually is#

Projects in ChatGPT are containers that hold files, chats and custom instructions together in one place. You upload reference material — PDFs, spreadsheets, documents, images — and every conversation inside that project can draw on it, along with the previous conversations in the same project. Projects are available on the paid tiers.

The design intent is a persistent working context. Instead of re-uploading the same five documents and re-explaining the same background at the start of every chat, you set it up once and the context persists. For that purpose it works well and the convenience is real.

Custom GPTs solve an adjacent problem: a configured assistant with fixed instructions and an attached knowledge set, shareable with other people.

Where it genuinely wins#

Being fair about this matters, because a comparison that finds no merit in the alternative is marketing rather than analysis.

Reasoning quality over a small set. Give it four documents and a hard analytical question, and the quality of the thinking is high. If you need something argued through rather than looked up, this is a good tool.

Drafting from context. "Write a client update based on these three files" plays directly to its strengths. Retrieval-first tools are often worse at this, because they are optimised to find and quote rather than to compose.

Iterative work. The persistence across chats is the feature. A piece of work that runs for three weeks benefits enormously from context that does not reset.

No setup cost. If your firm already pays for ChatGPT, a project exists the moment you make one.

For a consultant working through one client engagement, this is often the right answer and a second subscription would be waste.

Where it stops#

Five limits. None is a bug; each is a consequence of what the product is for.

1. Answers do not reliably cite a source you can open#

This is the big one for professional services. An answer may mention which file it drew on, but you do not get a dependable citation to a specific document and page on every answer, presented as something to click and verify.

That gap is the difference between a tool that drafts and a tool that can be relied on. A paralegal who cannot verify where a clause came from has to go and find it manually, which was the original job. We set out the full argument in why AI document search needs citations.

2. It is built to be helpful, not to refuse#

Ask a general assistant something your uploaded documents do not cover and it will typically answer anyway — from general knowledge, from plausible inference, from the shape of what an answer should look like. That is correct behaviour for an assistant and dangerous behaviour for a system of record.

The failure is quiet. You do not get a warning that the answer came from outside your files, so the only defence is already knowing the answer, which defeats the purpose. A system designed for document Q&A treats refusing to answer as a feature and measures how often it gets that right.

Test it in five minutes. Upload a handful of your documents. Ask something adjacent that they genuinely do not cover — a policy you know is not in there. See whether you get a refusal or a confident paragraph.

3. File caps make it a working set, not a library#

OpenAI documents limits on uploads, including caps on file size and on how many files a project holds, varying by plan. There are also limits on how much of a very large file is read, so an enormous document can be partially rather than fully available.

Projects are sized for a working set — the documents for this engagement. They are not sized for a firm's document library, and no amount of organisation gets around a per-project cap. Once you are splitting one matter across three projects because of a file limit, you have left the intended use case.

4. There is no per-matter isolation model#

Everything in a project is available to everything else in that project, and the boundaries are the ones you remember to draw.

For a firm handling multiple clients this is an information-barrier problem rather than a features problem. Nothing stops a conversation in a shared project surfacing material from a matter the person asking should not see, because the separation exists only as a filing convention. A tool built for client work treats isolation as an enforced property rather than a habit.

5. Confidentiality depends on the tier, and people get this wrong#

The data handling differs between consumer and business tiers, and the assumption that "we pay for it, so it is fine" is where firms get into trouble. This deserves its own treatment and has one: can law firms use ChatGPT with client data works through what confidentiality actually requires before AI touches client files.

A straight comparison#

| | ChatGPT Projects | Citation-first document workspace | |---|---|---| | Best at | Reasoning and drafting over a working set | Finding and proving a specific fact | | Citations | Not a dependable per-answer source link | Every answer names its source document | | Behaviour when the answer is absent | Usually answers anyway | Designed to refuse | | Scale | Per-project file caps | Built for a firm's library | | Matter separation | A filing convention | Enforced isolation | | Setup | Immediate | Roughly five minutes | | Right for | One engagement, one person | A team answering document questions daily |

A fuller three-way version, including Notion AI, is in SureCiteAI vs ChatGPT vs Notion AI.

What a week of real use looks like#

Abstract comparisons are easy to nod along to, so here is the concrete version. Imagine a six-person consultancy running three client engagements.

Monday. A consultant makes a project for the largest engagement, uploads the brief, two discovery decks and a data pack, and spends the morning working through implications. Projects is excellent at this and the alternative would be worse. No complaint.

Wednesday. A colleague asks what the contract says about the notice period for a different client. It is not in this project. They either make a second project and upload a second set of files, or open a one-off chat and paste the contract in. Both work. Neither leaves a searchable trace anyone else can find next month.

Thursday. Someone asks a question the files do not cover. The answer arrives fluent, reasonable and sourced from nothing in particular. Nobody notices, because there is no citation to check and no refusal to prompt a second look.

The following month. A new joiner asks where the escalation procedure is documented. The knowledge exists — it is spread across four projects belonging to three people, plus a chat someone deleted.

Nothing in that week is a malfunction. It is what happens when a tool designed for one person's working context is asked to be a firm's memory. The failure is that the knowledge never accumulates anywhere queryable, and the second failure is that Thursday's answer is indistinguishable from a good one.

How to decide without taking anyone's word for it#

Four questions about your own situation.

Is this one piece of work, or an ongoing service? A project with an end date suits Projects. "Where is the clause" asked daily by six people is a different product category.

Would a wrong answer cost you something? If a wrong answer means a mild embarrassment, optimise for convenience. If it means a filing, a malpractice exposure or a customer commitment, optimise for verifiability.

How many documents, really? Under twenty for one piece of work, Projects is fine. Several hundred across many clients, the caps decide it for you.

Does anyone need to check the answer? If answers are acted on directly, citations are a nice-to-have. If a partner reviews before anything goes out, the citation is the product — it is what makes review take thirty seconds instead of thirty minutes.

The honest recommendation#

Most small firms should use both, for different jobs.

Keep ChatGPT for thinking, drafting and analysis over a handful of documents you have deliberately chosen. Use something built for document Q&A when the question is "what do our files say about this" and somebody is going to rely on the answer.

The mistake is not choosing ChatGPT. The mistake is letting a tool designed to be helpful become the place your firm looks up facts it will act on, and only finding out at the point where an answer that read perfectly turns out to have come from nowhere.

Frequently asked questions#

Does ChatGPT Projects remember my documents between chats?#

Yes — that is the central feature. Files and chats inside a project stay available to new conversations in the same project, so you are not re-uploading and re-explaining each time.

Can I share a project with my team?#

Yes, on the plans that support it, and shared projects work well as a context hub. What you do not get is per-matter access control inside the shared space: anyone with access to the project has access to everything in it.

Will ChatGPT tell me when an answer is not from my files?#

Not reliably. It is built to be useful, which means it will often draw on general knowledge without flagging that it has done so. If that distinction matters for your work, test it explicitly with a question your documents cannot answer rather than assuming.

Are custom GPTs better than Projects for a document set?#

They solve a different problem. Custom GPTs give a configured, shareable assistant with fixed instructions — good for a repeatable task. Projects give a persistent workspace for evolving work. Neither changes the citation and refusal behaviour discussed above.

Is this article just saying ChatGPT is bad?#

No. The reasoning quality is excellent and for single-engagement work it is frequently the better choice. The argument is narrower: a general assistant optimised for helpfulness is the wrong shape for a firm's document system of record, and the gap shows up in citations, refusal and isolation rather than in intelligence.

What should I ask a vendor before switching?#

Ask for a citation on every answer and open one. Ask a question your documents do not cover and see whether the tool declines. Ask where the separation between clients is enforced. The full list of questions worth asking goes further, but those three separate the categories quickly.


If the "what do our files say" question comes up daily at your firm, SureCiteAI is a private per-company workspace where every answer cites the exact source document and the system refuses when your documents do not contain the answer.

Stop Searching. Start Finding.

Upload your documents and get AI-powered answers in minutes. No coding, no IT department, no complex setup.

No credit card required. Setup takes less than 5 minutes.