ArticlesPrivacy

Should you paste unpublished research into ChatGPT?

What actually happens to a paragraph you send to a cloud AI tool, which of the risks are real, what publishers and funders currently say about it, and how to tell whether a tool is genuinely local.

11 min readLocal AI Series
Diagram comparing where an unpublished draft travels with a cloud AI assistant versus local on-device processing

For most drafting help on work you intend to publish anyway, the risk is low and mostly about disclosure. For unpublished data, participant material, anything patentable, and anything you received under confidentiality, the answer is no, and in the peer review case it is a rule rather than advice.

Start here

The short answer, with the caveat that matters

The question people actually mean when they ask this is some version of “will I get in trouble?” The honest response has two halves. Pasting your own already-public literature review into a chatbot to tighten the prose is a small risk that mostly comes down to whether your institution wants it disclosed. Pasting a participant transcript, an unfiled invention, a manuscript you are reviewing for a journal, or the results section nobody has seen yet is a different act with different consequences.

The confusion comes from treating these as one question. They are not. What separates them is not how clever the tool is. It is whether the text crosses a network boundary and who is on the other side of it.

Mechanics

What actually happens to the text you paste

When you send a paragraph to a hosted assistant, it travels over the network to the provider’s infrastructure, gets processed there, and is stored for some period afterwards. That storage is normal engineering. Providers keep recent conversations to run abuse detection, to debug, and to let you scroll up. None of that is sinister and most of it is unavoidable if the model runs on someone else’s hardware.

The part worth understanding is that the defaults differ sharply by product tier, and people generalise from the wrong one. Consumer chat plans have historically used conversations to improve models unless you turn that off, with the control buried in settings. Business, enterprise and API tiers generally do not train on customer content by default, and say so contractually. Same company, same model, very different data posture, and the version your lab has a licence for is often not the version you opened at home.

So “is ChatGPT safe for my thesis” has no single answer, because ChatGPT is not one thing. Check which tier you are actually using, then read that tier’s current terms rather than a blog post about them. These policies get revised, sometimes quietly, and anything I write here has a shelf life.

Assessment

Which risks are real and which are folklore

The fear that a model will regurgitate your thesis verbatim to a stranger is the one people raise most and the one that matters least. Memorisation of a single document seen once is not how these systems typically behave, and the training pipelines that might ingest your text are separated from any given user’s session by a great deal of machinery. Possible in principle. Vanishingly unlikely to be what actually hurts you.

Four things are genuinely worth your attention.

Confidentiality you agreed to keep

This is the sharpest one, and the only one where the answer is a flat prohibition rather than a judgement call. If you are peer reviewing a manuscript, that document is not yours. You received it in confidence. Uploading it to a third-party service is a breach of that confidence regardless of what the service then does with it, which is why major publishers and funders have written explicit rules about it rather than leaving it to discretion.

The same logic covers participant data under an ethics approval, material under an NDA or a data use agreement, and unpublished work a colleague shared with you. The obligation attaches to the document, not to your intentions.

Novelty, if there is any chance of a patent

Public disclosure before filing can destroy patentability. Most jurisdictions outside the United States apply absolute novelty with no grace period, and the US grace period is limited and narrower than people assume. Whether a prompt to a commercial service counts as public disclosure is genuinely unsettled and I have not seen it tested cleanly. If your work might be commercialised, that ambiguity is the whole problem. Ask your tech transfer office before, not after.

Disclosure obligations

Nearly every publisher now requires authors to declare substantive AI use, and none of them will list a model as a co-author. This is the risk that most often actually bites people, and it is entirely avoidable. It is an administrative step, not a moral failing. Find your target journal’s statement, follow it, move on.

The provider having a bad day

Retained data can be exposed by a bug or a breach. This has happened to large providers before and will happen again, to someone, eventually. Your thesis is not a high-value target on its own. It just happens to be sitting in the same building as things that are.

The rules

What publishers and funders currently say

The specifics shift, so treat what follows as a map of where to look rather than a citation you can lean on. Check the current version before you rely on any of it.

WhoRoughly what they requireWhere to check
Most large publishers, Elsevier among themAuthors may use AI to improve readability and must disclose it. AI cannot be an author. Reviewers must not upload manuscripts to generative AI tools, on confidentiality grounds.The publisher's generative AI policy page for authors, and the separate one for reviewers
Funders and grant reviewersSeveral, including NIH, have prohibited generative AI in peer review of applications outright.The funder's peer review integrity guidance
Your universityVaries enormously. Some have institutional licences with no-training terms; some ban tools for thesis work entirely.Research office or graduate school AI guidance, not the IT helpdesk
Your ethics approvalSending participant data to a third-party processor is usually a change to your data handling, which usually means an amendment.The data management section of your own approved protocol

The pattern across all four is the same. Nobody objects to the technology. They object to documents leaving the custody they were entrusted to.

Verification

How to tell whether a tool is genuinely local

“Privacy first” on a landing page means nothing on its own. Plenty of tools marketed that way run a local interface over a cloud API, and some let you choose, defaulting to the cloud. Three checks settle it, and none need much technical skill.

Pull the network cable, or turn off wifi, and use the feature. If it still works, the model is on your machine. If it spins and fails, it was never local. This single test beats any amount of marketing copy.

Then look at what the app downloaded. Local models are large files, usually somewhere between one and several gigabytes, and they have to live on your disk. An app claiming on-device AI that installed at forty megabytes is doing the work somewhere else.

Last, watch the traffic. macOS tools like Little Snitch or LuLu will show you every outbound connection an app attempts. Run the AI feature and see whether anything leaves. This is more work than most people will do, and it is also the only proof that does not depend on trusting anyone.

In practice

A rule you can actually apply

Ask one question before you paste anything: if this text appeared on a public website tomorrow, what would break? For a paragraph of your own literature review, nothing much breaks, so use whatever helps you write it. For an interview transcript, an unfiled result, or a manuscript somebody trusted you to review, something breaks badly, and no convenience is worth it.

Most of what a researcher actually needs from AI is retrieval rather than generation. Finding the paper that said the thing. Pulling method and sample size out of forty PDFs. Checking whether your own library supports a sentence you just wrote. None of that requires a frontier model and none of it requires the text to leave your machine, which means for the largest slice of the work the whole dilemma is optional.

Where it is not optional, disclose it and stop worrying. The people setting these rules are not trying to catch you out. They are trying to keep confidential documents confidential, which is the same thing you want.

Run it on your Mac.

Everything in this article ships inside the app. Private, fast, and free for the individual creator.

Download on the App StoreFree on the App Store