When AI Starts Inventing Colleagues
A new preprint from Samsung and the University of Warsaw has a very odd message for anyone watching tech news and ai policy: generative models are no longer just producing sloppy paragraphs and confident nonsense. They’re also inventing people, then assigning those people jobs inside academic publishing.
That sounds absurd until you see the pattern. A model will conjure a name, attach it to a paper as an expert and then, in another file, reuse the same person as a coauthor or an author. The fake identities don’t stay in one lane. They wander across machine-made documents like they’ve got a badge and a lunch calendar.
Elena Vasquez’s one of the names that keeps popping up in this kind of output. In one place. She appears as the sort of specialist who would be cited for authority. In another, she shows up as a contributing author. Marcus Chen does the same sort of work, except he seems to get drafted into different academic roles depending on the prompt, the model and whatever mood the system’s in that day. The names feel oddly ordinary, which is part of the problem. They don’t read like sci-fi code names. They read like people you might actually email.
Once the same fake names start collecting bylines, the problem stops being a typo and starts looking like a paper trail.
That’s what makes this story more than a funny little hallucination story. A model inventing a random person in a draft’s one thing. It repeatedly recycling the same nonexistent people across documents is another. At that point, the machine isn’t merely failing to remember reality. It’s building a reusable cast list for scholarship that never happened.
And yes, the joke almost writes itself. Academia’s spent years worrying about plagiarism software, citation farming, predatory journals and the usual mess of publish-or-perish behavior. Now it has to contend with synthetic colleagues who never existed, never reviewed anything and apparently still got their names on the paper. The whole thing has the energy of a bad faculty directory, except the directory keeps generating more faculty.
The Samsung and University of Warsaw preprint treats this as more than a curiosity. Once fake identities leak into scholarly publishing, the issue stops being about one model making one bizarre mistake. Not ideal. It becomes a question of trust. Who wrote the paper? Who reviewed it? Who’s the expert being quoted? If the same invented person can turn up in different roles across machine-made documents, then the line between a harmless fabrication and a contaminated record gets thinner than anyone in publishing would like.
That’s where digital culture meets power and politics in a way nobody invited. Academic writing depends on names as a form of proof. Names sort authority, assign credit and tell readers which claims belong to which humans. They can also muddy the records that journals, repositories and search tools rely on, if AI systems can invent a convincing-sounding colleague on demand. A fake person with a real-sounding CV is funny for about five seconds. After that, it starts to look like infrastructure trouble.
But the weirdest part’s how normal the names feel. Elena Vasquez. Marcus Chen. “ That’s probably why they work so well. They don’t trigger suspicion. And they blend in. The next question is no longer whether a model can make up a person, and once synthetic identities blend in. It’s whether the academic system can keep pretending those people are harmless.
The names keep coming back, though and that’s where the story gets even stranger.

The Same Fake Names Keep Coming Back
After the first wave of surprise passes, the stranger part’s that these names don’t seem random at all. In the preprint, the authors describe something they call “name priors,” which is their way of saying some language models keep defaulting to the same personal names in the same kinds of prompts instead of inventing fresh ones each time. Ask for a scientist and you may get one phantom. Ask again later, and the same phantom’s back, with a new title and a tidier affiliation. That’s less like invention and more like a casting director with a very narrow contact list.
The paper methods study on recurring AI names and model fingerprints treats those names as more than stray hallucinations. The authors found correlated clusters, meaning some invented people tend to travel together. A name that appears in one synthetic paper often shows up beside the same handful of other names in another. So the fake author is not floating alone in the void. They arrive with colleagues, coauthors, and sometimes a whole little office of paper people.
That pattern matters because it suggests the models are not improvising from scratch every time. They appear to have default preferences. In fiction prompts, Claude repeatedly surfaced names like Elias Thorne, the sort of solemn, faintly literary name that sounds ready to brood in a corner with a chipped mug. For software roles, Marcus Chen kept popping up, usually in the clean, résumé-friendly way that makes the text look plausible for a few seconds longer than it should. Gemini and ChatGPT had their own recurring picks too, with different clusters showing up across academic-style prose, code-adjacent writing, and character sketches. A separate report on the same pattern noted that those repeating names can sometimes function like a rough model fingerprint, at least until the whole thing gets noisier news report on recurring AI names.
The oddity is not that AI invents names. It’s that it prefers the same ones, over and over, like a badly cast ensemble that never got new headshots.
There’s a practical reason to care about that, and it’s nothing to do with trivia-night bragging rights. If the same invented people keep appearing, editors and reviewers may eventually spot a manuscript that smells off before they spend time chasing references. “Elias Thorne” in a fabricated short story draft’s one thing. “Elias Thorne” showing up again, with a matching partner name and a similar institutional sheen, is another. The repetition gives the machine away a little.
The researchers seem to think that clue could be useful. A weirdly familiar cluster of names might help identify which system produced a piece of AI slop, especially when the surrounding text looks polished enough to fool a quick glance. That’s handy for now, and it could become a mess later. The signal may weaken, once the same names circulate widely enough. Marcus Chen can only serve as a fingerprint for so long before he turns into the AI version of a generic placeholder.
There’s also a cultural wrinkle here. These names don’t come out of nowhere. They sit inside training data shaped by internet habits, publishing conventions and the usual genre clichés. So part of the sameness may come from the model reaching for the most average idea of a scientist, engineer, or fictional lead it can find. The result’s oddly consistent. AI ghost authors end up sounding like they were assembled by committee, except the committee was made of autocomplete and bad instincts.
That consistency matters in academic publishing because it makes fake work feel strangely standardized. One model reaches for the same author names, another picks a different small set, and those habits leave traces. You can almost see the pattern before you know what you’re looking at: a familiar name in a fictional bio, then the same surname in a faux conference abstract, then the same tidy, synthetic colleague in a paper that never should’ve existed, if you work in tech news.
The catch, of course, is that this only helps while the list stays small and somewhat stable. Once these names spill into enough prompts, they stop feeling like a clue and start feeling like background noise. If Marcus Chen shows up everywhere, he becomes less of a fingerprint and more of a default. That’s the part that should make editors uneasy. A detection trick that works today can fade fast when the underlying habit gets copied, recycled and fed back into the next round of model training.
For now, though, the repetition is still useful. It gives reviewers a way to look twice at a draft and ask whether the byline, the abstract and the author list are all coming from the same machine. In a field already dealing with AI-generated papers, that kind of pattern recognition’s better than trusting a polished PDF to mean what it says.
And once those names escape the prompt and land in the wider record, the next problem is no longer spotting them. It’s tracing where they went.
How Those Ghosts Get Real Citations
Once large language models start inventing people, the next step is less funny than it sounds: the made-up names don’t stay inside one weird document. They get pushed into systems that are built to treat paperwork as proof. The preprint PDF lays out the mechanics in plain terms, and the pattern is almost cartoonish in its simplicity. A fake author name turns up on a fake paper, that paper is given a DOI, and suddenly the whole thing wears the costume of a legitimate scholarly record. You can see that chain in the preprint PDF, where the researchers connect the dots between machine-written text and the infrastructure that gives it a home.
Zenodo’s where the loophole gets especially awkward. With a free account, anyone can mint a DOI, which means a file doesn’t need to pass through an editor, an editorial assistant, or even a mildly suspicious graduate student before it starts looking formally published. That matters because a DOI isn’t just a string of characters. It’s the sort of label that gets copied into citation managers, search indexes and reference lists without much ceremony. Plenty of systems will accept it at face value, if the front of the document looks tidy enough. The record can then circulate as if it had earned its place in the literature, even when the journal name on the page’s made up and the authors never existed outside the model that produced them.
A DOI can make a file look archived before anyone checks whether the archive should have taken it.
In this case, the scale wasn’t a handful of odd uploads. It was well over 1,600 ghost-authored records tied to nonexistent journals and invented publication dates. That is not a stray glitch tucked away in a corner of the web. It is a pile of machine-generated paper trails large enough to be noticed by anyone who cares to look, and the repeated names give the game away. The interactive version of the findings walks through that pattern in a more visual way, showing how the same fake identities keep surfacing across records that are supposed to look unrelated. It makes the whole thing feel less like one bad upload and more like a system that keeps stamping the same blank passport over and over. The interactive explainer lays out the trail.
A neat trick in the data is the backdating. Some of the uploads were dated as if they had appeared earlier than they actually did, so the timestamp on the record didn’t match the moment the file went up. That sounds fussy until you remember how often people trust dates without checking them. A backdated record can slip into a bibliography, then into a search result, then into a reference export, where the mismatch becomes harder to spot. By the time anyone notices, the document has already picked up a little sediment from the surrounding setup. It’s a DOI, a date, a title, and maybe even a few citations from someone who assumed the metadata had been checked somewhere upstream.
The spread gets wider once the records leave Zenodo. Scholarly indexes can harvest them automatically, because harvesting’s what they do. They pull metadata, ingest files and map the result into searchable records with only limited judgment about whether the content behind the metadata makes any sense (to put it mildly). Ghost names also appear on ResearchGate, where uploaded papers and profile pages can blur together for casual readers. From there, the trail can drift into Google Scholar and Semantic Scholar, both of which are useful tools but not courtroom clerks. They don’t inspect every author identity by hand, and they certainly don’t know that a given Marcus Chen or Elena Vasquez was conjured out of a model’s statistical habit rather than a lab, a department, or a real inbox.
A separate report tracks how those records keep spreading once they’ve been wrapped in respectable-looking metadata, and the path is annoyingly ordinary. That report shows how the same ghost names can move from repository to profile to index with very little friction. Nobody has to announce the fraud loudly. The systems mostly do the work for it. A file lands, a DOI is assigned, a crawler notices, and a handful of search products do what they’re designed to do: make the record easier to find.
That’s the part that should make editors and librarians wince. The problem’s That large language models can write believable nonsense. It’s that the surrounding machinery can turn believable nonsense into something that looks archived, citable, and discoverable. Once that happens, the ghost stops being a novelty and starts acting like paperwork.
Why Academic Publishing May Need Better Locks
For editors and AI policy watchers, the weirdest part of this whole mess may be the temporary advantage it gives them. If a manuscript keeps turning up with the same invented names, that pattern can act like a fingerprint. A suspicious cluster of Elena Vasquez or Marcus Chen references can tip off a reviewer faster than a tired proofreader on deadline.
That trick may not age well.
So once those names spill across the open web, they stop looking like odd little tells and start looking like ordinary text. They get repeated, quoted, copied and folded into the same web pages, PDFs, slide decks and scraped datasets that future models learn from. At that point, a name that once looked like a giveaway can turn into background noise. The machine learns its own ghost stories. Cute, if you’re writing fiction. Less cute if you’re trying to sort real scholarship from synthetic filler.
A useful detection clue is only useful until the fake clue becomes part of the training set.
That’s the awkward loop here. Today’s tell can become tomorrow’s default. If AI systems keep seeing the same made-up people in enough places, they may treat those names as normal. Then the next batch of generated papers can recycle them with less hesitation, and editors lose one of the easiest ways to spot the slop.
The broader problem reaches beyond fake authorship. Journals have already published AI text in real articles, often without much fanfare and sometimes without anyone noticing until later. Peer review has also been dragged into the mess. Reviewers are now dealing with manuscripts, referee reports and cover letters that may have been drafted or polished by models, which means the process can get noisy in both directions. The people checking the work aren’t always checking human writing anymore.
ArXiv’s responded more aggressively than some traditional venues, handing out year-long bans for AI-generated submissions. That move says a lot about where the pressure’s coming from. If a preprint server has to police machine-written papers this hard, the old assumption that upload systems can sort themselves out starts to look shaky.
The real headache sits underneath all of that. Repositories, aggregators and search indexes often treat machine-made material as if it were ordinary scholarship once it clears a few mechanical hurdles. A file gets a DOI, a record gets indexed, a title gets scraped, and suddenly the thing has the same outward shape as a paper written by a human who spent three miserable months chasing citations. That’s the problem. The machinery rewards format, not truth.
Academic publishing was already built on trust in layers. Authors trust journals to vet papers, and journals trust reviewers to read carefully. Indexes trust repositories to hold real records. The system doesn’t always break loudly, when AI-generated content slips into each layer. It just gets a little fuzzier, a little easier to game, a little more willing to accept a neat-looking file at face value.
For anyone following tech news and AI policy, the lesson’s plain enough. The fight is no longer about whether a model can hallucinate a fake scholar. It can. The tougher question’s whether the academic record can still mean what it says when the pipes carrying it are happy to file machine-made work right next to the real thing. That’s the haunting here, and it’s less to do with bad papers than with whether scholarship still knows who, or what, wrote them.




