What happened
It was reported that a team at the University of Zurich ran a field experiment on the subreddit r/changemyview, a forum of about 3.8 million members, without the knowledge of its moderators or its users. According to the researchers’ own extended abstract, since withdrawn, the intervention ran from November 2024 to March 2025 and posted 1,783 comments on 1,061 unique posts from 34 accounts, of which 21 were shadowbanned.
Retraction Watch records the fabricated personas as including a victim of rape, a trauma counsellor specialising in abuse, and a Black man opposed to Black Lives Matter. In one condition the models were given personal attributes of the person they were replying to, inferred from their posting history by another model. The researchers told the moderators that they did not disclose that an AI was used, as this would have rendered the study unfeasible. The subreddit’s rules prohibited undisclosed AI-generated content and bots, and the moderators have stated that they were not contacted beforehand and would have declined.
It was reported that the researchers disclosed the experiment in March 2025, after it had finished, and that the moderators, having asked that the research not be published, published the matter on 26 April 2025. The University’s Ethics Commission replied that the project yields important insights, that the risks are minimal, and that suppressing publication is not proportionate.
The University stated that its Ethics Committee had advised rather than approved and that such recommendations are not legally binding. It was reported that the principal investigator was issued a formal warning and that the researchers decided not to publish. Reddit’s chief legal officer called the conduct deeply wrong on both a moral and legal level and prohibited by Reddit’s user agreement, and Reddit stated that it had banned all accounts associated with the effort.
The withdrawn draft reported persuasion rates of 0.18 personalised, 0.17 generic and 0.09 community aligned, against a human baseline of 0.03 over 478 analysed observations, and 137 deltas, the subreddit’s marker for a changed view. Those rates and that delta count use different denominators. All come from a draft the researchers withdrew and that was not peer reviewed, and are recorded here as claimed rather than established.
What an auditable version would have shown
The nature of what was running was not visible to the community while it ran. The moderators enforce a rule against undisclosed AI content, and 21 of the 34 accounts were shadowbanned during the intervention, which indicates that some accounts met platform visibility restrictions without establishing why those restrictions were applied. The community learned the nature of what had happened when the researchers chose to tell them, four months in and after the last comment had been posted. The figures for the study’s scale in the public record, the 1,783 comments, the 34 accounts, the 1,061 posts, come from the researchers describing their own work.
An auditable version does not depend on that disclosure. A record written for each generated message says which system produced it, under which declared identity, and against which target, and that record exists whether or not the operator later chooses to publish. Where such records are required of research conducted on a live platform, the question of what was deployed and at whom is answered from the record rather than from the operator’s own subsequent account of it. It does not decide whether the study was permissible, and it does not repair the absence of prior disclosure to the people being studied. What it changes is that the extent of an intervention stops being something the public can know only because the people who ran it decided to say.
Where the gap was
It was reported that AI-written comments were presented as the words of people who did not exist, including people presented as survivors of assault and as members of a racial group, and that the community was not told in advance. A PersonaGuard tests each reply against the identity an agent is authorised to present, so a message written in the voice of a trauma survivor or of a member of a group the operator does not belong to is refused rather than posted. An AuthorityGate requires a documented and authorised decision before a campaign of this kind runs, which is the step that makes visible whether a project moved from an ethics opinion that advised additional safeguards into a live intervention without documented evidence that those safeguards had been met, as a recorded event rather than a matter reconstructed afterwards. A ConductRecord preserves, per message, the model, the persona, the inferred attributes used to personalise it and the account it was posted from, which is what allows the scale of a deployment to be established from records. None of these controls decides whether a study is ethical. What they change is that the operator, the institution and the platform can each see what was actually sent.
What governance should have looked like
The University’s position is that its ethics committee gave a non-binding opinion and that responsibility for the conduct and publication of the project rested with the researchers. The moderators’ position is that they were not consulted and would have refused.
Where an organisation deploys an agent that writes as a person to people who have not agreed to take part, best practice would be for the identity the agent may present to be declared and enforced before each message rather than described in a protocol, for the platform’s own rules to be tested at the point of posting, and for every generated message and the identity it carried to be recorded whether or not the work is ever published.
The draft was withdrawn and was not peer reviewed, so this entry treats its persuasion claims as claimed rather than established. The University’s statement and the withdrawn abstract were both read through Retraction Watch reproducing them rather than at their original locations.
Failure Pattern: AI-written text was presented as the speech of people who did not exist, to users who had not been told they were taking part in an experiment, and no record on the platform or in the study distinguished the automated accounts from human ones while it was running.
Governance Principle: an organisation deploying an AI agent that speaks to people should be able to establish, before each message is sent, that the identity the agent presents is one it is authorised to present, and to show afterwards which messages an automated system produced.
The reference implementation of PersonaGuard, AuthorityGate and ConductRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.
Sources
- Ethics committee, University of Zurich, on the AI LLM Reddit changemyview experiment (Retraction Watch, 29 April 2025, reproducing the University’s statement)
- Experiment using AI-generated posts on Reddit draws fire for ethics concerns (Retraction Watch, 28 April 2025)
- Extended abstract, withdrawn and not peer reviewed (University of Zurich researchers, mirrored by Retraction Watch)
- Researchers secretly ran a massive unauthorized AI persuasion experiment on Reddit users (404 Media, 28 April 2025)
- Reddit issuing formal legal demands against researchers who conducted secret AI experiment on users (404 Media, April 2025)
- Swiss researchers admit to secretly running an AI experiment on Reddit users (The Register, 29 April 2025)
- A controversial experiment on Reddit reveals the persuasive powers of AI (NPR, 7 May 2025)
- Fake news AI study: when researchers do more harm than good (SWI swissinfo.ch)
- Researchers secretly experimented on Reddit users with AI-generated comments (Engadget, April 2025), carrying the subreddit membership figure and Reddit’s statement that it had banned the accounts
- Report carrying the comment, account, shadowban and delta counts (Decrypt, April 2025)
- Interview with a r/changemyview moderator on the disclosure and its timing (Community Signal)