What happened
It was reported that on 18 September 2024 LinkedIn, owned by Microsoft, updated its privacy policy to allow members’ data to be used to train its generative AI features, and that the setting enabling this had been switched on for users by default rather than requiring them to opt in. Digital-rights advocates and users criticised the move, arguing that people had been enrolled into AI training without being clearly asked. The company said the data would improve its AI features and that members could opt out through a setting.
The United Kingdom’s Information Commissioner’s Office raised concerns about LinkedIn’s approach to training generative AI on UK users’ data. Within days LinkedIn said it had stopped training its generative AI models on data from members in the United Kingdom, the European Economic Area and Switzerland, and would not offer the setting to members in those regions until further notice. In other regions, including the United States, the default remained in place unless a member turned it off.
What an auditable version would have shown
Whether data can be used to train AI turns on a clear legal basis, most often consent, and that basis is exactly the kind of thing that should be recorded rather than assumed. An auditable version keeps, for each member, the basis on which their data is being used: whether they were asked, what they were told, whether they agreed or were defaulted in, and when they changed the setting. With that record, the question the regulator asked, were people properly asked before their data was used, has an answer on file, and a member can see and prove what they did or did not agree to.
Where the gap was
Members’ data was put to a new use, training AI, before many of them knew it was happening, and the default did the deciding for them. An EgressGate governs personal data leaving its original purpose for a new one such as model training, and requires a valid, recorded basis before it can, so data is not repurposed by a quiet default. A ConductRecord keeps the account of each member’s consent state over time, so the platform can show, and the member can check, what was agreed and when. The gap was not that LinkedIn built AI features, but that a consequential use of personal data was switched on without a clear, recorded choice by the people it belonged to.
What governance should have looked like
When a company changes what it does with people’s data, the honest default is to ask, not to assume, and to keep a record of what each person actually agreed to. Opt-out by quiet policy update puts the burden on the user to notice and object, which regulators have repeatedly found is not valid consent under GDPR-style rules. The lesson is that the legitimacy of AI training rests on a recorded basis for using the data, and “you were opted in unless you found the setting” is not that.
Failure Pattern: personal data was repurposed to train AI by a default setting, without a clear, recorded consent from the people it belonged to.
Governance Principle: before personal data is used to train AI, there should be a valid, recorded basis for that use, and each person should be able to see and prove what they agreed to.
The reference implementation of EgressGate and ConductRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.
Sources
- LinkedIn has stopped grabbing UK users’ data for AI (TechCrunch)
- LinkedIn halts AI data processing in the UK amid privacy concerns raised by the ICO (The Hacker News)
- LinkedIn backtracks on controversial AI training rules after user backlash (IT Pro)