The community.lexicon.preference.ai lexicon

First, I think there are some issues with the proposal from a legal perspective in that the language is very mixed with regards to whether it is just a suggestion about what a user prefers or if it’s trying actually create some sort of legally enforceable set of constraints or allowing users to give affirmative permission for their data to be used if a law were to require that.

Saying they are “user preferences” to me implies that it’s simply what the user prefers and is not even trying to be legally binding in any way.

However, the phrasing “These schemas allow users to declare how their public data may be used by external consumers” seems much stronger and that it implies that it is legally binding.

I am not a lawyer, so I can’t really offer any advice on whether this sort of thing could actually place enforceable legal obligations on external data consumers (and if it can, it almost certainly depends on the country). However, if that is the goal, it seems pretty important to have a lawyer review this and add what is likely to be many pages of additional legal statements.

For example, when it comes to defaults or situations where a user has indication some preferences, but not all, it is possible that there is a meaningful legal distinction in some jurisdictions. Perhaps users explicitly opting out confers additional legal penalties when violated or legal use of the data requires users to take some affirmative action to consent beyond just accepting the defaults. It’s possible the inability to distinguish between those could create problems.

If this is just a sort of for fun set of suggestions that no one is even pretending might have any legal implications for the use of the data, then I think the language should make that more explicitly clear so that users understand this isn’t the basis for them to sue anyone.

Secondly, from the perspective of an AI builder who is trying in good faith to comply with these preferences, there is a lack of clarity that makes it extremely difficult. Most terms of service call for granting a license to the data in perpetuity to make things easy. That probably doesn’t work for this, but not having any listed time frame potentially makes it impossible to comply with. Most notably what happens when a user changes their preferences to opt out of something that they had previously opted into.

For example, being clear that a current user preference for their data to being used to train an AI model confers a license to train the model for 3 years from the date of access and then the right to use the model in perpetuity after it has been trained. Shorter times obviously react more quickly to changing user preferences, but if they are too short it makes certain applications, such as publishing research with an accompanying data set for someone else to replicate the results much tougher.

There are plenty of bad actors who will ignore the user preferences no matter what, so if there is any worthwhile goal of this, it should be to allow someone to build something with AI that they can affirmatively say is done completely with the consent of the users who generated the data. Right now, I think it still falls short of that.

3 Likes