The dos and don’ts of processing data in the age of AI
Beyond profits: Giving vs gaining in the digital age
The digital economy has been built on the wonderful promise of equal, fast, and free access to knowledge and information. It has been a long time since then. And instead of the promised equality, we got power imbalances amplified by network effects locking users to the providers of the most popular services. Yet, at first glance, it might appear the users are still not paying anything. But this is where throwing a second glance is worth it. Because they are paying. We all are. We are giving away our data (and a lot of it) to simply access some of the services in question. And all the while their providers are making astronomical profits at the back end of this unbalanced equation. And this applies not only to the present and well-established social media networks but also to the ever-growing number of AI tools and services available out there.
In this article, we will take a full ride on this wild slide and we will do it by considering both the perspective of the users and that of the providers. The current reality, where most service providers rely on dark patterned practices to get their hands on as much data as possible, is but one alternative. Unfortunately, the one we are all living in. To see what some of the other ones might look like, we’ll start off by considering the so-called technology acceptance model. This will help us determine whether the users are actually accepting the rules of the game or if they are just riding the AI hype no matter the consequences. Once we’ve cleared that up, we will turn to what happens in the aftermath with all the (so generously given away) data. Finally, we will consider some practical steps and best practice solutions for those AI developers wanting to do better.
a. Technology acceptance or sleazing your way to consent?
The technology acceptance model is by no means a new concept. Quite to the contrary, this theory has been the subject of public discussion since as early as 1989 when Fred D. Davis introduced it in his Perceived Usefulness, Perceived Ease of Use, and User Acceptance of Information Technology.y on
For example, as opposed to OpenAI actively finding every single person whose personal data is contained in the data sets used to train their models, it could definitely inform their active users that their chats will be used to improve the current and train new models. And here the disclaimer
“As noted above, we may use Content you provide us to improve our Services, for example to train the models that power ChatGPT. See here for instructions on how you can opt out of our use of your Content to train our models.”
does not make the cut for several reasons. And if these payment details have ended up in the models, we can safely assume that the accompanying names, email addresses and other account information are not excluded as well.
Of course, in this described context, the term data altruism can only be used with a significant amount of sarcasm and irony. However, even with providers that aren’t blatantly lying about which data they use and aren’t intentionally elusive with the purposes they use it for, we will again run into problems. Such as, for instance, the complexity of the processing operations that either leads to oversimplification of privacy policies, similar to that of OpenAI, or incomprehensible policies that no one wants to have a look at, let alone read. Both end with the same result, users agreeing to whatever is necessary just to be able to access the service.
Now, one very popular response to such observations happens to be that most of the data we give away is not that important to us, so why should it be to anyone else? Furthermore, who are we to be so interesting to the large conglomerates running the world? However, when this data is used to build nothing less than a business model that relies particularly on those small, irrelevant data points collected from millions across the globe, then the question gets a completely different perspective.
c. Stealing data as a business model?
To examine the business model built on these millions of unimportant consents thrown around every day, we need to examine just how altruistic the users are in giving away their data. Of course, when the users access the service and give away their data in the process, they also get that service in exchange for the data. But that is not the only thing they get. They also get advertisements, or maybe a second-grade service, as the first grade is reserved for subscription users. Not to say that these subscription users aren’t still giving away their Content (with a capital c), as well as (at least in the case of OpenAI) their account information.
And so, while the users are agreeing to just about anything being done with their data in order to use the tool or service, the data they give away is being monetized multiple times to serve them personalized ads and develop new models, which may again follow a freemium model of access. Leaving aside the more philosophical questions, such as why numbers on a bank account are so much more valuable than our life choices and personal preferences, it seems far from logical that the users would be giving away so much to get so little. Especially as the data we are discussing is essential for the service providers, at least if they want to remain competitive.
However, this does not have to be the case. We do not have to wait for new and specific AI regulations to tell us what to do and how to behave. At least when it comes to personal data, the GDPR is pretty clear on how it can be used and for which purposes, no matter the context.
What does the law have to say about it?
As opposed to copyright issues, where the regulations might need to be reinterpreted in light of the new technologies, the same cannot be said for data protection. Data protection has for the better part developed in the digital age and in trying to govern the practices of online service providers. Hence, applying the existing regulations and adhering to existing standards cannot be avoided. Whether and how this can be done is another question.
Here, a couple of things ought to be considered:
1. Consent is an obligation, not a choice.
Not informing the users (before they actually start using the tool) of the fact that their personal data and model inputs will be used for developing new and improving existing models is a major red flag. Basically as red as they get. Consent pop-ups, similar to those for collecting cookie consents are a must, and an easily programmable one.
On the other hand, the idea of pay-or-track (or in the context of AI models pay-or-collect), meaning that the choice is left to the users to decide if they are willing to have their data used by the AI developers, is heavily disputed and can hardly be lawfully implemented. Primarily, because the users still have to have a free choice of accepting or declining tracking, meaning that the price has to be proportionally low (read the service has to be quite cheap) to even justify contending the choice is free. Not to mention, you have to stick with this promise and not collect any subscription users’ data. As Meta has recently switched to this model, and the data protection authorities already received the first complaints because of it,
.
on Medium, where people are continuing the conversation by highlighting and responding to this story.
SOCIAL SHARE CARD GENERATOR