🎉 Get 20% OFF the first C-DPO cohort! Code:
20KSACDPO
ChatGPT Summarize Button

Problem with Building Compliant AI under the KSA PDPL

Training a large language model almost certainly violates data protection law. Here is how the world's leading regulators are trying to resolve it — and what it means for Saudi Arabia.

Table of Contents

Introduction

Vision 2030 in the Kingdom of Saudi Arabia has Artificial Intelligence as one of its main themes. The Kingdom is investing money, people and time into trying to become a regional leader.

The Kingdom is not just trying to become a user of AI – its ambition it to become a builder of AI products.

But any AI built must now also align with the Personal Data Protection Law (PDPL) – and this now causes the same challenge in the Kingdom as the GDPR causes in Europe.

The critical stumbling block come when finding data to train an AI model.

Sourcing data to train LLMs

In order to train most LLMs, you need to feed them data. This is commonly personal data if the product is to interact with humans. There are definitely some LLMs that do not need personal data – for example, if you are training a model to understand earthquakes, you are likely to be feeding in tectonic patterns. One of our clients has AI-enabled cameras to assess if a driver is wearing a seatbelt – and amazingly their training data was taken from game Grand Theft Auto.

But most things need personal data – so, where do AI companies get this personal data to feed their LLMs?

Scraping data

One starting point is scraping data from the web. This is literally deploying software to go and harvest data from websites such as Wikipedia, newspapers, blogs – absolutely everything that is published.

Is this a problem? A lot of people will say “It’s public, so it’s allowed” or “its public, so it’s not personal data”. Both suppositions are wrong.

The fact that something is public does not mean it is not covered by terms and conditions or copyright or intellectual property laws.

The fact that something is public does not stop it being personal data. Personal data does not mean private – it just means data about a person. In Arabic, mixing up شخصي (personal) and خاص (private) is so common that many assume personal data must be secret.

Under the PDPL, however, any information that identifies you—whether hidden on a secure server or openly shown on your LinkedIn profile for example—is personal data.

Buying data

It is possible to buy the right data sets, develop synthetic data or get permission to use public data under license. For example, Wikipedia publishes its data under a reuse license. But take BBC News’ terms and conditions. Section 8A says that you cannot take anything from the BBC website to train an AI or do computer analysis.

So, unless you are VERY strict about your data sources, a lot of the personal data you are using to train your LLM is likely to be obtained improperly.

Why does this matter? What has training an AI model got to do with data protection?

Like the GDPR, the PDPL required you to have a lawful basis for processing any personal data. The list of lawful basis bases if given and we can start considering them:

Consent – it is clearly not possible to contact all the people whose personal data you have scraped and ask for consent (imagine trying to contact all the people referred to on Wikipedia – does anyone think that you can email the US President to ask for consent?)

Contract – you have no contract with the data subjects – there is no agreement where either one of you is providing a service or good for payment. So this is not an option.

Legal obligation – there is clearly no legal obligation in a private company developing an LLM – although if you are a government entity, this may be worth exploring further

Interests of the data subject – there is clearly no interest of the data subject – you are doing this for the benefit of the company developing the LLM, not to benefit the specific individual whose data you are processing.

Legitimate interests – this is where you balance the interest of the data subject and the interests of the data controller (the company) – processing the personal data “Brad Pitt is an actor” is extremely unlikely to have any harmful effects on the data subject

Legitimate interest looks really promising – it looks like we have found our legal basis! Problem solved!

BUT… let’s read Article 16 of the Regulations to look into using legitimate interests in more detail:

16: 1- Except in cases where the Controller is a Public Entity, the Controller may process Personal Data to achieve a Legitimate Interest provided that the following conditions are met:

  1. a) Purpose shall not violate any of the laws in the Kingdom.
  2. b) A balance between the rights and interests of the Data Subject and the Legitimate Interest of the Controller, so that the interests of the Controller do not affect the rights and interests of the Data Subject.
  3. c) Processing shall not include Sensitive Data.
  4. d) Processing shall be within the reasonable expectations of the Data Subject

 

The problems with Legitimate Interests

You can overcome 16.1.b&d by doing the balancing test and attempting to be transparent (there are issues with the latter, but we can pretend).

The main issues are a) not violating any laws and c) not including sensitive data

Not Violating any Laws in the Kingdom

As we noted above, it is likely that data collection may breach copyright law. You would need a lawyer to advise on how the terms and conditions of the BBC website interact with KSA copyright law or intellectual property law. If there is any violation of a law, then you cannot use LIs.

Not include sensitive data

Brad Pitt is an actor – but if you read the Wikipedia page, it mentions that he has undiagnosed prosopagnosia (face blindness). So, if your LLM is scraping Wikipedia, you are likely to pick up health data – which is sensitive personal data. You will also likely pick up the fact that Sheikh Abdul-Aziz ibn Abdullah Al Al-Sheikh is the Grand Mufti of KSA. From this, we conclude that he is a Muslim – however uncontroversial or public this is, it remains sensitive personal data. If you scrape or collect or process sensitive personal data, then you cannot use LIs.

It feels like there is nowhere to go now – but even in the EU, with the mighty GDPR, they are developing LLMS – how is this happening?

The European Solution

Data protection is young in Saudi and all of the other GCC countries. It takes the EU, which is 27 countries and has the second largest economy in the world years to make decisions on critical and technical issues around data protection. So, we cannot expect regulators such as SDAIA to have sorted out these issues for us yet.

The 27 regulators comprising the European Data Protection Board published an interesting opinion in December 2024 on resolving data protection issues when building and using AI. This is the best document to turn to – and this is what contains a way forward.

The paper presents various scenarios – the most useful one is where an LLM has been trained by unlawful processing of personal data.

The paper separates the development and deployment stages of the LLM.

…a controller unlawfully processes personal data to develop the AI model, then ensures that it is anonymised, before the same or another controller initiates another processing of personal data in the context of the deployment… if it can be demonstrated that the subsequent operation of the AI model does not entail the processing of personal data, the EDPB considers that the GDPR would not apply. Hence, the unlawfulness of the initial processing should not impact the subsequent operation of the model. 

Analysis of the solution

While not recommending that LLMs are developed unlawfully, the EDPB recommends that if, after development, the data is anonymized, then it may be possible to deploy the LLM in a compliant way.

Anonymization is a difficult step – as many people have commented, this is not just about removing names. SDAIA has given some guidance on anonymization. But this document is not yet sufficiently developed. More guidance is available from the UK regulator, the ICO, which may prove more helpful.

You have still breached data protection laws in development, but you have washed and purified your LLM for further use.

The EDPB further goes on to say:

…when controllers subsequently process personal data collected during the deployment phase, after the model has been anonymised, the GDPR would apply in relation to these processing operations. In these cases, the Opinion considers that, as regards the GDPR, the lawfulness of the processing carried out in the deployment phase should not be impacted by the unlawfulness of the initial processing. 

Conclusion

When your now compliant LLM is in deployed and in action, you still need to comply with data protection laws. At this stage, you may need consent for processing sensitive personal data. You may be able to use a legitimate interest or contract. But you have a way forward.

This solution may strike you as elegant and helpful, a way of supporting tech development.

This solution may strike you as depressing, showing how the institutions of the European Union have been influenced by large tech lobbyists.

Depressing or elegant – while the GCC has not developed its own mechanisms for overcoming compliance issues for LLMs, it is likely that all regulators will borrow some elements from the EDPB opinion.