Subscribe to ConsentBit Newsletter

Thank you!
Your submission has been received!
Oops! Something went wrong while submitting the form.
Consent

Consent for AI: Do You Need Permission to Train AI on Customer Data?

By the Editorial Team
03
August
2026
02
September
2026

TL;DR:

If you want to use personal data, including data you collected on your website, to train or fine-tune an AI model, you need a lawful basis under GDPR. The two realistic options are consent and legitimate interest. Consent is the cleanest but rarely scales. Legitimate interest is the common practical basis, but the bar is high and regulators are scrutinizing it. As of 2026 the law here is still evolving. Whatever basis you choose, transparency, a genuine opt-out or right to object, safeguards, and documentation are non-negotiable.

Important: This article is general information, not legal advice. The rules for AI and personal data are actively evolving, and they vary by region. Consult qualified data protection counsel before using personal data to train AI.

Everyone is racing to use their data for AI. Support transcripts, CRM records, user behavior, product data, all of it looks like fuel for a model or a smarter product.

In the rush, one question tends to get skipped: are we actually allowed to do this? The honest answer is nuanced, and in 2026 it is still being written. Here is what is reasonably clear, what is not, and what to do in the meantime.

The core rule: personal data plus AI training needs a lawful basis

Under the GDPR (and the UK GDPR), you cannot process personal data without a lawful basis. Training an AI model on personal data is processing. So the first question is not "which AI tool," it is "what is our lawful basis for using this personal data this way?"

For AI training, two lawful bases come up in practice: consent and legitimate interest. Understanding the tradeoff between them is most of the decision.

Option 1: Consent

Consent means asking people for clear, specific, informed, opt-in permission to use their data to train AI.

  • The upside: it is the cleanest basis. If someone genuinely agreed, your footing is strong.
  • The problem: consent rarely scales. If you are training on large datasets, getting specific consent from every individual is often impractical or impossible. Consent also carries a right of withdrawal, and "un-training" a model once someone's data is baked into it is extremely difficult to operationalize.

Consent works best for narrow, first-party, clearly bounded uses where you can actually ask and honor the answer.

Option 2: Legitimate interest

Legitimate interest lets you process data for a genuine business interest, without opt-in consent, provided that interest is not overridden by the individual's rights, and provided people can object.

This is the basis many organizations lean on for AI training, and in December 2024 the European Data Protection Board (EDPB) confirmed, in Opinion 28/2024, that legitimate interest can potentially be a lawful basis for developing and deploying AI models. But it set a high bar. You have to pass a three-part assessment:

  • Legitimate interest: a real, specific, documented interest (not speculative).
  • Necessity: the processing must be genuinely necessary. If you could achieve the goal without personal data, the test fails.
  • Balancing: your interest must not override people's rights, judged partly on their reasonable expectations, the context the data was collected in, and the safeguards you apply.

Crucially, legitimate interest comes with a right to object. People must be able to say no, and you must honor it.

What the regulators have actually said

A few reference points, because this is where accuracy matters:

  • The EDPB's Opinion 28/2024 is the key statement. It leaves the door open to legitimate interest but stresses that the bar is high, that safeguards (transparency, pseudonymization, measures to prevent a model from regurgitating personal data) matter to the balancing test, and that training a model on unlawfully processed data can taint its later deployment.
  • In 2026, the EDPB has continued to issue guidance, including draft guidelines touching on web scraping for generative AI and on anonymisation, signs that scrutiny is increasing, not decreasing.
  • Enforcement and interpretation are left substantially to individual data protection authorities, which is part of why outcomes are still uncertain.

What is still unsettled (the honest part)

It would be misleading to present this as settled. It is not.

  • There is no single, binding rule that neatly answers "consent or legitimate interest for AI training." The EDPB's opinion is influential but non-binding.
  • Regulators retain wide discretion, so the same practice can be viewed differently across jurisdictions.
  • The framework is actively evolving, with ongoing debate and proposed clarifications at the EU level.
  • Outside the EU, the picture is different again. The US has no single federal rule, a patchwork of state privacy laws, and emerging AI-specific regulation, so "it depends where your users are" is a real answer.

Anyone claiming there is a simple, universal yes-or-no here is overstating what the law currently provides.

What every business should do now

Regardless of how the finer points settle, this is the defensible path:

  • Identify and document your lawful basis before you train on personal data. If you rely on legitimate interest, complete and keep a legitimate interest assessment.
  • Be transparent. Tell people, in your privacy policy and at the point of collection, if their data may be used to train or improve AI. Surprising people is where trust and compliance both break.
  • Offer a genuine opt-out or honor the right to object, and make it real, not buried.
  • Apply safeguards: minimize the data, pseudonymize where you can, and use technical measures to prevent a model from exposing personal data.
  • Keep records of your decisions and assessments. Documentation is often what defends you under scrutiny.
  • Get legal advice for anything at scale, anything involving special-category (sensitive) data, or anything crossing borders.

Where consent management fits

Much of the work above is about transparency, permission, and honoring people's choices, and that is exactly what a consent management platform handles. ConsentBit helps you capture and record consent and preferences, present clear choices at the point of collection, and honor opt-outs and objections consistently across your site.

To be clear about the boundary: a consent platform manages the consent and transparency layer. It does not, on its own, make an AI training program lawful. That still requires choosing and documenting the right lawful basis, applying safeguards, and, for anything significant, legal review. A consent platform supports that work. It does not replace it.

Frequently asked questions

1. Do I need consent to train AI on customer data?

Not necessarily consent specifically, but you do need a lawful basis under GDPR. The two practical options are consent (clean but hard to scale) and legitimate interest (common, but it requires passing a high-bar three-part assessment and offering a right to object). The right choice depends on your situation, and the law is still evolving.

2. Is legitimate interest a valid basis for AI training?

Potentially. The EDPB confirmed in Opinion 28/2024 that legitimate interest can be a lawful basis for AI models, but only if you pass the legitimate interest, necessity, and balancing tests and provide a right to object. The bar is high and regulators are scrutinizing it.

3. Can I use data people gave me for one purpose to train AI?

Not automatically. Using data for a new purpose (AI training) it was not collected for raises questions about lawful basis, transparency, and reasonable expectations. You generally need to establish a basis for the new use and be transparent about it. Get legal advice.

4. What about US customer data?

The US has no single federal rule. A patchwork of state privacy laws and emerging AI regulation applies, so requirements depend on where your users are. Treat it as its own analysis, separate from GDPR.

5. Does a consent tool make my AI training compliant?

No. A consent management platform handles capturing consent, presenting choices, and honoring opt-outs, which is an important part of the picture. It does not by itself make an AI training program lawful. That requires a documented lawful basis, safeguards, and legal review.

Want the consent and transparency layer handled properly?

If you are collecting data that might feed AI, the transparency and opt-out side has to be solid. ConsentBit helps you capture consent, present clear choices, and honor objections across your site, so that part is covered while your team and counsel handle the lawful-basis decisions.