


An AI company is only as sound as its right to use the data behind it. Training data, customer data, third-party content and model licences all carry legal obligations, and getting them wrong can lead to lawsuits, regulatory action, lost customers or a failed acquisition.
The risk is no longer theoretical. High-profile copyright lawsuits, including The New York Times' 2023 case against OpenAI and Microsoft, have put data provenance at the centre of AI investing, and new rules such as the EU AI Act, which entered into force in 2024, add transparency and governance obligations. This guide explains the main data and licensing risks and how investors can diligence them.
Where did the data used to train or fine-tune models come from? Scraped web content, licensed datasets, public-domain material and synthetic data all carry different rights and risks.
Many AI products learn from the data customers put into them. Whether the company can use that data to improve its models, and for which customers, depends on contracts and privacy law.
Privacy laws such as the EU's GDPR and California's CCPA set rules on collecting, processing and retaining personal information, including in AI systems.
Commercial model providers and open-source models come with terms of use that can restrict commercial use, certain applications or competition with the provider.
Who owns what the AI generates, and what happens if an output infringes someone else's rights or causes harm?
Acquirers diligence data rights closely, because they inherit the risk. Unclear data provenance can reduce valuation, add escrow or special indemnities, or stop a deal entirely. See our guide to selling your startup: escrow and reps and warranties.
Publicly accessible does not always mean free to use. Copyright, website terms and privacy law may still apply, and the law is still developing through litigation.
It depends on the jurisdiction, the model provider's terms and the company's customer contracts. Copyright protection for purely AI-generated content is limited in several jurisdictions.
It can, if their AI systems are placed on the EU market or their outputs are used in the EU.
Some offer indemnities under certain conditions. Read the terms carefully, as coverage and exclusions vary.
Data rights have become a core part of AI due diligence. Investors who map data sources, licences, customer terms and regulatory exposure can avoid inheriting hidden legal risk; founders who document their rights from the start make fundraising and exit far smoother. For related AI topics, see evaluating an AI startup's moat.
Global Capital Network connects AI founders with investors through our events and investor network. Get in touch to learn more.
This article is general information, not legal advice. Data, copyright and AI laws are evolving quickly; take specialist legal advice.



.png)




