Competition Law And Synthetic Data Providers And Competition Issues .
Competition Law and Synthetic Data Providers: Competition Issues
1. Introduction
Synthetic data refers to artificially generated data designed to reproduce selected statistical, structural, behavioural, or technical characteristics of real-world datasets without necessarily reproducing the original observations. It is increasingly used for AI training, software testing, healthcare research, autonomous systems, financial modelling, cybersecurity, and privacy-preserving analytics.
Synthetic-data providers can therefore occupy several positions in the competitive chain:
Original data holders → Data processors/curators → Synthetic-data generators → AI/model developers → Downstream applications
From a competition-law perspective, the important question is not merely whether synthetic data is “artificial.” The central issue is whether control over underlying data, generation technology, models, APIs, quality benchmarks, distribution channels, or downstream ecosystems enables a provider to restrict competition.
Existing competition jurisprudence concerning data markets provides useful analogies, although there is not yet a large body of reported decisions specifically concerning synthetic-data markets. Modern digital-market regulation increasingly treats data access, portability, interoperability and data-sharing as competition issues. For example, the EU Digital Markets Act contains specific data-access and portability obligations, while the UK CMA has imposed a data-portability conduct requirement on Google.
2. Relevant Competition-Law Framework
Synthetic-data providers may potentially be examined under:
A. Abuse of dominance
A provider with substantial market power could potentially engage in:
- refusal to supply synthetic data;
- discriminatory access;
- excessive pricing;
- tying;
- self-preferencing;
- exclusionary rebates;
- interoperability restrictions;
- API restrictions;
- degradation of data quality for competitors;
- exclusive licensing;
- leveraging from an upstream data market into a downstream AI market.
In the EU, Article 102 TFEU is particularly relevant.
B. Restrictive agreements
Article 101 TFEU and equivalent national provisions may become relevant where synthetic-data providers enter agreements concerning:
- exclusivity;
- territorial restrictions;
- customer allocation;
- data-sharing restrictions;
- collective refusal to supply;
- price coordination;
- standard-setting;
- restrictions on interoperability.
C. Merger control
A merger between:
large data owner + synthetic-data generator + AI developer
could produce substantial vertical and conglomerate effects even where the parties do not appear to compete directly.
Competition authorities may examine whether the combined entity can:
- deny competitors access to important data;
- degrade competitors' synthetic-data inputs;
- bundle synthetic data with AI models;
- foreclose rival model developers;
- use proprietary data advantages to reinforce an existing ecosystem.
D. Essential-facility considerations
In exceptional circumstances, highly differentiated datasets or synthetic-data infrastructure could raise questions analogous to essential-facility doctrine.
However, merely possessing a valuable dataset does not automatically create a legal obligation to supply it. The stringent requirements developed in cases such as Magill, Bronner and IMS Health remain important.
3. Relevant Market Definition
Synthetic data creates unusually difficult market-definition questions.
A competition authority could potentially examine several separate markets.
3.1 Synthetic-data generation market
The relevant product could be:
software and services that generate synthetic datasets from specified source information.
Possible competitors include:
- specialist synthetic-data companies;
- cloud providers;
- AI companies;
- database vendors;
- internal enterprise systems.
3.2 Synthetic data by sector
Synthetic data may not constitute one homogeneous market.
Separate markets could potentially emerge for:
- healthcare synthetic data;
- financial synthetic data;
- autonomous-driving datasets;
- robotics data;
- cybersecurity data;
- geospatial data;
- telecommunications data;
- consumer-behaviour datasets.
A healthcare synthetic-data provider, for example, may possess capabilities that cannot readily be substituted by a general-purpose synthetic-data generator.
3.3 Synthetic-data generation technology
The relevant market could instead concern the technology used to generate synthetic data:
- GAN-based systems;
- diffusion models;
- probabilistic models;
- agent-based simulations;
- digital twins;
- privacy-preserving generation;
- LLM-generated datasets.
3.4 Data-quality dimensions
Competition authorities may also need to consider whether two synthetic datasets are genuinely substitutable.
Important parameters include:
- statistical fidelity;
- diversity;
- bias;
- privacy;
- temporal relevance;
- geographic coverage;
- accuracy;
- representativeness;
- rare-event representation;
- regulatory compliance.
Thus, a dataset with a lower price may not be an effective competitive substitute if it cannot reproduce the relevant rare events or technical characteristics.
4. Major Competition Issues
4.1 Control over underlying data
The first competition concern is often not the synthetic data itself but the source data used to create it.
Suppose Provider A controls:
- millions of medical records;
- proprietary clinical datasets;
- specialised annotations;
- longitudinal patient information;
and uses these to generate synthetic medical datasets.
If competitors cannot obtain equivalent source material, Provider A may acquire an important structural advantage.
The competitive issue becomes:
Does control over the underlying dataset allow the synthetic-data provider to foreclose rival synthetic-data providers or downstream AI developers?
This resembles established competition concerns surrounding data advantages in digital markets.
5. Data Network Effects
Synthetic-data providers can benefit from a feedback loop:
More customers → more feedback → better generation models → higher-quality synthetic data → more customers
This can produce data-driven network effects.
A dominant provider could potentially use these advantages to:
- improve model accuracy;
- reduce generation costs;
- identify customer requirements;
- develop sector-specific synthetic datasets;
- attract additional users.
The resulting advantage may become self-reinforcing.
However, competition law should distinguish between legitimate innovation-based advantages and exclusionary conduct.
6. Synthetic Data as a Competitive Bottleneck
A synthetic-data provider may become a bottleneck if downstream firms depend heavily on its datasets.
For example:
Synthetic medical data → AI diagnostic model → hospital software
If a synthetic-data provider supplies most high-quality training data to competing AI developers, it could potentially influence downstream competition through:
- selective access;
- delayed access;
- discriminatory pricing;
- inferior-quality datasets;
- contractual restrictions.
This makes data access an important competition-law issue.
7. Refusal to Supply
A dominant synthetic-data provider might refuse to provide data to a competitor.
Ordinarily, competition law does not impose a general duty on firms to assist competitors.
The exceptional-facility cases demonstrate that compulsory access requires particularly demanding conditions.
Relevant considerations include:
- whether the input is indispensable;
- whether duplication is realistically possible;
- whether refusal eliminates effective competition;
- whether legitimate business justification exists;
- whether access is technically and commercially feasible.
The doctrine should therefore be applied cautiously to synthetic datasets because synthetic-data alternatives may sometimes be generated independently.
8. Discriminatory Access
A provider could potentially offer:
high-quality synthetic data to its own AI subsidiary
while offering:
lower-quality or delayed data to competing AI developers.
This could become particularly significant where the provider is vertically integrated.
The competition authority would need to examine whether the difference is objectively justified or instead has exclusionary effects.
9. Self-Preferencing
Self-preferencing is another major concern.
Consider:
Synthetic-data platform
↓
Own AI model
↓
Own downstream application
If the platform systematically provides its own AI business with:
- earlier datasets;
- richer datasets;
- more granular datasets;
- higher-frequency updates;
- better API access;
while restricting competitors, the conduct could potentially constitute leveraging or discriminatory treatment.
The concern is similar to competition cases involving dominant platforms favouring their own downstream services.
10. Tying and Bundling
Synthetic-data providers may bundle:
synthetic data + cloud computing + AI model + analytics platform.
Bundling can create efficiencies, but competition concerns may arise if a dominant provider makes access to one product conditional on purchasing another.
For example:
“To obtain our premium synthetic healthcare dataset, you must use our AI development platform.”
Potential concerns include:
- foreclosure of rival cloud services;
- foreclosure of rival AI platforms;
- raising competitors' costs;
- ecosystem lock-in;
- reduced interoperability.
11. Exclusive Dealing
Synthetic-data companies may negotiate exclusive arrangements with major data holders.
For example:
Hospital consortium → exclusive synthetic-data agreement → Provider A
If Provider A obtains exclusive access to a uniquely valuable dataset and competitors cannot realistically replicate it, the arrangement could increase entry barriers.
The analysis would depend on:
- duration;
- coverage;
- market share;
- availability of alternatives;
- importance of the dataset;
- foreclosure percentage;
- efficiencies.
12. Synthetic Data and Interoperability
Interoperability is particularly important because synthetic-data systems may use proprietary:
- APIs;
- schemas;
- metadata;
- ontologies;
- model formats;
- quality metrics.
A provider could potentially make it difficult for customers to transfer their synthetic-data workflows to competitors.
This can produce switching costs.
Modern digital-market rules increasingly address these concerns. The European Commission describes DMA data-access rights as intended to facilitate innovation and contestability, including access to valuable data held by gatekeepers.
13. Switching Costs and Lock-In
Synthetic-data customers may invest heavily in:
- data pipelines;
- validation systems;
- model architectures;
- APIs;
- software integration;
- compliance procedures.
Consequently, switching providers may be expensive.
A provider could potentially exploit this through:
- restrictive contracts;
- proprietary formats;
- termination fees;
- non-portable datasets;
- technical barriers;
- incompatible APIs.
The competitive concern becomes greater where switching costs are artificially increased rather than naturally arising from technological investment.
14. Quality Degradation
A particularly novel issue concerns quality discrimination.
Suppose a dominant provider supplies:
| Customer | Data quality |
|---|---|
| Own AI subsidiary | 99% statistical fidelity |
| Preferred partners | 95% |
| Independent competitors | 80% |
If the difference cannot be objectively justified, this could potentially function as an exclusionary strategy.
Competition authorities would therefore need sophisticated technical evidence concerning:
- accuracy;
- statistical similarity;
- bias;
- representativeness;
- rare-event coverage;
- privacy guarantees.
15. Synthetic Data and Algorithmic Collusion
Synthetic-data providers could potentially create competition concerns if competing firms use common systems to generate pricing or strategic information.
For example:
Provider A + Provider B + Provider C
↓
Common synthetic-data platform
↓
Common pricing predictions
↓
Coordinated commercial behaviour
Competition law could become relevant if the system facilitates:
- exchange of competitively sensitive information;
- coordination;
- signalling;
- algorithmic price alignment.
The fact that the information is synthetic does not automatically eliminate antitrust risk if the underlying system facilitates coordination.
16. Data Accuracy and Competitive Harm
Synthetic data creates an unusual evidentiary problem.
A dataset may be:
- privacy-preserving;
- statistically realistic;
- commercially valuable;
but still systematically biased.
For competition analysis, authorities may therefore need to investigate whether a provider manipulates synthetic data in ways that disadvantage competitors.
Examples include:
- systematically omitting competitor products;
- generating biased market simulations;
- suppressing rare competitor events;
- altering demand patterns;
- manipulating consumer preferences.
17. Merger Control and Synthetic Data
Synthetic-data companies could become strategically important acquisition targets.
A transaction involving:
major data holder + synthetic-data generator
could combine complementary assets that were previously separate.
Similarly:
synthetic-data company + foundation-model developer
could create vertical foreclosure risks.
Authorities may examine:
- input foreclosure;
- customer foreclosure;
- data concentration;
- innovation effects;
- interoperability;
- access to datasets;
- potential competitors;
- nascent competition.
18. Six Important Case Laws
Because reported decisions specifically concerning synthetic-data providers remain limited, the following cases are particularly useful analogies for analysing synthetic-data competition.
Case 1: Magill TV Guide/ITP, BBC and RTÉ
Principle
The European Court developed important principles concerning refusal to license protected information.
A refusal involving an indispensable input can, in exceptional circumstances, constitute abuse of dominance.
Relevance to synthetic data
A synthetic-data provider controlling a unique dataset might face analogous questions where:
- the data is indispensable;
- competitors cannot realistically reproduce it;
- refusal excludes effective competition;
- downstream innovation depends upon access.
However, synthetic-data providers cannot automatically be treated as having an obligation to license their datasets.
19. Case 2: Bronner v Mediaprint
Principle
The Court of Justice applied a stringent test to refusal of access to an infrastructure allegedly indispensable for competition.
The mere fact that access would be advantageous or economically preferable was insufficient.
Synthetic-data relevance
This is important because competitors may argue:
“We need Provider A's synthetic dataset to compete.”
The legal question is more demanding:
Is the dataset genuinely indispensable, or can competitors develop alternative datasets?
If alternative generation methods exist, compulsory access becomes much harder to justify.
20. Case 3: IMS Health GmbH & Co. OHG v NDC Health
Principle
The case concerned access to a data structure and intellectual-property-related refusal to license.
The Court identified stringent conditions for compulsory licensing, including indispensability and elimination of competition.
Synthetic-data relevance
This is particularly relevant to:
- proprietary synthetic-data architectures;
- specialised data formats;
- unique datasets;
- data-generation standards;
- proprietary schemas.
A provider could potentially have intellectual-property rights in aspects of its technology, but intellectual property does not provide unlimited immunity from competition law.
21. Case 4: Microsoft Corp. v Commission
Principle
The European Commission and EU courts considered Microsoft's refusal to provide interoperability information to competing work-group server products.
The case demonstrated how interoperability restrictions can reinforce dominance.
Synthetic-data relevance
The analogy is strong where a dominant synthetic-data platform controls:
- API access;
- data formats;
- interoperability information;
- technical documentation;
- data-transfer mechanisms.
If competitors cannot interoperate with a dominant platform, competition may be weakened.
The broader EU data-policy environment now expressly recognises the importance of interoperability and data portability in digital markets.
22. Case 5: Google Shopping
Principle
The European Commission and EU courts examined Google's treatment of competing comparison-shopping services and its use of dominance in general search to favour its own specialised service.
Synthetic-data relevance
The important analogy is leveraging and self-preferencing.
Suppose:
Google-type platform → synthetic-data service → own AI product
and the platform systematically gives its own synthetic-data business superior access to:
- user-generated information;
- search information;
- computational resources;
- APIs;
- data updates.
The competitive analysis could resemble the broader principle that dominance in one market can be leveraged to disadvantage competitors in another.
23. Case 6: Facebook/WhatsApp Merger Decision
The European Commission's examination of the Facebook/WhatsApp transaction is important for data-driven merger analysis.
Principle
The transaction raised questions concerning the competitive importance of data and the ability of a large platform to combine data resources.
Synthetic-data relevance
A merger between:
large consumer-data platform + synthetic-data provider
could create a significantly expanded data ecosystem.
Potential concerns include:
- combining datasets;
- increasing entry barriers;
- improving AI training capabilities;
- reducing rivals' access to comparable information;
- creating a vertically integrated data ecosystem.
Data concentration therefore can matter even when the parties' products are not straightforward substitutes.
24. Case 7: Bundeskartellamt v Facebook/Meta
Germany's competition authority examined Facebook's combination of data from different services in the context of Facebook's dominant position.
The case is significant because it demonstrated how data collection and data combination can intersect with competition law.
The subsequent EU litigation concerning Meta's data practices further illustrates the interaction between competition law, personal-data rules and market power.
Synthetic-data relevance
A synthetic-data provider might attempt to combine:
- customer data;
- behavioural data;
- third-party datasets;
- proprietary datasets;
to create a synthetic dataset unavailable to competitors.
The competition analysis may therefore consider whether the accumulation and combination of data reinforce market power.
25. Case 8: Aspen Skiing Co. v Aspen Highlands Skiing Corp.
Principle
The U.S. Supreme Court considered a refusal to continue a prior cooperative arrangement where the conduct was found to have exclusionary characteristics.
Synthetic-data relevance
The case is useful when analysing a synthetic-data provider that:
- previously supplied data to competitors;
- abruptly terminates access;
- sacrifices profitable transactions;
- uses the termination to disadvantage a rival.
The case should not be interpreted as establishing a general duty to deal. Its relevance is tied to its specific factual circumstances.
26. Case 9: United States v. Google
The U.S. Google litigation illustrates broader concerns concerning exclusionary conduct in digital ecosystems.
The DOJ's Google litigation remains active in the remedies phase as of 2026.
Synthetic-data relevance
The case provides an important framework for examining whether control over a strategically important digital input or distribution channel can be used to reinforce market power.
For synthetic-data providers, the analogous concern would be:
control over data-generation infrastructure → exclusion of competing AI/data services.
27. Case-Law Principles Compared
| Case | Core principle | Synthetic-data relevance |
|---|---|---|
| Magill | Exceptional refusal to license | Access to indispensable datasets |
| Bronner | Strict indispensability test | Refusal to supply synthetic data |
| IMS Health | Data/IP access and exceptional circumstances | Proprietary datasets and formats |
| Microsoft | Interoperability | APIs and data portability |
| Google Shopping | Leveraging/self-preferencing | Favouring own AI/data products |
| Facebook/WhatsApp | Data concentration in merger analysis | Combining data ecosystems |
| Facebook/Meta | Data combination and dominance | Data accumulation strategies |
| Aspen Skiing | Exceptional refusal to deal | Termination of established data access |
28. Synthetic Data and Essential-Facility Doctrine
An essential-facility claim involving synthetic data would likely require careful analysis of:
1. Indispensability
Can competitors create equivalent synthetic data themselves?
2. Replicability
Can the underlying dataset or generation technology be reproduced?
3. Cost
Is replication technically or economically unrealistic?
4. Competition elimination
Would refusal eliminate effective competition?
5. Justification
Does the provider have legitimate reasons for refusing access?
6. Remedy feasibility
Can access be provided without destroying legitimate incentives to innovate?
This is particularly important because excessive compulsory-access obligations could reduce incentives to invest in expensive data-generation technology.
29. Privacy and Competition Law
Synthetic data is often promoted as a privacy-enhancing technology.
That does not mean privacy and competition are unrelated.
The FTC has expressly recognised the interaction between data control, privacy and competition in digital markets.
Competition authorities may therefore ask:
- Who owns the underlying data?
- How was it collected?
- Can it legally be combined?
- Is the synthetic dataset actually anonymous?
- Can individuals be re-identified?
- Does privacy regulation prevent rivals from obtaining equivalent data?
- Does the dominant firm gain a competitive advantage from privileged access?
The distinction between privacy-preserving data and competition-preserving data access is therefore crucial.
30. Re-identification Risk
A synthetic dataset may still create competitive and regulatory concerns if it permits reconstruction of sensitive information.
For example:
Original dataset → generative model → synthetic dataset
If the generated dataset contains memorised or reconstructable information, the provider may possess a competitive advantage based on access to sensitive underlying information.
This also creates potential asymmetry:
Incumbent: access to real data + synthetic generation
Entrant: synthetic generation only
That asymmetry can become a significant entry barrier.
31. Data Portability
Data portability can reduce switching costs.
The EU already recognises individual data-portability rights under GDPR in specified circumstances, while the DMA contains additional competition-oriented portability and access obligations for designated gatekeepers.
For synthetic-data markets, portability could concern:
- generated datasets;
- metadata;
- schemas;
- prompts;
- generation parameters;
- validation records;
- model outputs;
- API configurations.
The competitive objective would be to prevent technical incompatibility from unnecessarily locking customers into one provider.
32. Recent Regulatory Development: Google Search Data
The developing EU DMA approach is particularly relevant to future synthetic-data competition.
In 2026, the European Commission specified measures requiring Google to share anonymised search information—including ranking, query, click and view data—with qualifying competing search providers under fair, reasonable and non-discriminatory conditions.
The significance for synthetic data is conceptual:
Data generated by a dominant platform can become a competitively significant input for downstream innovation.
This may become increasingly relevant where AI developers use search-derived or platform-derived information to construct synthetic training datasets.
33. Data Portability as a Competitive Remedy
The UK CMA has also imposed a data-portability conduct requirement on Google, requiring tools that facilitate effective portability of search data where authorised by consumers.
For synthetic-data markets, analogous remedies could potentially include:
- standardized APIs;
- machine-readable datasets;
- continuous data transfer;
- interoperability requirements;
- non-discriminatory access;
- transparent technical standards.
34. Competition Risks Specific to Synthetic-Data Providers
The principal risks can be classified as follows:
Structural risks
- high entry costs;
- proprietary datasets;
- computational requirements;
- specialised expertise;
- economies of scale.
Behavioural risks
- refusal to supply;
- discriminatory access;
- tying;
- bundling;
- exclusive dealing;
- self-preferencing;
- API restrictions.
Data risks
- data accumulation;
- data combination;
- data quality manipulation;
- re-identification;
- insufficient portability.
Technological risks
- proprietary formats;
- interoperability restrictions;
- model lock-in;
- technical switching costs.
Merger risks
- vertical integration;
- data concentration;
- elimination of nascent competitors;
- foreclosure of downstream AI firms.
35. Possible Competition-Law Remedies
Authorities could potentially consider:
Structural remedies
- divestiture;
- separation of business units;
- restrictions on vertical integration.
Behavioural remedies
- FRAND access;
- non-discrimination obligations;
- API access;
- interoperability requirements;
- data portability;
- prohibition of tying;
- prohibition of exclusive dealing.
Transparency remedies
- disclosure of access conditions;
- quality standards;
- technical specifications;
- audit mechanisms.
Data remedies
- controlled data sharing;
- anonymisation;
- secure data rooms;
- independent data trustees;
- standardized formats.
Remedies must nevertheless protect legitimate privacy, cybersecurity and intellectual-property interests.
36. Compliance Framework for Synthetic-Data Providers
A synthetic-data provider should establish:
Competition compliance
- Identify relevant markets.
- Assess market power.
- Document objective access criteria.
- Avoid discriminatory treatment.
- Review exclusivity arrangements.
- Monitor tying and bundling.
- Maintain interoperability where commercially appropriate.
Data governance
- Document data provenance.
- Establish lawful collection procedures.
- Test re-identification risk.
- Validate synthetic-data quality.
- Maintain audit trails.
Merger compliance
- Screen acquisitions for data concentration.
- Assess vertical foreclosure.
- Examine potential-competition effects.
- Identify strategic datasets.
37. Hypothetical Example
Assume Company A operates the largest healthcare synthetic-data platform.
It controls:
- 80% of a specialised synthetic medical-data segment;
- proprietary hospital datasets;
- a generation model;
- an API;
- an AI-development platform.
Company A supplies its own AI diagnostic subsidiary with real-time high-quality synthetic data but provides independent AI developers with delayed and lower-quality datasets.
Potential competition concerns could include:
Dominance
↓
Control of important healthcare data
↓
Preferential supply to own AI subsidiary
↓
Disadvantage to competing AI developers
↓
Potential leveraging / discriminatory access
The authority would still need to establish the relevant market, dominance, actual or likely exclusionary effects, and the absence of objective justification.
38. Key Doctrinal Tension
Synthetic-data competition law presents a fundamental tension:
Innovation argument
Strong protection of proprietary datasets and generation technologies may:
- reward investment;
- encourage innovation;
- protect trade secrets;
- promote better synthetic-data technologies.
Competition argument
Excessive control may:
- increase entry barriers;
- entrench incumbents;
- prevent interoperability;
- restrict AI innovation;
- create data bottlenecks.
Competition law must therefore distinguish competition on the merits from strategic exclusion of competitors.
39. Conclusion
Synthetic-data providers are likely to become increasingly important participants in data-driven markets. Their competitive significance may arise not only from the synthetic datasets they sell, but from their control over underlying data, generation models, computational infrastructure, APIs, quality standards and downstream AI ecosystems.
The principal competition-law questions concern:
- market definition;
- data-driven dominance;
- refusal to supply;
- essential-facility arguments;
- self-preferencing;
- discriminatory data access;
- tying and bundling;
- exclusive agreements;
- interoperability and switching costs;
- algorithmic coordination;
- data-driven mergers;
- vertical foreclosure.
The most useful existing jurisprudence comes from Magill, Bronner, IMS Health, Microsoft, Google Shopping, Facebook/WhatsApp, Facebook/Meta, and Aspen Skiing. These cases do not establish a standalone doctrine for synthetic data; rather, they provide established principles that can be adapted to the competitive characteristics of synthetic-data markets.

comments