Competition Law And Synthetic Data Providers And Competition Issues .

Competition Law and Synthetic Data Providers: Competition Issues

1. Introduction

Synthetic data refers to artificially generated data designed to reproduce selected statistical, structural, behavioural, or technical characteristics of real-world datasets without necessarily reproducing the original observations. It is increasingly used for AI training, software testing, healthcare research, autonomous systems, financial modelling, cybersecurity, and privacy-preserving analytics.

Synthetic-data providers can therefore occupy several positions in the competitive chain:

Original data holders → Data processors/curators → Synthetic-data generators → AI/model developers → Downstream applications

From a competition-law perspective, the important question is not merely whether synthetic data is “artificial.” The central issue is whether control over underlying data, generation technology, models, APIs, quality benchmarks, distribution channels, or downstream ecosystems enables a provider to restrict competition.

Existing competition jurisprudence concerning data markets provides useful analogies, although there is not yet a large body of reported decisions specifically concerning synthetic-data markets. Modern digital-market regulation increasingly treats data access, portability, interoperability and data-sharing as competition issues. For example, the EU Digital Markets Act contains specific data-access and portability obligations, while the UK CMA has imposed a data-portability conduct requirement on Google.

2. Relevant Competition-Law Framework

Synthetic-data providers may potentially be examined under:

A. Abuse of dominance

A provider with substantial market power could potentially engage in:

  • refusal to supply synthetic data;
  • discriminatory access;
  • excessive pricing;
  • tying;
  • self-preferencing;
  • exclusionary rebates;
  • interoperability restrictions;
  • API restrictions;
  • degradation of data quality for competitors;
  • exclusive licensing;
  • leveraging from an upstream data market into a downstream AI market.

In the EU, Article 102 TFEU is particularly relevant.

B. Restrictive agreements

Article 101 TFEU and equivalent national provisions may become relevant where synthetic-data providers enter agreements concerning:

  • exclusivity;
  • territorial restrictions;
  • customer allocation;
  • data-sharing restrictions;
  • collective refusal to supply;
  • price coordination;
  • standard-setting;
  • restrictions on interoperability.

C. Merger control

A merger between:

large data owner + synthetic-data generator + AI developer

could produce substantial vertical and conglomerate effects even where the parties do not appear to compete directly.

Competition authorities may examine whether the combined entity can:

  1. deny competitors access to important data;
  2. degrade competitors' synthetic-data inputs;
  3. bundle synthetic data with AI models;
  4. foreclose rival model developers;
  5. use proprietary data advantages to reinforce an existing ecosystem.

D. Essential-facility considerations

In exceptional circumstances, highly differentiated datasets or synthetic-data infrastructure could raise questions analogous to essential-facility doctrine.

However, merely possessing a valuable dataset does not automatically create a legal obligation to supply it. The stringent requirements developed in cases such as Magill, Bronner and IMS Health remain important.

3. Relevant Market Definition

Synthetic data creates unusually difficult market-definition questions.

A competition authority could potentially examine several separate markets.

3.1 Synthetic-data generation market

The relevant product could be:

software and services that generate synthetic datasets from specified source information.

Possible competitors include:

  • specialist synthetic-data companies;
  • cloud providers;
  • AI companies;
  • database vendors;
  • internal enterprise systems.

3.2 Synthetic data by sector

Synthetic data may not constitute one homogeneous market.

Separate markets could potentially emerge for:

  • healthcare synthetic data;
  • financial synthetic data;
  • autonomous-driving datasets;
  • robotics data;
  • cybersecurity data;
  • geospatial data;
  • telecommunications data;
  • consumer-behaviour datasets.

A healthcare synthetic-data provider, for example, may possess capabilities that cannot readily be substituted by a general-purpose synthetic-data generator.

3.3 Synthetic-data generation technology

The relevant market could instead concern the technology used to generate synthetic data:

  • GAN-based systems;
  • diffusion models;
  • probabilistic models;
  • agent-based simulations;
  • digital twins;
  • privacy-preserving generation;
  • LLM-generated datasets.

3.4 Data-quality dimensions

Competition authorities may also need to consider whether two synthetic datasets are genuinely substitutable.

Important parameters include:

  • statistical fidelity;
  • diversity;
  • bias;
  • privacy;
  • temporal relevance;
  • geographic coverage;
  • accuracy;
  • representativeness;
  • rare-event representation;
  • regulatory compliance.

Thus, a dataset with a lower price may not be an effective competitive substitute if it cannot reproduce the relevant rare events or technical characteristics.

4. Major Competition Issues

4.1 Control over underlying data

The first competition concern is often not the synthetic data itself but the source data used to create it.

Suppose Provider A controls:

  • millions of medical records;
  • proprietary clinical datasets;
  • specialised annotations;
  • longitudinal patient information;

and uses these to generate synthetic medical datasets.

If competitors cannot obtain equivalent source material, Provider A may acquire an important structural advantage.

The competitive issue becomes:

Does control over the underlying dataset allow the synthetic-data provider to foreclose rival synthetic-data providers or downstream AI developers?

This resembles established competition concerns surrounding data advantages in digital markets.

5. Data Network Effects

Synthetic-data providers can benefit from a feedback loop:

More customers → more feedback → better generation models → higher-quality synthetic data → more customers

This can produce data-driven network effects.

A dominant provider could potentially use these advantages to:

  • improve model accuracy;
  • reduce generation costs;
  • identify customer requirements;
  • develop sector-specific synthetic datasets;
  • attract additional users.

The resulting advantage may become self-reinforcing.

However, competition law should distinguish between legitimate innovation-based advantages and exclusionary conduct.

6. Synthetic Data as a Competitive Bottleneck

A synthetic-data provider may become a bottleneck if downstream firms depend heavily on its datasets.

For example:

Synthetic medical data → AI diagnostic model → hospital software

If a synthetic-data provider supplies most high-quality training data to competing AI developers, it could potentially influence downstream competition through:

  • selective access;
  • delayed access;
  • discriminatory pricing;
  • inferior-quality datasets;
  • contractual restrictions.

This makes data access an important competition-law issue.

7. Refusal to Supply

A dominant synthetic-data provider might refuse to provide data to a competitor.

Ordinarily, competition law does not impose a general duty on firms to assist competitors.

The exceptional-facility cases demonstrate that compulsory access requires particularly demanding conditions.

Relevant considerations include:

  1. whether the input is indispensable;
  2. whether duplication is realistically possible;
  3. whether refusal eliminates effective competition;
  4. whether legitimate business justification exists;
  5. whether access is technically and commercially feasible.

The doctrine should therefore be applied cautiously to synthetic datasets because synthetic-data alternatives may sometimes be generated independently.

8. Discriminatory Access

A provider could potentially offer:

high-quality synthetic data to its own AI subsidiary

while offering:

lower-quality or delayed data to competing AI developers.

This could become particularly significant where the provider is vertically integrated.

The competition authority would need to examine whether the difference is objectively justified or instead has exclusionary effects.

9. Self-Preferencing

Self-preferencing is another major concern.

Consider:

Synthetic-data platform

↓

Own AI model

↓

Own downstream application

If the platform systematically provides its own AI business with:

  • earlier datasets;
  • richer datasets;
  • more granular datasets;
  • higher-frequency updates;
  • better API access;

while restricting competitors, the conduct could potentially constitute leveraging or discriminatory treatment.

The concern is similar to competition cases involving dominant platforms favouring their own downstream services.

10. Tying and Bundling

Synthetic-data providers may bundle:

synthetic data + cloud computing + AI model + analytics platform.

Bundling can create efficiencies, but competition concerns may arise if a dominant provider makes access to one product conditional on purchasing another.

For example:

“To obtain our premium synthetic healthcare dataset, you must use our AI development platform.”

Potential concerns include:

  • foreclosure of rival cloud services;
  • foreclosure of rival AI platforms;
  • raising competitors' costs;
  • ecosystem lock-in;
  • reduced interoperability.

11. Exclusive Dealing

Synthetic-data companies may negotiate exclusive arrangements with major data holders.

For example:

Hospital consortium → exclusive synthetic-data agreement → Provider A

If Provider A obtains exclusive access to a uniquely valuable dataset and competitors cannot realistically replicate it, the arrangement could increase entry barriers.

The analysis would depend on:

  • duration;
  • coverage;
  • market share;
  • availability of alternatives;
  • importance of the dataset;
  • foreclosure percentage;
  • efficiencies.

12. Synthetic Data and Interoperability

Interoperability is particularly important because synthetic-data systems may use proprietary:

  • APIs;
  • schemas;
  • metadata;
  • ontologies;
  • model formats;
  • quality metrics.

A provider could potentially make it difficult for customers to transfer their synthetic-data workflows to competitors.

This can produce switching costs.

Modern digital-market rules increasingly address these concerns. The European Commission describes DMA data-access rights as intended to facilitate innovation and contestability, including access to valuable data held by gatekeepers.

13. Switching Costs and Lock-In

Synthetic-data customers may invest heavily in:

  • data pipelines;
  • validation systems;
  • model architectures;
  • APIs;
  • software integration;
  • compliance procedures.

Consequently, switching providers may be expensive.

A provider could potentially exploit this through:

  • restrictive contracts;
  • proprietary formats;
  • termination fees;
  • non-portable datasets;
  • technical barriers;
  • incompatible APIs.

The competitive concern becomes greater where switching costs are artificially increased rather than naturally arising from technological investment.

14. Quality Degradation

A particularly novel issue concerns quality discrimination.

Suppose a dominant provider supplies:

CustomerData quality
Own AI subsidiary99% statistical fidelity
Preferred partners95%
Independent competitors80%

If the difference cannot be objectively justified, this could potentially function as an exclusionary strategy.

Competition authorities would therefore need sophisticated technical evidence concerning:

  • accuracy;
  • statistical similarity;
  • bias;
  • representativeness;
  • rare-event coverage;
  • privacy guarantees.

15. Synthetic Data and Algorithmic Collusion

Synthetic-data providers could potentially create competition concerns if competing firms use common systems to generate pricing or strategic information.

For example:

Provider A + Provider B + Provider C

↓

Common synthetic-data platform

↓

Common pricing predictions

↓

Coordinated commercial behaviour

Competition law could become relevant if the system facilitates:

  • exchange of competitively sensitive information;
  • coordination;
  • signalling;
  • algorithmic price alignment.

The fact that the information is synthetic does not automatically eliminate antitrust risk if the underlying system facilitates coordination.

16. Data Accuracy and Competitive Harm

Synthetic data creates an unusual evidentiary problem.

A dataset may be:

  • privacy-preserving;
  • statistically realistic;
  • commercially valuable;

but still systematically biased.

For competition analysis, authorities may therefore need to investigate whether a provider manipulates synthetic data in ways that disadvantage competitors.

Examples include:

  • systematically omitting competitor products;
  • generating biased market simulations;
  • suppressing rare competitor events;
  • altering demand patterns;
  • manipulating consumer preferences.

17. Merger Control and Synthetic Data

Synthetic-data companies could become strategically important acquisition targets.

A transaction involving:

major data holder + synthetic-data generator

could combine complementary assets that were previously separate.

Similarly:

synthetic-data company + foundation-model developer

could create vertical foreclosure risks.

Authorities may examine:

  • input foreclosure;
  • customer foreclosure;
  • data concentration;
  • innovation effects;
  • interoperability;
  • access to datasets;
  • potential competitors;
  • nascent competition.

18. Six Important Case Laws

Because reported decisions specifically concerning synthetic-data providers remain limited, the following cases are particularly useful analogies for analysing synthetic-data competition.

Case 1: Magill TV Guide/ITP, BBC and RTÉ

Principle

The European Court developed important principles concerning refusal to license protected information.

A refusal involving an indispensable input can, in exceptional circumstances, constitute abuse of dominance.

Relevance to synthetic data

A synthetic-data provider controlling a unique dataset might face analogous questions where:

  • the data is indispensable;
  • competitors cannot realistically reproduce it;
  • refusal excludes effective competition;
  • downstream innovation depends upon access.

However, synthetic-data providers cannot automatically be treated as having an obligation to license their datasets.

19. Case 2: Bronner v Mediaprint

Principle

The Court of Justice applied a stringent test to refusal of access to an infrastructure allegedly indispensable for competition.

The mere fact that access would be advantageous or economically preferable was insufficient.

Synthetic-data relevance

This is important because competitors may argue:

“We need Provider A's synthetic dataset to compete.”

The legal question is more demanding:

Is the dataset genuinely indispensable, or can competitors develop alternative datasets?

If alternative generation methods exist, compulsory access becomes much harder to justify.

20. Case 3: IMS Health GmbH & Co. OHG v NDC Health

Principle

The case concerned access to a data structure and intellectual-property-related refusal to license.

The Court identified stringent conditions for compulsory licensing, including indispensability and elimination of competition.

Synthetic-data relevance

This is particularly relevant to:

  • proprietary synthetic-data architectures;
  • specialised data formats;
  • unique datasets;
  • data-generation standards;
  • proprietary schemas.

A provider could potentially have intellectual-property rights in aspects of its technology, but intellectual property does not provide unlimited immunity from competition law.

21. Case 4: Microsoft Corp. v Commission

Principle

The European Commission and EU courts considered Microsoft's refusal to provide interoperability information to competing work-group server products.

The case demonstrated how interoperability restrictions can reinforce dominance.

Synthetic-data relevance

The analogy is strong where a dominant synthetic-data platform controls:

  • API access;
  • data formats;
  • interoperability information;
  • technical documentation;
  • data-transfer mechanisms.

If competitors cannot interoperate with a dominant platform, competition may be weakened.

The broader EU data-policy environment now expressly recognises the importance of interoperability and data portability in digital markets.

22. Case 5: Google Shopping

Principle

The European Commission and EU courts examined Google's treatment of competing comparison-shopping services and its use of dominance in general search to favour its own specialised service.

Synthetic-data relevance

The important analogy is leveraging and self-preferencing.

Suppose:

Google-type platform → synthetic-data service → own AI product

and the platform systematically gives its own synthetic-data business superior access to:

  • user-generated information;
  • search information;
  • computational resources;
  • APIs;
  • data updates.

The competitive analysis could resemble the broader principle that dominance in one market can be leveraged to disadvantage competitors in another.

23. Case 6: Facebook/WhatsApp Merger Decision

The European Commission's examination of the Facebook/WhatsApp transaction is important for data-driven merger analysis.

Principle

The transaction raised questions concerning the competitive importance of data and the ability of a large platform to combine data resources.

Synthetic-data relevance

A merger between:

large consumer-data platform + synthetic-data provider

could create a significantly expanded data ecosystem.

Potential concerns include:

  • combining datasets;
  • increasing entry barriers;
  • improving AI training capabilities;
  • reducing rivals' access to comparable information;
  • creating a vertically integrated data ecosystem.

Data concentration therefore can matter even when the parties' products are not straightforward substitutes.

24. Case 7: Bundeskartellamt v Facebook/Meta

Germany's competition authority examined Facebook's combination of data from different services in the context of Facebook's dominant position.

The case is significant because it demonstrated how data collection and data combination can intersect with competition law.

The subsequent EU litigation concerning Meta's data practices further illustrates the interaction between competition law, personal-data rules and market power.

Synthetic-data relevance

A synthetic-data provider might attempt to combine:

  • customer data;
  • behavioural data;
  • third-party datasets;
  • proprietary datasets;

to create a synthetic dataset unavailable to competitors.

The competition analysis may therefore consider whether the accumulation and combination of data reinforce market power.

25. Case 8: Aspen Skiing Co. v Aspen Highlands Skiing Corp.

Principle

The U.S. Supreme Court considered a refusal to continue a prior cooperative arrangement where the conduct was found to have exclusionary characteristics.

Synthetic-data relevance

The case is useful when analysing a synthetic-data provider that:

  1. previously supplied data to competitors;
  2. abruptly terminates access;
  3. sacrifices profitable transactions;
  4. uses the termination to disadvantage a rival.

The case should not be interpreted as establishing a general duty to deal. Its relevance is tied to its specific factual circumstances.

26. Case 9: United States v. Google

The U.S. Google litigation illustrates broader concerns concerning exclusionary conduct in digital ecosystems.

The DOJ's Google litigation remains active in the remedies phase as of 2026.

Synthetic-data relevance

The case provides an important framework for examining whether control over a strategically important digital input or distribution channel can be used to reinforce market power.

For synthetic-data providers, the analogous concern would be:

control over data-generation infrastructure → exclusion of competing AI/data services.

27. Case-Law Principles Compared

CaseCore principleSynthetic-data relevance
MagillExceptional refusal to licenseAccess to indispensable datasets
BronnerStrict indispensability testRefusal to supply synthetic data
IMS HealthData/IP access and exceptional circumstancesProprietary datasets and formats
MicrosoftInteroperabilityAPIs and data portability
Google ShoppingLeveraging/self-preferencingFavouring own AI/data products
Facebook/WhatsAppData concentration in merger analysisCombining data ecosystems
Facebook/MetaData combination and dominanceData accumulation strategies
Aspen SkiingExceptional refusal to dealTermination of established data access

28. Synthetic Data and Essential-Facility Doctrine

An essential-facility claim involving synthetic data would likely require careful analysis of:

1. Indispensability

Can competitors create equivalent synthetic data themselves?

2. Replicability

Can the underlying dataset or generation technology be reproduced?

3. Cost

Is replication technically or economically unrealistic?

4. Competition elimination

Would refusal eliminate effective competition?

5. Justification

Does the provider have legitimate reasons for refusing access?

6. Remedy feasibility

Can access be provided without destroying legitimate incentives to innovate?

This is particularly important because excessive compulsory-access obligations could reduce incentives to invest in expensive data-generation technology.

29. Privacy and Competition Law

Synthetic data is often promoted as a privacy-enhancing technology.

That does not mean privacy and competition are unrelated.

The FTC has expressly recognised the interaction between data control, privacy and competition in digital markets.

Competition authorities may therefore ask:

  • Who owns the underlying data?
  • How was it collected?
  • Can it legally be combined?
  • Is the synthetic dataset actually anonymous?
  • Can individuals be re-identified?
  • Does privacy regulation prevent rivals from obtaining equivalent data?
  • Does the dominant firm gain a competitive advantage from privileged access?

The distinction between privacy-preserving data and competition-preserving data access is therefore crucial.

30. Re-identification Risk

A synthetic dataset may still create competitive and regulatory concerns if it permits reconstruction of sensitive information.

For example:

Original dataset → generative model → synthetic dataset

If the generated dataset contains memorised or reconstructable information, the provider may possess a competitive advantage based on access to sensitive underlying information.

This also creates potential asymmetry:

Incumbent: access to real data + synthetic generation
Entrant: synthetic generation only

That asymmetry can become a significant entry barrier.

31. Data Portability

Data portability can reduce switching costs.

The EU already recognises individual data-portability rights under GDPR in specified circumstances, while the DMA contains additional competition-oriented portability and access obligations for designated gatekeepers.

For synthetic-data markets, portability could concern:

  • generated datasets;
  • metadata;
  • schemas;
  • prompts;
  • generation parameters;
  • validation records;
  • model outputs;
  • API configurations.

The competitive objective would be to prevent technical incompatibility from unnecessarily locking customers into one provider.

32. Recent Regulatory Development: Google Search Data

The developing EU DMA approach is particularly relevant to future synthetic-data competition.

In 2026, the European Commission specified measures requiring Google to share anonymised search information—including ranking, query, click and view data—with qualifying competing search providers under fair, reasonable and non-discriminatory conditions.

The significance for synthetic data is conceptual:

Data generated by a dominant platform can become a competitively significant input for downstream innovation.

This may become increasingly relevant where AI developers use search-derived or platform-derived information to construct synthetic training datasets.

33. Data Portability as a Competitive Remedy

The UK CMA has also imposed a data-portability conduct requirement on Google, requiring tools that facilitate effective portability of search data where authorised by consumers.

For synthetic-data markets, analogous remedies could potentially include:

  • standardized APIs;
  • machine-readable datasets;
  • continuous data transfer;
  • interoperability requirements;
  • non-discriminatory access;
  • transparent technical standards.

34. Competition Risks Specific to Synthetic-Data Providers

The principal risks can be classified as follows:

Structural risks

  • high entry costs;
  • proprietary datasets;
  • computational requirements;
  • specialised expertise;
  • economies of scale.

Behavioural risks

  • refusal to supply;
  • discriminatory access;
  • tying;
  • bundling;
  • exclusive dealing;
  • self-preferencing;
  • API restrictions.

Data risks

  • data accumulation;
  • data combination;
  • data quality manipulation;
  • re-identification;
  • insufficient portability.

Technological risks

  • proprietary formats;
  • interoperability restrictions;
  • model lock-in;
  • technical switching costs.

Merger risks

  • vertical integration;
  • data concentration;
  • elimination of nascent competitors;
  • foreclosure of downstream AI firms.

35. Possible Competition-Law Remedies

Authorities could potentially consider:

Structural remedies

  • divestiture;
  • separation of business units;
  • restrictions on vertical integration.

Behavioural remedies

  • FRAND access;
  • non-discrimination obligations;
  • API access;
  • interoperability requirements;
  • data portability;
  • prohibition of tying;
  • prohibition of exclusive dealing.

Transparency remedies

  • disclosure of access conditions;
  • quality standards;
  • technical specifications;
  • audit mechanisms.

Data remedies

  • controlled data sharing;
  • anonymisation;
  • secure data rooms;
  • independent data trustees;
  • standardized formats.

Remedies must nevertheless protect legitimate privacy, cybersecurity and intellectual-property interests.

36. Compliance Framework for Synthetic-Data Providers

A synthetic-data provider should establish:

Competition compliance

  1. Identify relevant markets.
  2. Assess market power.
  3. Document objective access criteria.
  4. Avoid discriminatory treatment.
  5. Review exclusivity arrangements.
  6. Monitor tying and bundling.
  7. Maintain interoperability where commercially appropriate.

Data governance

  1. Document data provenance.
  2. Establish lawful collection procedures.
  3. Test re-identification risk.
  4. Validate synthetic-data quality.
  5. Maintain audit trails.

Merger compliance

  1. Screen acquisitions for data concentration.
  2. Assess vertical foreclosure.
  3. Examine potential-competition effects.
  4. Identify strategic datasets.

37. Hypothetical Example

Assume Company A operates the largest healthcare synthetic-data platform.

It controls:

  • 80% of a specialised synthetic medical-data segment;
  • proprietary hospital datasets;
  • a generation model;
  • an API;
  • an AI-development platform.

Company A supplies its own AI diagnostic subsidiary with real-time high-quality synthetic data but provides independent AI developers with delayed and lower-quality datasets.

Potential competition concerns could include:

Dominance

↓

Control of important healthcare data

↓

Preferential supply to own AI subsidiary

↓

Disadvantage to competing AI developers

↓

Potential leveraging / discriminatory access

The authority would still need to establish the relevant market, dominance, actual or likely exclusionary effects, and the absence of objective justification.

38. Key Doctrinal Tension

Synthetic-data competition law presents a fundamental tension:

Innovation argument

Strong protection of proprietary datasets and generation technologies may:

  • reward investment;
  • encourage innovation;
  • protect trade secrets;
  • promote better synthetic-data technologies.

Competition argument

Excessive control may:

  • increase entry barriers;
  • entrench incumbents;
  • prevent interoperability;
  • restrict AI innovation;
  • create data bottlenecks.

Competition law must therefore distinguish competition on the merits from strategic exclusion of competitors.

39. Conclusion

Synthetic-data providers are likely to become increasingly important participants in data-driven markets. Their competitive significance may arise not only from the synthetic datasets they sell, but from their control over underlying data, generation models, computational infrastructure, APIs, quality standards and downstream AI ecosystems.

The principal competition-law questions concern:

  1. market definition;
  2. data-driven dominance;
  3. refusal to supply;
  4. essential-facility arguments;
  5. self-preferencing;
  6. discriminatory data access;
  7. tying and bundling;
  8. exclusive agreements;
  9. interoperability and switching costs;
  10. algorithmic coordination;
  11. data-driven mergers;
  12. vertical foreclosure.

The most useful existing jurisprudence comes from Magill, Bronner, IMS Health, Microsoft, Google Shopping, Facebook/WhatsApp, Facebook/Meta, and Aspen Skiing. These cases do not establish a standalone doctrine for synthetic data; rather, they provide established principles that can be adapted to the competitive characteristics of synthetic-data markets.

LEAVE A COMMENT