Competition Law And Synthetic Data Market Competition Issues
Competition Law and Synthetic Data Market Competition Issues
1. Introduction
Synthetic data refers to artificially generated data designed to reproduce relevant statistical, structural, behavioural, or technical characteristics of real-world datasets without necessarily reproducing the underlying individuals or records. It is increasingly used in artificial intelligence, healthcare, autonomous vehicles, financial services, cybersecurity, robotics, digital advertising, manufacturing, and scientific research.
From a competition-law perspective, synthetic data creates a distinctive problem: although it can reduce dependence on scarce personal or proprietary datasets, the companies controlling the systems that generate, validate, distribute, or benchmark synthetic data may themselves become important competitive bottlenecks.
There is not yet a large body of reported competition cases dealing expressly with a "synthetic data market." Consequently, the most relevant jurisprudence comes from data-access, digital-platform, interoperability, tying, self-preferencing, exclusionary conduct, and data-concentration cases. Those cases provide principles that can be applied to synthetic-data markets.
2. What Is the Synthetic Data Market?
The synthetic-data ecosystem can be divided into several potentially distinct markets:
- Synthetic-data generation
- Generative models
- GANs
- Diffusion models
- Large language models
- Simulation engines
- Synthetic-data validation
- Accuracy testing
- Statistical similarity
- Bias testing
- Privacy testing
- Utility benchmarking
- Synthetic-data platforms
- Data marketplaces
- Data-as-a-Service platforms
- Industry-specific repositories
- Synthetic-data infrastructure
- Cloud computing
- GPU infrastructure
- Model hosting
- Storage
- Synthetic-data applications
- Healthcare AI
- Autonomous driving
- Financial modelling
- Defence and cybersecurity
- Robotics
- Hybrid data markets
- Real data + synthetic data
- Synthetic data + proprietary datasets
- Synthetic data + foundation models
A competition authority therefore has to determine which market is actually affected rather than assuming that all synthetic-data activities constitute one market.
3. Why Synthetic Data Can Create Competition Problems
A. Data Concentration
Synthetic data does not necessarily eliminate data concentration.
A company possessing:
- superior generation models,
- proprietary validation datasets,
- enormous computing resources,
- specialised domain knowledge,
- customer feedback,
may produce substantially better synthetic datasets than smaller competitors.
This can produce a data-generation feedback loop:
More users → more real-world feedback → better models → better synthetic data → more users → more feedback.
The result can be increasing returns and barriers to entry.
4. Market Definition
Competition authorities may need to determine whether synthetic data is:
- a separate product from real data;
- a substitute for real data;
- complementary to real data;
- a distinct data-generation service; or
- part of a broader AI/cloud/data-services market.
The answer may differ by industry.
For example, synthetic medical records may not be sufficiently substitutable for clinical records where actual clinical validation is required.
Conversely, synthetic images may compete closely with real images for training certain computer-vision systems.
Therefore, functional substitutability rather than the label "synthetic" should ordinarily be central to market definition.
5. Market Power in Synthetic Data
Traditional market shares may be insufficient.
A company may have substantial competitive significance because it controls:
- a unique generation model;
- proprietary validation technology;
- a dominant cloud ecosystem;
- important APIs;
- industry benchmarks;
- specialised datasets;
- customer switching infrastructure.
Competition analysis should therefore consider:
Traditional factors
- market share;
- entry barriers;
- pricing;
- customer alternatives;
- network effects.
Data-specific factors
- uniqueness of datasets;
- quality;
- accuracy;
- representativeness;
- interoperability;
- portability;
- retraining costs;
- validation requirements.
6. Major Competition Issues
I. Refusal to Supply Synthetic Data
Suppose a dominant synthetic-data provider supplies essential training datasets to competing AI developers but refuses access to selected competitors.
The issue may arise under the essential-facilities doctrine where the relevant legal requirements are satisfied.
Important questions include:
- Is the data genuinely indispensable?
- Are viable alternatives available?
- Is replication technically or economically feasible?
- Does refusal eliminate effective competition?
- Is there an objective justification?
- Would access undermine privacy, security, or intellectual-property rights?
Synthetic data is unusual because, unlike an irreplaceable natural resource, it may sometimes be recreated.
Consequently, indispensability may be harder to establish.
7. Data Access Discrimination
A dominant synthetic-data platform might provide:
- high-quality datasets to its own AI subsidiary;
- inferior datasets to independent competitors;
- delayed API access;
- discriminatory licensing terms;
- preferential validation;
- different technical limits.
This could resemble self-preferencing or discriminatory-access theories developed in digital-platform competition law.
8. Self-Preferencing
A vertically integrated company could operate:
Synthetic-data generation → data marketplace → AI model → application platform.
It might rank its own synthetic datasets above rival datasets.
For example:
Third-party synthetic dataset
↓ lower ranking
Platform's own dataset
↓ preferential ranking
Platform's AI model
Such conduct could disadvantage competing data providers.
The competition concern becomes particularly strong when the platform controls an important distribution channel.
9. Bundling and Tying
A dominant cloud provider might require customers purchasing synthetic-data generation to also purchase:
- its cloud infrastructure;
- its AI model;
- its storage;
- its analytics system.
For example:
"Synthetic-data generation is available only with our cloud-hosting package."
Such conduct could foreclose independent synthetic-data generators.
The legal analysis would consider:
- dominance;
- separate products;
- coercion;
- foreclosure;
- objective justification;
- efficiencies.
10. Exclusive Dealing
A synthetic-data provider could require customers to agree:
"Synthetic datasets purchased from this platform may not be used with competing AI models."
This could raise concerns under rules governing:
- exclusive supply;
- exclusive purchasing;
- foreclosure;
- platform dependence.
The concern increases where the provider has substantial market power and customers cannot readily obtain comparable datasets elsewhere.
11. Synthetic Data and Algorithmic Collusion
Synthetic data can also facilitate algorithmic coordination.
Competitors might use the same:
- synthetic-data generator;
- pricing algorithm;
- market simulator;
- demand model.
If competing firms rely upon a common algorithmic infrastructure, there could be risks of parallel pricing or reduced competitive uncertainty.
The legal question is whether the common system merely facilitates independent decision-making or becomes part of an unlawful agreement or coordinated practice.
12. Synthetic Data and AI Training
Synthetic datasets may become a critical input for AI models.
A vertically integrated firm controlling:
Data → Synthetic data → Foundation model → Cloud → Application
may possess several layers of market power.
This creates the possibility of vertical leveraging.
For example:
Synthetic-data provider
→ gives preferential access to its own AI subsidiary
→ restricts competitors
→ competitors obtain less effective training data
→ AI-market competition is weakened.
13. Quality as a Dimension of Competition
Synthetic data creates an important competition concept: data quality.
Two datasets may have the same number of records but dramatically different competitive value.
Relevant quality dimensions include:
- accuracy;
- diversity;
- statistical fidelity;
- temporal relevance;
- geographic coverage;
- bias;
- privacy;
- representativeness;
- label quality.
A dominant company could theoretically degrade competitors' access to high-quality data while maintaining nominal access.
Thus, competition authorities may need to examine quality discrimination, not merely price discrimination.
14. Interoperability
Synthetic-data providers may use proprietary:
- formats;
- APIs;
- metadata;
- schemas;
- validation protocols.
If competitors cannot easily import or export datasets, customers may become locked into one ecosystem.
This creates switching costs.
A competitive synthetic-data market therefore benefits from:
- interoperable formats;
- open APIs;
- portability;
- transparent documentation;
- common validation standards.
15. Switching Costs
Suppose an AI company trains its systems around one provider's synthetic-data architecture.
Switching providers may require:
- retraining models;
- rewriting pipelines;
- revalidating datasets;
- regulatory reapproval;
- changing APIs;
- reconstructing metadata.
The effective switching cost can therefore be much higher than the subscription price.
Competition authorities should distinguish legitimate investment-related switching costs from artificially imposed restrictions designed to prevent customers from changing suppliers.
16. Network Effects
Synthetic-data platforms can exhibit network effects.
More customers may generate:
- more feedback;
- more validation;
- more domain-specific improvements;
- more model optimisation;
- more benchmarking information.
This can create a data-network-effect cycle.
However, synthetic-data markets can also have economies of scale without classical network effects because the marginal cost of generating additional datasets can be relatively low once the underlying infrastructure has been built.
17. Mergers and Acquisitions
Synthetic-data acquisitions can raise difficult merger questions.
Consider:
Dominant AI company acquires leading synthetic-data generator.
Even if the synthetic-data company has modest current revenue, the acquisition may remove a significant future competitor.
Authorities may therefore examine:
- innovation competition;
- potential competition;
- access to unique data;
- vertical foreclosure;
- interoperability;
- elimination of nascent rivals;
- control over AI inputs.
This is particularly relevant where conventional turnover thresholds fail to capture the strategic value of the target.
18. Killer Acquisition Risk
A large AI/cloud company might acquire a small synthetic-data startup before the startup becomes a meaningful competitor.
The target could possess:
- superior synthetic-data generation;
- specialised healthcare data technology;
- autonomous-driving simulation;
- privacy-preserving generation;
- unique validation technology.
The acquisition may therefore have competitive significance despite relatively low present revenue.
19. Privacy and Competition Law
Synthetic data is often promoted as a privacy-enhancing technology.
But synthetic data is not automatically anonymous.
Poorly designed models can:
- reproduce identifiable information;
- leak training data;
- enable membership inference;
- reproduce rare records;
- permit re-identification.
This has competition implications where privacy is a non-price dimension of competition.
The Meta/Facebook litigation is particularly relevant because the CJEU recognised that data-protection considerations can be relevant to competition-law assessment in appropriate circumstances.
20. Six Important Case Laws
Because reported cases specifically concerning synthetic-data markets remain limited, the following cases are important analogical precedents for analysing synthetic-data competition.
1. Meta Platforms v Bundeskartellamt — C-252/21
Facts
The German competition authority objected to Meta's combination of Facebook data with data obtained from other Meta services and third-party sources.
The CJEU considered the relationship between:
- dominance;
- data processing;
- GDPR;
- competition law.
The Court confirmed that a competition authority may, in appropriate circumstances, take GDPR compliance into account when assessing abuse of dominance.
Relevance to synthetic data
This is highly relevant where a synthetic-data provider combines:
- real-world personal data;
- third-party datasets;
- behavioural information;
- generated data.
Competition analysis cannot necessarily treat data practices as completely separate from competitive conditions.
Principle
Data governance can form part of competition-law analysis where it is sufficiently connected to the competitive conduct under examination.
2. Bundeskartellamt – Facebook Data Combination Case
The Bundeskartellamt found that Facebook's practice of making use of the social network conditional upon extensive combination of data from Facebook, Instagram, WhatsApp and third-party sources could constitute an abuse of dominance. The case ultimately led to extensive litigation and implementation measures.
Synthetic-data relevance
A synthetic-data platform could similarly create competitive concerns by requiring users to provide extensive data to generate or improve synthetic datasets.
For example:
"Use our synthetic-data service only if you permit us to combine your operational data with our entire ecosystem."
The competition issue would concern whether the data requirement is legitimately connected to the service or is being used to reinforce dominance.
3. Google Shopping — Google and Alphabet v Commission
The Google Shopping litigation concerned Google's preferential treatment of its own comparison-shopping service within general search results.
Synthetic-data relevance
The underlying principle is important for a synthetic-data platform operating as both:
- marketplace;
- data generator;
- downstream AI provider.
If the platform systematically places its own synthetic datasets or services above competing datasets, the conduct may raise self-preferencing concerns.
Principle
Vertical integration becomes competitively significant when a platform controlling an important intermediary position gives its downstream service preferential treatment.
4. Google Android — Google and Alphabet v Commission
The Android case involved several contractual and ecosystem practices concerning Google's position in mobile operating systems, including tying-related conduct.
Synthetic-data relevance
The case illustrates how competition law can address strategies through which a powerful ecosystem provider uses control over one layer of technology to reinforce another.
A comparable synthetic-data ecosystem might involve:
Cloud infrastructure + synthetic-data generation + AI model + application layer.
If customers cannot realistically purchase the relevant components independently, tying or leveraging concerns can arise.
5. Amazon Marketplace Proceedings
The Bundeskartellamt's Amazon proceedings have examined Amazon's role as both marketplace operator and seller, including price-parity provisions and other conditions affecting marketplace sellers. The authority has also examined Amazon's pricing mechanisms and its use of marketplace controls.
Synthetic-data relevance
The analogy is particularly important where a synthetic-data platform simultaneously:
- operates the marketplace;
- sells its own synthetic data;
- hosts competitors' data;
- controls ranking;
- controls access.
This creates potential conflict-of-interest and self-preferencing risks.
6. SAMR v Alibaba
China's 2021 Alibaba decision concerned Alibaba's practice of requiring merchants to choose between Alibaba and competing platforms, commonly described as an "either-or" exclusivity arrangement.
SAMR imposed a substantial penalty for abuse of dominance. Academic analysis also notes that the enforcement programme addressed data, algorithms and platform practices affecting competition.
Synthetic-data relevance
The case provides an important framework for examining:
- exclusivity;
- platform power;
- data advantages;
- ecosystem foreclosure.
A synthetic-data platform could potentially engage in similar conduct by requiring customers to use its generated datasets exclusively with its own AI models or infrastructure.
7. Additional Relevant Case: Microsoft/LinkedIn Data Issues
The Microsoft/LinkedIn merger is also useful as a data-driven merger precedent.
The transaction illustrated the importance of considering:
- datasets;
- access to data;
- platform ecosystems;
- potential competitive advantages generated by combining data assets.
Synthetic-data relevance
Synthetic-data mergers may similarly require analysis beyond turnover and conventional product overlaps.
A small synthetic-data company can possess strategically important technology even when its current sales are modest.
21. China-Specific Competition Perspective
China's Anti-Monopoly Law is particularly relevant because China's digital-platform enforcement has increasingly considered:
- data;
- algorithms;
- platform rules;
- technological means;
- exclusive dealing;
- abuse of dominance.
Chinese enforcement materials have specifically addressed the need for platforms to use data and algorithms in ways that do not exclude or restrict competition.
The Alibaba decision therefore provides an important foundation for analysing synthetic-data platforms in China.
Potential Chinese competition issues include:
Article 17-type abuse concerns
- refusal to supply;
- discriminatory treatment;
- unreasonable conditions;
- tying;
- exclusive arrangements.
Article 22-type dominance concerns
A synthetic-data company with substantial market power could face scrutiny if it uses that position to restrict downstream AI competition.
Merger control
Acquisitions of synthetic-data startups may raise concerns regarding:
- potential competition;
- innovation;
- data advantages;
- ecosystem foreclosure.
22. Synthetic Data and Essential Facilities
An important hypothetical is:
A dominant company controls the only commercially viable synthetic dataset capable of training a particular safety-critical AI system.
A refusal to license the dataset could potentially trigger essential-facility analysis.
However, mere importance is not enough.
The claimant would generally need to establish the relevant legal requirements, which may include:
- control by a dominant undertaking;
- indispensability;
- absence of realistic alternatives;
- inability to reproduce the input;
- potential elimination of effective competition;
- lack of objective justification.
Synthetic data may make the indispensability requirement particularly difficult because competitors may be able to generate alternative datasets.
23. Synthetic Data and Predatory Pricing
Because synthetic data can have low marginal reproduction costs, a large provider could potentially price datasets:
- below cost;
- at zero;
- bundled with another service.
Free synthetic data is not automatically anti-competitive.
The competition question is whether the pricing strategy is being used to:
eliminate competitors → establish dominance → subsequently exploit customers.
Cost standards and evidence of exclusionary intent/effect would therefore become important.
24. Synthetic Data and Price Discrimination
A dominant provider could charge:
- large enterprises: high price;
- startups: low price;
- competitors: very high price;
- affiliated companies: zero or near-zero price.
Such differential pricing is not inherently unlawful.
Competition authorities would need to determine whether the differences amount to prohibited discriminatory treatment or facilitate exclusionary conduct.
25. Algorithmic Transparency
Synthetic-data generation may involve complex algorithms whose operation is difficult for competitors and regulators to understand.
Competition authorities may need information about:
- model architecture;
- training sources;
- data-generation methodology;
- validation methodology;
- API restrictions;
- pricing algorithms;
- ranking systems.
The challenge is balancing competition enforcement against:
- trade secrets;
- cybersecurity;
- intellectual property;
- privacy.
26. Synthetic Data and Data Portability
Portability can reduce competitive lock-in.
A customer should, where technically and legally appropriate, be able to migrate:
Dataset → metadata → schemas → validation records → model pipeline
to another provider.
Where a dominant provider intentionally makes migration difficult, the resulting switching costs may reinforce market power.
27. Competition Between Synthetic-Data Generators
Competition should not focus solely on price.
Relevant dimensions include:
| Competition parameter | Importance |
|---|---|
| Price | Direct purchasing cost |
| Accuracy | Reliability of generated data |
| Diversity | Coverage of different cases |
| Privacy | Risk of information leakage |
| Bias | Fairness and representativeness |
| Scalability | Ability to generate large datasets |
| Interoperability | Ease of switching |
| Validation | Reliability of synthetic records |
| Customisation | Industry-specific usefulness |
| Security | Protection against attacks |
This means synthetic-data competition is likely to be multidimensional and innovation-driven.
28. Possible Anticompetitive Scenarios
Scenario 1 — Data foreclosure
Dominant company refuses access to critical synthetic datasets.
Issue: exclusionary refusal to supply.
Scenario 2 — Self-preferencing
Platform ranks its own synthetic datasets above competitors.
Issue: leveraging/platform neutrality.
Scenario 3 — Bundling
Synthetic-data generation is available only with the provider's cloud service.
Issue: tying.
Scenario 4 — Exclusivity
Customers are prohibited from using rival synthetic-data providers.
Issue: exclusive dealing.
Scenario 5 — Data discrimination
Competitors receive lower-quality datasets or restricted APIs.
Issue: discriminatory access.
Scenario 6 — Acquisition
Dominant AI company purchases the leading independent synthetic-data startup.
Issue: loss of potential competition and innovation.
29. Possible Pro-Competitive Justifications
Not every restriction should be regarded as anticompetitive.
Synthetic-data providers may legitimately impose restrictions because of:
- cybersecurity;
- privacy;
- intellectual-property protection;
- data integrity;
- model security;
- regulatory compliance;
- prevention of misuse;
- quality assurance.
A competition assessment should therefore distinguish legitimate technical restrictions from exclusionary restrictions.
30. Remedies
Competition authorities could potentially consider:
Structural remedies
- divestiture;
- separation of businesses;
- prohibition of certain acquisitions.
Behavioural remedies
- non-discriminatory access;
- API access;
- interoperability;
- data portability;
- transparent ranking;
- prohibition of exclusivity;
- fair licensing.
Merger remedies
- data-access commitments;
- interoperability commitments;
- licensing commitments;
- restrictions on combining datasets;
- preservation of independent development teams.
31. Six-Core Case-Law Principles at a Glance
| Case | Core competition principle | Synthetic-data relevance |
|---|---|---|
| Meta Platforms v Bundeskartellamt, C-252/21 | Data practices can intersect with abuse-of-dominance analysis | Data combination and privacy |
| Bundeskartellamt Facebook proceedings | Dominant platform's data practices may constitute abuse | Data accumulation and ecosystem power |
| Google Shopping | Self-preferencing can raise exclusionary concerns | Ranking own synthetic datasets |
| Google Android | Ecosystem leveraging/tying can restrict competition | Bundling synthetic data with cloud/AI |
| Amazon Marketplace | Platform neutrality and seller-access conditions matter | Synthetic-data marketplace discrimination |
| SAMR v Alibaba | Platform exclusivity can constitute abuse of dominance | Exclusive synthetic-data arrangements |
The first two are particularly relevant to data as a competitive asset, while the Google, Amazon and Alibaba cases provide useful frameworks for analysing platform conduct, vertical integration, access, ranking and exclusivity.
32. Key Legal Tests
A competition authority examining synthetic-data conduct should ask:
Market
- What is the relevant product market?
- Is synthetic data substitutable for real data?
- Is the market local, national, regional or global?
Power
- Does the undertaking possess market power?
- Does it control a unique dataset, model or infrastructure?
- Are there meaningful alternatives?
Conduct
- Is there exclusionary conduct?
- Is there discriminatory access?
- Is there tying or bundling?
- Is there exclusivity?
- Is there self-preferencing?
- Is interoperability being restricted?
Effects
- Are competitors being foreclosed?
- Are innovation incentives reduced?
- Are switching costs increased?
- Is consumer choice reduced?
Justification
- Is the restriction technically necessary?
- Is it proportionate?
- Does it protect privacy or security?
- Are there less restrictive alternatives?
33. Conclusion
Synthetic data can reduce traditional data scarcity without necessarily eliminating competition problems. The critical competitive asset may shift from ownership of raw data to control over the technology, infrastructure, validation systems, distribution channels and ecosystems used to create synthetic data.
The most important competition-law concerns are therefore:
- data concentration;
- refusal of access;
- discriminatory licensing;
- self-preferencing;
- tying and bundling;
- exclusive dealing;
- interoperability restrictions;
- switching costs;
- algorithmic coordination;
- vertical foreclosure;
- data-driven mergers; and
- control over innovation inputs.
The existing jurisprudence does not yet provide a mature, standalone doctrine of "synthetic-data competition law." Instead, cases such as Meta v Bundeskartellamt, Google Shopping, Google Android, Amazon Marketplace and SAMR v Alibaba provide the principal analytical building blocks. Meta is especially important because the CJEU confirmed that data-protection considerations can, in appropriate circumstances, be relevant to a competition authority's assessment of data-driven conduct.
Accordingly, future synthetic-data competition law is likely to develop around a central question:
Who contro

comments