Global Knowledge Graphs And Epistemic Infrastructure Concentration .
Global Knowledge Graphs And Epistemic Infrastructure Concentration
Introduction
Global knowledge graphs are large-scale systems that connect entities, facts, concepts, documents, people, organizations, locations, events, and relationships into machine-readable structures. They increasingly underpin search engines, recommendation systems, AI assistants, scientific discovery, financial intelligence, mapping, healthcare, government services, and automated decision-making.
Epistemic infrastructure is broader. It means the technological, institutional, and informational systems through which society determines:
- what information is available;
- how information is classified and connected;
- which sources are considered authoritative;
- what facts can be discovered;
- how conflicting information is ranked;
- which datasets AI systems can access;
- and ultimately what information is capable of influencing decisions.
The competition-law concern arises when a small number of firms control critical layers of this infrastructure. Concentration may therefore produce epistemic market power: the ability to influence not merely prices or output, but the information environment within which consumers, businesses, researchers, governments and AI systems make decisions.
1. Meaning Of Knowledge-Graph Concentration
A knowledge graph generally contains:
Entities → attributes → relationships → sources → inference → outputs
For example:
Company A → owns → Company B → operates → Platform C → collects → Dataset D.
A large knowledge graph can therefore become a connective layer between enormous quantities of information.
Concentration can occur at several levels:
- Data acquisition
- Data aggregation
- Entity resolution
- Knowledge-graph construction
- Ranking and retrieval
- AI-model access
- API distribution
- User-interface control
- Institutional integration
The competitive problem becomes particularly serious when the same undertaking controls several layers simultaneously.
2. What Is Epistemic Infrastructure?
Traditional infrastructure consists of roads, electricity grids, telecommunications networks and payment systems.
Epistemic infrastructure performs a comparable function for information.
It may include:
- search indexes;
- knowledge graphs;
- digital libraries;
- scientific databases;
- citation networks;
- mapping databases;
- identity databases;
- data brokers;
- cloud repositories;
- AI training datasets;
- model-access APIs;
- content-discovery systems;
- fact-checking systems;
- recommendation engines;
- metadata standards.
The distinction is important because an undertaking can possess substantial market power even where its service is nominally free.
The relevant competitive asset may be control over information flows rather than monetary transactions.
3. Sources Of Concentration
A. Data Scale
Knowledge systems benefit from enormous datasets.
The more entities and relationships a platform contains, the more useful its graph may become.
This can generate:
scale → better results → more users → more data → better results.
That feedback loop may create substantial entry barriers.
B. Network Effects
Users often prefer systems that contain the greatest number of entities and relationships.
A knowledge graph therefore exhibits a form of data/network effect.
An entrant may possess excellent technology but still struggle because it lacks:
- historical data;
- entity mappings;
- user-generated corrections;
- behavioural signals;
- source relationships;
- structured metadata.
C. Switching Costs
Businesses and public institutions may build workflows around a particular:
- API;
- identifier system;
- ontology;
- database;
- knowledge graph;
- cloud environment.
Once integrated, migration becomes expensive.
Switching costs can therefore transform technological superiority into durable market power.
D. Interoperability Barriers
If competing graphs cannot easily exchange:
- identifiers;
- metadata;
- ontologies;
- provenance;
- semantic relationships;
then users may become locked into one ecosystem.
Interoperability is consequently an important competition-law remedy.
4. Epistemic Gatekeeping
The most significant concern is gatekeeping over discoverability.
Suppose information exists publicly but a dominant intermediary controls:
- indexing;
- ranking;
- entity classification;
- recommendation;
- AI retrieval.
The information technically remains available, but its practical accessibility may depend upon the intermediary.
This creates a distinction between:
formal availability and effective availability.
Competition authorities may therefore need to consider whether exclusion from a dominant knowledge system substantially reduces an organization's ability to reach users.
5. Knowledge Graphs As Essential Inputs
A knowledge graph can become an important input where downstream competitors depend upon it.
Potential downstream markets include:
- AI assistants;
- search;
- travel;
- financial analysis;
- healthcare information;
- scientific research;
- mapping;
- advertising;
- cybersecurity;
- enterprise intelligence.
If competitors cannot realistically reproduce the relevant graph, refusal to provide access may raise essential-facility or refusal-to-deal concerns, depending on the jurisdiction.
The analysis would normally require factors such as:
- indispensability;
- absence of realistic alternatives;
- elimination of effective competition;
- technical feasibility of access;
- objective justification;
- proportionality of the requested remedy.
6. Data Advantages And Competition
Data itself does not automatically constitute market power.
The relevant question is whether the particular data advantage is:
- difficult to replicate;
- commercially important;
- continuously refreshed;
- exclusive;
- interoperable;
- necessary for downstream competition.
A dataset becomes particularly powerful when combined with:
proprietary identifiers + behavioural data + search data + transactional data + computational infrastructure.
This creates a compound data advantage rather than a simple database advantage.
7. Vertical Integration
Knowledge infrastructure can create vertical competition problems.
For example:
Data collection → knowledge graph → AI model → search engine → advertising platform
If one undertaking controls the entire chain, it may have incentives to:
- favour its downstream products;
- deny competitors access;
- degrade interoperability;
- impose discriminatory API conditions;
- self-preference its services;
- use proprietary data unavailable to rivals.
This resembles traditional vertical foreclosure but operates through information infrastructure.
8. Self-Preferencing
A dominant platform may rank its own:
- knowledge panels;
- databases;
- AI answers;
- products;
- services;
- vertical search results
above competing sources.
The competitive concern is not necessarily that the preferred result is inaccurate.
The concern is that the platform may use control over the discovery layer to disadvantage competitors.
This makes self-preferencing especially significant where the platform is also the infrastructure upon which those competitors depend.
9. Epistemic Bias As A Competition Issue
Competition law traditionally asks whether conduct harms:
- price;
- quality;
- output;
- innovation;
- consumer choice.
Knowledge infrastructure adds another dimension:
informational quality and diversity.
A dominant graph could influence:
- which businesses are visible;
- which scientific claims are discoverable;
- which sources are treated as authoritative;
- which products are recommended;
- which facts AI systems retrieve.
Consequently, competition authorities may increasingly consider plurality of information sources as an element of quality and innovation.
10. AI Makes Knowledge-Graph Concentration More Important
Generative AI systems increasingly rely on structured and unstructured information.
A concentrated knowledge layer can therefore influence multiple AI systems simultaneously.
For example:
Knowledge Graph
↓
Retrieval system
↓
AI model
↓
Answer
↓
User decision
If a dominant infrastructure provider controls the first layer, it can indirectly influence downstream AI competition.
This creates a new form of upstream epistemic leverage.
11. Key Competition-Law Issues
1. Relevant Market Definition
Possible markets include:
- general search;
- specialized search;
- knowledge-graph services;
- structured data;
- data brokerage;
- AI retrieval;
- enterprise knowledge services;
- identity resolution.
Market definition may need to account for zero-price services and quality dimensions.
2. Dominance
Indicators may include:
- data scale;
- query volume;
- graph coverage;
- API dependency;
- switching costs;
- interoperability;
- technological advantages;
- institutional adoption.
3. Exclusionary Conduct
Potential theories include:
- refusal to supply;
- discriminatory access;
- self-preferencing;
- tying;
- bundling;
- exclusive dealing;
- data foreclosure;
- interoperability restrictions;
- exploitative contractual terms.
4. Merger Control
Acquisitions of:
- datasets;
- specialized databases;
- search technologies;
- graph providers;
- scientific information platforms;
- AI retrieval companies
may create strategic data concentration even where traditional revenue-based thresholds underestimate the transaction's importance.
12. At Least 6 Important Case Laws
1. United States v. Google LLC — Search
The Google search litigation is highly relevant because it concerns control over the search-distribution ecosystem.
The case demonstrates how dominance can be reinforced through agreements affecting access points to search.
Relevance
For knowledge graphs, the analogous concern is whether control over distribution and default access points allows an undertaking to preserve its informational infrastructure advantage and prevent rivals from reaching sufficient scale.
Principle: Control over an important gateway can reinforce an already powerful information intermediary.
2. Google Search (Shopping) — European Commission
The European Commission's Google Shopping decision concerned preferential treatment of Google's own comparison-shopping service within its general search results.
Relevance
This is directly relevant to knowledge infrastructure because a dominant information intermediary can use its ranking architecture to favour an affiliated downstream service.
The case illustrates the competition-law importance of:
- ranking;
- visibility;
- search neutrality;
- self-preferencing;
- traffic foreclosure.
Principle: Control over a discovery infrastructure can be used to disadvantage downstream competitors.
3. Google Android — European Commission
The Android case concerned Google's contractual arrangements surrounding Android, including restrictions involving search and application distribution.
Relevance
Knowledge infrastructure often operates through an ecosystem rather than a single product.
Control over:
- operating systems;
- app stores;
- search;
- defaults;
- APIs
can reinforce control over information-discovery systems.
Principle: Competition analysis may examine how contractual restrictions across complementary technological layers reinforce dominance.
4. Microsoft Corp. v. United States
The Microsoft litigation remains a foundational technology-antitrust case.
Microsoft's conduct involving Windows and web browsers demonstrated how control over a dominant platform can be leveraged into adjacent technological markets.
Relevance
The same economic logic can apply to knowledge infrastructure:
dominant infrastructure → control over access → exclusion of complementary technologies → preservation of ecosystem power.
Principle: A dominant technological platform can use control over an important infrastructure layer to impede adjacent competition.
5. United States v. Microsoft Corp. — D.C. Circuit
The appellate decision is particularly important for its treatment of exclusionary conduct and network effects.
The court recognized that practices that appear individually small can become competitively significant when they protect a dominant position in a networked market.
Relevance
Knowledge graphs exhibit comparable feedback effects.
A dominant graph can become stronger because:
- more users generate more signals;
- more sources improve coverage;
- greater coverage attracts more users;
- greater usage increases downstream dependence.
Principle: Network effects can make exclusionary conduct particularly consequential.
6. Aspen Skiing Co. v. Aspen Highlands Skiing Corp.
The U.S. Supreme Court recognized an exceptional refusal-to-deal theory where a dominant firm terminated a previously profitable course of cooperation and thereby harmed competition.
Relevance
The case provides an important conceptual foundation for analyzing refusal to provide access to information infrastructure.
If a dominant knowledge provider historically supplied information or interoperability to rivals and later withdraws access strategically, the conduct may raise questions concerning exclusionary intent and competitive harm.
Principle: A strategically motivated termination of cooperation can, in exceptional circumstances, constitute unlawful exclusionary conduct.
7. Verizon Communications Inc. v. Law Offices of Curtis V. Trinko, LLP
Trinko limited the scope of compulsory-dealing theories under U.S. antitrust law.
Relevance
This is particularly important for knowledge graphs because it prevents competition law from automatically converting every proprietary information system into a compulsory public resource.
The case emphasizes the importance of:
- preserving incentives to innovate;
- distinguishing competition law from sector regulation;
- establishing genuine anticompetitive conduct.
Principle: Dominance alone does not generally create a universal obligation to share proprietary assets.
8. Bronner v. Mediaprint
The Court of Justice of the European Union established a stringent framework for refusal-to-supply claims involving potentially indispensable infrastructure.
Relevance
A knowledge graph seeking essential-facility treatment would need to demonstrate circumstances approaching genuine indispensability.
This prevents competition law from requiring dominant firms to provide access merely because their information resource is commercially valuable.
Principle: Indispensability and elimination of effective competition are central to exceptional refusal-to-supply cases.
9. IMS Health GmbH & Co. OHG v NDC Health
IMS Health is especially relevant to data infrastructure.
The dispute concerned access to a pharmaceutical data structure protected through intellectual-property rights.
The CJEU recognized circumstances in which refusal to license an intellectual-property asset could constitute abuse of dominance.
Relevance
The case is highly relevant to proprietary:
- ontologies;
- classification systems;
- identifiers;
- databases;
- knowledge structures.
Principle: Intellectual-property protection does not necessarily immunize a dominant infrastructure from competition law where exceptional refusal-to-license conditions are satisfied.
10. Magill
The Magill cases established the foundational European doctrine concerning exceptional compulsory licensing of intellectual property.
Relevance
Where a proprietary knowledge system becomes indispensable for a downstream market, competition law may examine whether withholding access prevents the emergence of new products or services.
This is particularly significant for AI systems that require access to structured information.
Principle: Exceptional circumstances can justify intervention where refusal to license prevents downstream innovation and competition.
13. Comparative Case-Law Lessons
| Case | Core issue | Knowledge-infrastructure relevance |
|---|---|---|
| Google Search (Shopping) | Self-preferencing | Preferential treatment within discovery infrastructure |
| Google Android | Ecosystem restrictions | Leveraging across technological layers |
| Microsoft | Platform foreclosure | Infrastructure-based network effects |
| Aspen Skiing | Refusal to cooperate | Strategic withdrawal of access |
| Trinko | Refusal to deal | Limits of compulsory access |
| Bronner | Essential facilities | Indispensability |
| IMS Health | Data/IP access | Proprietary information structures |
| Magill | Compulsory licensing | Downstream innovation |
14. Global Competition-Law Framework
The issue can be analysed under several legal systems.
European Union
Key tools include:
- Article 101 TFEU;
- Article 102 TFEU;
- EU merger control;
- Digital Markets Act.
Article 102 is particularly relevant to:
- refusal to supply;
- tying;
- discrimination;
- exclusionary conduct;
- self-preferencing.
The DMA adds a more regulatory approach to major digital gatekeepers.
United States
The principal framework includes:
- Sherman Act §1;
- Sherman Act §2;
- Clayton Act;
- merger enforcement.
U.S. doctrine generally requires careful proof of competitive harm and is cautious about imposing mandatory access obligations.
United Kingdom
The Competition Act 1998 and UK merger-control regime provide tools for examining:
- abuse of dominance;
- exclusionary conduct;
- digital platform concentration;
- strategic acquisitions.
The UK's digital-markets regime also provides a more ex ante approach for firms designated with substantial and entrenched market power.
India
The Competition Act 2002 is relevant to:
- abuse of dominant position;
- denial of market access;
- discriminatory conditions;
- leveraging;
- tying/bundling;
- combinations.
Knowledge infrastructure can therefore become relevant where control over data or digital gateways creates an appreciable competitive advantage.
15. Merger-Control Risks
Knowledge-graph concentration may arise through acquisitions rather than conduct.
A large platform could acquire:
niche database + scientific repository + mapping dataset + identity-resolution company + AI retrieval provider.
Individually, each acquisition might appear modest.
Collectively, however, they could create a global epistemic infrastructure conglomerate.
Competition authorities should therefore consider:
- data complementarities;
- future competition;
- innovation competition;
- interoperability;
- access foreclosure;
- ecosystem effects;
- acquisition of nascent competitors;
- control of unique datasets.
Traditional turnover thresholds may be insufficient where strategically important data assets generate little current revenue.
16. Data Portability As A Remedy
One potential remedy is enhanced portability.
Users or businesses could move:
- entity identifiers;
- metadata;
- records;
- relationships;
- annotations;
- historical information
between systems.
However, portability alone may not solve the problem where the dominant firm possesses unique proprietary data that cannot be replicated.
17. Interoperability Remedies
A stronger remedy may require technical interoperability.
Possible obligations include:
- open APIs;
- standardized identifiers;
- common metadata formats;
- machine-readable exports;
- interoperability protocols;
- non-discriminatory access;
- provenance standards.
The objective is not necessarily to make every proprietary graph completely open.
Instead, the objective is to prevent proprietary architecture from becoming an artificial barrier to competition.
18. Data Access Remedies
Competition authorities may consider:
FRAND-style access
Access could be provided on:
- fair;
- reasonable;
- non-discriminatory
terms.
Data trusts
Sensitive or commercially important datasets could be administered by independent intermediaries.
Clean rooms
Competitors could access specific information without obtaining commercially sensitive underlying datasets.
API access
Controlled technical access may allow downstream innovation without transferring the entire database.
19. Epistemic Infrastructure And Consumer Welfare
Consumer harm may occur even where prices remain zero.
Potential harms include:
- reduced informational diversity;
- lower search quality;
- reduced innovation;
- inaccurate information;
- manipulation of visibility;
- reduced choice;
- discriminatory ranking;
- weaker privacy;
- reduced ability to challenge dominant information sources.
Thus:
consumer welfare ≠ price alone.
For knowledge infrastructure, quality, diversity, reliability and contestability become economically significant.
20. Innovation Competition
A concentrated knowledge infrastructure can suppress innovation indirectly.
A start-up may be unable to compete because it cannot obtain:
- sufficiently comprehensive data;
- standardized identifiers;
- reliable entity resolution;
- API access;
- historical information;
- source metadata.
The incumbent may therefore not need to copy the entrant's technology.
It can simply prevent the entrant from obtaining the informational inputs necessary to scale.
21. Epistemic Monoculture
The most serious long-term concern is epistemic monoculture.
If many downstream systems rely upon the same knowledge source:
one graph → many AI systems → many applications → many decisions.
An error or bias in the upstream graph can consequently propagate through multiple markets.
From a competition perspective, this creates concerns about:
- systemic dependency;
- common-input concentration;
- correlated errors;
- innovation suppression;
- loss of informational diversity.
22. Relationship With AI Governance
Knowledge graphs may become part of the infrastructure used by:
- AI agents;
- autonomous economic systems;
- regulatory technology;
- financial AI;
- healthcare AI;
- scientific AI;
- procurement systems.
The entity controlling the knowledge layer could therefore influence the information available to autonomous systems.
This creates a novel competition concern:
control over the information environment of competing AI systems.
23. Possible Competition Theories
Future enforcement could develop around several theories.
Theory 1 — Knowledge Infrastructure Monopoly
A firm controls an indispensable structured information layer.
Theory 2 — Epistemic Gatekeeping
The firm controls which information becomes discoverable.
Theory 3 — Data Leveraging
The firm transfers upstream informational advantages into downstream markets.
Theory 4 — Graph-Based Self-Preferencing
The firm uses its graph to favour affiliated products.
Theory 5 — Interoperability Foreclosure
The firm prevents competing systems from communicating effectively.
Theory 6 — AI Input Foreclosure
The firm restricts access to information required for competing AI models.
Theory 7 — Epistemic Merger Concentration
Acquisitions combine complementary information assets into an infrastructure monopoly.
24. Regulatory Challenges
Competition authorities face several difficulties.
Defining the relevant market
Knowledge infrastructure does not fit neatly into conventional product markets.
Measuring market power
Traditional market shares may underestimate:
- data advantages;
- API dependency;
- switching costs;
- network effects.
Separating accuracy from competition
A dominant system may legitimately rank information according to quality.
Competition law must distinguish legitimate quality improvement from exclusionary ranking.
Protecting innovation incentives
Mandatory access can weaken incentives to invest in proprietary data systems.
Privacy conflicts
Data-access remedies must comply with privacy and data-protection requirements.
International coordination
Knowledge graphs operate globally while competition authorities remain largely jurisdictional.
25. Future Direction Of Competition Law
The evolution is likely to move from:
market concentration
towards:
infrastructure concentration
and ultimately toward:
control over critical informational ecosystems.
Competition authorities may increasingly examine not merely whether one firm sells the dominant product, but whether it controls the information architecture upon which competing markets depend.
Conclusion
Global knowledge graphs represent a potentially fundamental layer of the digital economy. Their competitive importance derives from their ability to connect vast quantities of information and make that information usable by search engines, businesses, governments and AI systems.
The central competition-law risk is therefore not simply a conventional monopoly over a database.
It is concentration of epistemic infrastructure.
Where one undertaking controls data acquisition, entity resolution, graph construction, ranking, APIs and downstream applications, it can potentially influence both economic competition and the informational conditions under which competition occurs.
The principal legal lessons from Google Shopping, Google Android, Microsoft, Aspen Skiing, Trinko, Bronner, IMS Health and Magill are that competition law already contains many of the conceptual tools required to address this problem—self-preferencing, leveraging, refusal to deal, essential facilities, interoperability and exceptional access to proprietary information.
The emerging challenge is to adapt those doctrines to an environment in which knowledge itself becomes infrastructure. The decisive question for future antitrust enforcement will increasingly be:

comments