Our Methodology
The LAVF framework is a synthesis of peer-reviewed academic research on Generative Engine Optimization — not a proprietary black-box algorithm. Every dimension, signal, and weighting decision is traceable to published findings, primarily from Princeton's GEO research group (ACM SIGKDD 2024) and subsequent academic work.
This methodology describes our current production measurement system. Planned capabilities are described separately and are not included in the current dataset.
LAVF — LLMsNetwork AI Visibility Framework
AI citation behaviour is measurable and systematically improvable. Aggarwal et al. demonstrated this at ACM SIGKDD 2024, showing up to 40% improvement in AI visibility through structured content optimisation. LAVF operationalises this finding across five empirically grounded dimensions.
Framework Dimensions
Authority
Does AI trust your brand as a credible source?
Authority captures the signals AI models use to determine whether a source is trustworthy and worth citing. Unlike PageRank-based domain authority, AI citation authority is shaped by earned-media presence, brand search volume, editorial coverage in authoritative publications, and the quality of inbound links from sources that AI models are known to draw from.
Measured Signals
- ·Brand search volume (correlates 0.334 with AI citation rate)
- ·Earned-media mentions in AI-trusted publications
- ·Inbound link quality from cited domains
- ·LinkedIn and professional-platform presence
Citation Quality
When AI cites you, is it accurate and consistent?
A brand being mentioned by AI is not the same as being cited accurately. Citation Quality measures factual fidelity (does the AI describe you correctly?), citation consistency (does the same brand appear across equivalent queries?), and absence of hallucinated or misattributed claims. Research shows 25–30% inconsistency in LLM citations across semantically equivalent queries — making this dimension critical for brand integrity.
Measured Signals
- ·Factual accuracy of AI-generated brand descriptions
- ·Citation recurrence rate across equivalent queries
- ·Absence of hallucinated claims or misattributed data
- ·Consistency across repeated observations within the current measurement system (Perplexity Sonar Pro)
Content Structure
Is your content structured for LLM consumption?
LLMs do not read content the same way humans or traditional crawlers do. Content Structure measures whether your web presence uses formatting, schema markup, clear heading hierarchies, and semantic organisation that AI models are statistically more likely to retrieve and cite. GEO-SFE research identifies specific structural features that measurably increase citation rates — this pillar operationalises those findings.
Measured Signals
- ·Heading hierarchy and semantic HTML usage
- ·Schema.org structured data completeness
- ·Formatting conventions (lists, tables, definitions)
- ·Answer-shaped content aligned to common AI query patterns
Knowledge Coverage
How broadly does AI know what you do?
Knowledge Coverage measures the range of topics, use cases, and query intents for which AI models consider a brand relevant. A company with deep coverage on its core category but no coverage on adjacent topics will be under-cited relative to competitors with broader topical presence. Research shows that higher-traffic sites achieve 6.4× more citations per query.
Measured Signals
- ·Topic breadth across core and adjacent categories
- ·Query intent coverage (informational, commercial, navigational)
- ·Recency of content indexed by AI training data
- ·Topical depth relative to category competitors
AI Accessibility
Can AI systems actually reach and process your content?
AI Accessibility addresses the technical and structural barriers that prevent AI crawlers, indexers, and retrieval systems from accessing content. This includes robots.txt configuration for AI bots, page speed, JavaScript rendering issues, and LLM-friendly content delivery. Notably, 88% of Google AI Mode citations come from outside the regular SERP top results — meaning technical accessibility is a distinct ranking factor from traditional SEO.
Measured Signals
- ·AI crawler access (robots.txt, crawl budgets for AI bots)
- ·Core Web Vitals and page speed
- ·JavaScript dependency and server-side rendering
- ·Content accessible without authentication or paywalls
How Scores Are Calculated
LAVF scores are a structured synthesis of publicly available research — not outputs from a proprietary model. Each step in the scoring process is grounded in empirical findings.
Pillar-level signals identified
Each LAVF pillar maps to a set of measurable signals drawn directly from GEO research literature. Signals are selected based on empirical evidence of correlation with AI citation rates.
Signals weighted by research evidence
Signals supported by stronger empirical evidence (e.g., the Princeton ACM SIGKDD 2024 findings) receive higher weighting than signals from early-stage or single-study research.
Pillar scores normalised (0–100)
Each pillar produces a normalised score from 0 to 100, benchmarked against the distribution of scores observed across the LLMs Network directory.
Composite LAVF Score calculated
The overall LAVF Score is a weighted composite of the five pillar scores. Weights reflect the relative strength of research evidence for each dimension.
Scoring Framework
Each LAVF assessment produces a composite score from 0 to 100. Score bands indicate relative AI visibility positioning — not an absolute measure of citation probability.
Leading
Strong, consistent presence across query types in the current measurement system.
Established
Solid visibility with room to extend coverage into adjacent topics and query types.
Developing
Emerging presence; key authority and structural gaps identified.
Early Stage
Limited AI citation presence; foundational signals require attention.
Scores are subject to change as LLMs update their training data and citation behaviour. LLMs Network scores reflect observed AI citation patterns and do not constitute a guarantee of future visibility or business outcomes.
Data Collection
Visibility assessments draw from four categories of data sources.
Public web data
Publicly accessible web pages, structured data, and content signals that AI crawlers and retrieval systems can access.
AI model queries
A fixed set of 20 queries is submitted daily to Perplexity Sonar Pro. We measure brand appearances and observed citation sources over time. At present, Perplexity Sonar Pro is the only production measurement system. Support for additional AI systems is planned but is not yet part of the published methodology.
Third-party datasets
Industry research from SparkToro, SE Ranking, BrightEdge, Ahrefs, Moz, and ConvertMate, cited with source and date wherever referenced.
GEO research literature
Peer-reviewed papers from arXiv and ACM, primarily from the Princeton GEO research group and subsequent academic work in the field.
Limitations & Transparency
Openly disclosedWe believe methodological transparency is essential in an emerging field. The following limitations apply to all LAVF assessments.
AI models change frequently
LLM training data, retrieval systems, and citation behaviour evolve with each model update. LAVF scores reflect a point-in-time measurement and should be re-assessed regularly.
Citation inconsistency is inherent
Research confirms 25–30% inconsistency in LLM citations across equivalent queries. No methodology can fully account for this stochastic variability — scores are probabilistic, not deterministic.
Training data opacity
AI model training datasets are not publicly disclosed. We cannot directly verify what content each model has indexed. Our signals are proxies derived from published research, not direct model inspection.
No guaranteed outcomes
Improving LAVF scores indicates improved conditions for AI citation — it does not guarantee citation. AI model behaviour is complex, contextual, and not fully predictable.
One measurement system only
All production measurements are collected from a single AI system (Perplexity Sonar Pro) using one fixed query set. Results should not be read as representative of other AI systems.
Responses are truncated
API responses are capped at 500 tokens. Companies mentioned later in a long answer are structurally excluded from our counts. This is a directional bias, not random error — an absence in our data does not establish that a company was not mentioned.
Coverage is a fixed list
We track a fixed list of 71 companies. Companies outside that list may appear in AI answers without being counted.
Framework is a synthesis, not a proprietary algorithm
LAVF is a structured synthesis of publicly available academic research. It is not a trade-secret algorithm. We publish our methodology transparently so practitioners and researchers can scrutinise and improve it.
Explore the Directory
Browse AI companies currently tracked in LLMs Network, or submit your company for listing and visibility assessment.
LAVF is a synthesis of publicly available academic research. LLMs Network does not guarantee specific AI visibility outcomes. Scores are probabilistic assessments, not deterministic rankings.