This website requires JavaScript.
> Back

The 2026 AI Reputation Barometer: the machine behind the measurement

Nearly 40,000 questions administered to four leading generative AI models, some 573,000 source citations traced and analyzed, nine reputation dimensions scored for every CAC 40 company: the 2026 edition of the AI Reputation Barometer, co-produced with Vectors, is a change of scale. That change has a name: PromptForge, the tool Trickstr built to analyze LLM responses at scale.

Measuring an algorithmic perception is an engineering problem

Asking a chatbot what it thinks of a company is not a measurement. The same question, asked twice to the same model, produces two different answers. An LLM's perception is not a fact you look up, it is a distribution you sample.

A stable measurement therefore requires volume and protocol: questions declined into multiple, neutral phrasings, spread across every dimension of corporate reputation, repeated, administered under controlled conditions and fully logged. An experimental protocol in the scientific sense of the term, not a conversation.

This is what separates AI reputation, a discipline of multidimensional measurement, from visibility-driven approaches such as GEO, whose goal is above all to appear in answers. The object here is to quantify how models perceive an organization, dimension by dimension, and to identify what drives that perception.

PromptForge, an end-to-end orchestrator

Trickstr has been interrogating models at scale since early 2025: its study on the political biases of AI already rested on 41,410 responses obtained from each of the 14 models tested, 579,740 responses analyzed in total. The first edition of the barometer followed in October 2025. Edition after edition, every step of the protocol has been industrialized, converging into a proprietary tool: PromptForge, now integrated into the Kosmos platform.

PromptForge orchestrates the full chain:

  • Protocol generation. Question templates are declined into tens of thousands of unique questions (company × dimension × phrasing), written so as not to lead the answer.
  • Mass administration. Questions are submitted to the different models in parallel, with throughput, session and context management, and full logging of every exchange. A campaign can be replayed identically, the precondition for any comparison over time.
  • Structured collection. Answers and cited sources are captured, normalized and archived. For the 2026 edition, nearly 573,000 source citations were traced and classified by type (owned websites, media, forums and blogs, encyclopedias, social networks).
  • Scoring and semantic analysis. Each answer is scored on the nine dimensions of the framework, entities are extracted (executives, brands, controversies), and statistical aggregation produces rankings, correlations and weak signals.

2026: the change of scale

The October 2025 edition rested on 4,536 questions and seven dimensions. The 2026 edition administers nearly 40,000, almost nine times more, and extends the framework to nine dimensions. The difference is not cosmetic, it is statistical: at this scale, confidence intervals tighten, gaps between companies become significant, and the ranking, led this year by Schneider Electric, gains a robustness that artisanal approaches cannot offer.

Above all, volume makes visible phenomena that no manual polling of models can detect.

What only volume reveals

Four examples from the 2026 edition:

  • The frozen memory of models. Analysis of the freshness of cited sources shows knowledge massively concentrated on the 2020-2024 window. AIs judge companies on a past frozen at training time.
  • Awareness is not reputation. The volume of mentions a company generates and the quality of its perception diverge sharply. Being known to the models does not mean being well rated by them.
  • Media overexposure comes at a price. Counterintuitively, a company's media over-representation relative to its economic weight correlates negatively with its ranking.
  • The hyperconcentration of executives. A handful of leaders captures most of the executive mentions produced by the models, leaving most others in an algorithmic blind spot.

Each of these findings requires processing hundreds of thousands of answers and citations through an identical analytical chain across models. That is precisely what PromptForge makes possible.

A tool proven on global stakes

The barometer is PromptForge's public showcase, not its only field of operation. The tool is deployed for major global companies, among them L'Oréal, Carrefour and Vinci: multi-market, multi-language perception mappings, measurement of image in the models before and after strategic sequences, competitive benchmarks, identification of the sources that shape AI answers on sensitive issues.

With every new generation of models, perceptions shift. The question is no longer whether AIs hold an opinion on organizations: they do, argued, sourced, and consulted more every day. The question is how to measure it rigorously, at regular intervals, with a constant instrument. That is what PromptForge is for, and the reason this barometer exists.