Methodology

How the Primacy Score is measured.

A score only means something if it is measured the same way every time. A number you cannot trace back to a method is an opinion. A number with a published, repeatable method behind it is a measurement.

So here is exactly how the Primacy Score is produced. Nothing here is proprietary mystery. The rigor is the point.

What the Score answers

Three questions at the center, six components underneath.

When someone asks an AI tool for a recommendation in your category, three things can happen. Your brand shows up, or not. It is backed by a source, or mentioned in passing. It is the actual recommendation, or just one name in a crowd. The Primacy Score measures all three, across the AI tools people actually use, and gives you one certified, trackable number — reported with the two breakout reads behind it: a Discovery Score, measured only on the questions that don't name you (the search a new customer runs), and a Reputation Score for how you're treated on the questions that do. The reads are never blended into each other, so the headline can never hide the read that predicts new demand.

Presence

Do you appear at all when buyers ask?

Citation

When you appear, is a source attached, the kind of backing that makes an AI tool trust and repeat your name?

Prominence

Are you the recommendation, or just present in the background?

Underneath those questions, the Score is built from six measured components. The first three are the pillars above; the other three capture the context that decides whether your presence is actually winning you business.

  • 01Presence RateHow often you appear in relevant answers.
  • 02Citation RateHow often, when you appear, a source is attached.
  • 03ProminenceWhether you are recommended, listed, or passed over.
  • 04Competitive Share of VoiceHow much of the recommendation space you hold versus named rivals.
  • 05SentimentHow favorably you are described when you come up.
  • 06Engine BreadthHow evenly you show up across AI tools, rather than carrying one and missing the rest.

Your live report shows all six, plus the exact weights, so you can see what is driving each read. All six feed Discovery; Reputation reads citation, prominence, and sentiment on the questions that name you.

How we measure it

A standard, not a screenshot.

EVERY

line

you sell is an answer surface, and your question set covers all of them, across awareness, consideration, decision, and trust. Set and locked with you at setup. Certification asks whether the complete set ran, not how large it was.

MEASURED ON

8

AI tools on a certified Score, none hidden: the six core tools (ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, Perplexity) plus Copilot Search and Grok, reported alongside. Shown per tool, unweighted.

SAMPLED

3-7

samples per question, allocated by disagreement rather than bought as a flat number: every cell starts at three, and only the cells where engines disagree, and where the disagreement could move the score, go deeper. Depth is always disclosed. A single benchmark can produce well over two thousand individual responses.

Real buyer questions

We ask the questions a real buyer asks.

Your question set follows what you sell rather than a fixed number: each product line or service is an answer surface, and the set covers every one of them. Certification asks whether the complete set ran, not how large it was. The questions span the full path a buyer travels, from early awareness (how do I choose between motor types), through consideration (who makes custom servo motors for medical devices), to decision and trust (most reliable brand, is this company reliable). The exact set is chosen and locked with you during setup, so the test does not drift between cycles. Need more coverage for extra products, regions, or audiences? Expansion packs add questions on top.

Many times, not once

We ask many times, not once.

A single check is a lucky screenshot. AI tools give different answers to the same question depending on the moment, so one look tells you almost nothing. We ask each question many times and average the results. How many is not a fixed number, and it is not the same for every question: sampling starts at three and deepens only where the engines disagree with each other, because a question every engine answers the same way does not get more true by asking it seven more times. Depth goes where the disagreement is. The range printed beside every score covers sampling variation only — how much the number could move because each question was asked a limited number of times rather than infinitely many. What we disclose is what was actually collected and whether every material disagreement was resolved, never a sample count that measures how much computing was purchased.

Two reads, never blended

Discovery and Reputation are separate numbers.

Appearing in an answer when the question already names you is the floor, not an achievement. So the report splits. The Discovery Score is computed only on unbranded and competitor-comparison questions — the surface where a new buyer actually shops — using all six components. The Reputation Score is computed on the questions that name you, from the quality components only: citation, prominence, and sentiment, reweighted 40/40/20. They are reported side by side and never blended into each other, because a blend is how an excellent brand nobody can find ends up with a passing grade. Your headline Score reports with both reads beside it, so it can never hide the read that predicts new demand.

A floor, not a loophole

Branded questions cannot inflate Discovery.

Branded questions are excluded from the Discovery read entirely, so the number a new buyer's search yields cannot be lifted by asking more questions that already name you. Be plain about the other direction: the headline composite is a blend across ALL prompts, branded and unbranded, so a branded-heavy set CAN lift it. That is exactly why the composite is never reported alone and why every report prints Discovery beside it. Failing to appear when the prompt names you is never averaged into either read: it is surfaced as a standalone red-flag finding. A certified set must contain a minimum share of Discovery-eligible questions, at least two thirds of the set, and a set below that floor cannot be approved or run. Every report states the number of Discovery-eligible questions behind its Discovery read.

Every major tool, none hidden

We test every major AI tool, and we hide none of them.

The Score covers the six tools buyers use most: ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, and Perplexity. Your report shows your result on each one, unweighted, so a strong showing on one tool can never quietly cover for a weak one. Each headline read gives more weight to the tools more people actually use, so it reflects real-world reach rather than treating a tiny tool the same as a dominant one.

Two more tools, reported never blended

Copilot Search and Grok are included, and reported beside the Score.

Every certified Score also reads Microsoft Copilot Search and Grok, at no extra charge. They appear beside your Score as separate numbers and are never folded into the headline. The certified composite is always computed on the same six core tools for every client, which is precisely what lets any two Primacy Scores be compared — including your own from last cycle.

Certification

Three tests, no shortcuts.

A measurement earns the name certified Primacy Score because it achieved a stated precision, not because a fixed number of samples was bought. All three of these must hold.

Certification, test one

We collected enough, and we tell you the bar.

Every planned question is asked on every engine. Engines do occasionally fail to answer, so the bar is stated rather than implied: a certified Score needs at least 95% of all planned answers collected, at least 80% on every single engine, and 100% on your protected-core questions — the ones you marked as the ones that matter. Miss any of those and the Score is downgraded rather than quietly computed from the survivors. We publish the numbers this rests on: the response total on your report reconciles against questions × engines × sampling, so you can check it yourself.

Certification, test two

We resolved every contested answer.

AI engines mostly agree about a brand: it is either clearly present in their answers or clearly not. Where they disagree, we automatically ask again, up to a hard cap, until the disagreement is settled or explicitly flagged. No contested answer is silently averaged away.

Certification, test three

The range is computed and printed, whatever it comes out at.

Every Score is published with a stated range, for example 58.4 plus or minus 3.2, meaning that if we ran the same measurement again the number would land inside it. Certification does not require the range to clear a bar. It asserts that the range was computed by the disclosed method and printed beside the number, so you can see for yourself how much the result can be trusted to hold still. A wide range is disclosed, never refused and never hidden.

Why not a sample count

A sample count measures spend. The range measures trust.

Notice what is not on the list: a sample count. We used to certify any measurement that ran ten samples per question per engine. Then we ran a calibration study across our own live measurements and found that identical sampling can produce different precision, because some question sets are simply noisier than others. A sample count measures how much computing was purchased, not what the measurement is worth. We briefly replaced it with a precision threshold, then retired that too: a threshold turns an honest wide range into a failure, which is an incentive to measure easy things. So certification now asserts that the protocol ran in full, and the range is published for you to judge.

Citation counts where you appeared

We ask whether the mention carried a source, not whether a source exists in the abstract.

Citation rate is measured across the answers you actually appeared in. Doing it the other way — dividing by every answer we collected — would record a zero for an answer where you were never mentioned at all, which asserts that you appeared and were not sourced. That did not happen. It also confuses two different problems: a brand cited well but mentioned half as often as a rival has a presence problem, not a citation problem, and a denominator of everything hides exactly that distinction.

Why some ranges are wider

A brand that appears less often can be measured less precisely.

Several components can only be measured where you actually appear: whether the mention carried a source, how prominently you were placed, how you were described. A brand that AI assistants mention rarely gives us fewer of those observations, so its stated range is wider. This is a property of the measurement, not a penalty, and we would rather say it plainly than have it look like inconsistency: a widely mentioned brand and a rarely mentioned one measured the same way will not carry the same range, and the rarely mentioned one is where a tight range would be least believable.

What the range means

The range is reproducibility, not accuracy.

The range beside the score measures how much the number could move because each question was asked a limited number of times rather than infinitely many. It shrinks predictably as sampling deepens, and we publish it at every depth. It does not mean the score is right. A badly chosen question set can produce a precise measurement of the wrong thing. That is why question sets are locked with the client and versioned, and why every cycle states on the report whether its exact configuration was frozen and can be re-derived. Precision and honesty of the instrument are separate disciplines, and Primacy Score publishes both.

Where the depth goes

Depth is spent where it resolves something.

Disagreement between engines concentrates in a narrow contested band, and that is where measurement actually needs depth. So sampling adapts: every question starts at three samples per engine, and only cells where engines disagree, and where the disagreement could move the score, escalate deeper.

The scorer, measured

We test the judge, not just the engines.

In Primacy's engine-calibration study (completed 2026-07-28 UTC, version 2026-07-cal1), each of 717 scored engine answers included a deliberately fictional brand among the judged subjects. The scoring model recorded 0 false appearances in 717 judgments (0.00%): it never attributed presence to a brand that does not exist. This measures false-appearance precision only. It says nothing about recall or the other judgment dimensions.

Comparability

What makes one Score comparable to another.

Two Scores in different categories are not on the same scale.

Each company runs its own question set against its own five named competitors, and one of the six components is share of voice, which is relative to whoever those competitors are. So a 64 in one category and a 72 in another do not mean one company is eight points behind the other. What compares across companies is the method, which is identical, and rank within each company's own competitive set: leading your field on citation rate means the same thing in any category. Portfolio owners should read the ranks and the gaps, not a league table of absolute Scores.

The method is versioned.

Every report is stamped with the methodology version, the scoring rubric version, and the question-set version. If we improve the method, the version number changes and you can see exactly what changed. Your old Scores stay readable against the method that produced them.

Bad news is surfaced, not averaged away.

Negative mentions, the moments an AI tool raises a real concern about you, are pulled out and shown on their own, never blended into the Score. A genuine reputation problem should never be hidden by good results elsewhere. You see the exact wording and which tool said it.

The contested ground is scored separately.

Alongside the headline, we report a Protected-Core Score: your standing on only the questions where you and named competitors are fighting for the same recommendation, with easy wins stripped out. It tells you whether you are winning where it is actually hard to win.

The Score is tied to outcomes.

Visibility is only worth measuring if it moves the business. Reports track AI-referred website sessions, conversions from those sessions, and AI crawler activity, cycle over cycle, so you can see whether a rising Score is reaching your site and your pipeline.

Honest about its edges

What the Primacy Score is not.

  • A measurement, not a guarantee. It describes how AI tools answer today, and that behavior shifts as the tools change.
  • A point-in-time reading with normal variation. That is exactly why we sample many times and publish cycles, so you read the trend, not a single twitch.
  • Not a replacement for your search tools. Use a platform like Semrush or Ahrefs for day-to-day search operations. Use Primacy Score for the one question those tools cannot answer cleanly: does AI actually recommend us, and if not, why.
  • Not a single number to be read alone. The Score points you in a direction. Discovery, Reputation, the six components, and the per-tool breakdown tell you what to do about it.

In short

Same buyer questions. Asked many times. Across every major AI tool. Reported with its Discovery and Reputation reads, never blended into each other. Surfaced fairly, scored against named competitors, tied to real traffic, and stamped with a version, so it holds up over time and against anyone else's Score.

That is what a standard requires. That is how we measure it.