An answer names a company. A source list links to its website. Another answer describes the product without naming it. All three may matter to a brand team, but putting them in the same column hides the difference between being visible, being used as a source and being recommended.
A workable definition begins with the decision the record will support. If the team wants to know whether its name appeared, count explicit name references. If it wants to understand attribution, record citations separately. If it wants to investigate product accuracy, extract the factual claims. One answer can contain several of these signals at once.
The taxonomy below is a proposed working method, not an industry-wide standard. All example companies and answers are illustrative. Use the categories to build a consistent review process, then document any changes before comparing results over time.
Start with the observation, not the total
The basic record should be a single completed answer under known conditions. Save the exact question, answer text, visible links, collection date and interface. Record language and location settings when available. Note whether the answer followed earlier conversation, because a brand introduced in a previous turn may explain its presence later.
Do not reconstruct the original question from memory. “Which tools suit a small repair business?” and “Is Example Workshop suitable for a small repair business?” create different opportunities for a mention. In the second question, the brand has already been supplied. A mention there cannot be treated as equivalent to the system independently including it in a shortlist.
Keep the raw observation before classification. A screenshot can preserve presentation, while copied text supports review and comparison. If collection conditions are unknown, say so. An incomplete record can still be useful for investigating an error, but it should not quietly enter a tightly controlled measurement series.
Use separate labels for separate facts
An explicit mention occurs when the answer names the target entity in its own visible response. That includes a company name, a clearly identified product name or an unambiguous alias. Count the answer once for presence, even if the name appears several times, unless the metric is specifically intended to count occurrences.
A citation is a visible reference to a source. Record its URL and where it appears. A company’s domain can be cited without the company being discussed as a product choice. Conversely, a company can be recommended without a link. Those are separate fields, not competing definitions of one field.
A recommendation adds a selection judgment. “Example Workshop offers scheduling software” describes a company. “Example Workshop is a suitable choice for a small repair business because…” recommends it for a stated purpose. The recommendation may be qualified, conditional or one of several options. Preserve those conditions rather than reducing the answer to a positive sentiment label.
| Observation | Suggested field | What it establishes |
|---|---|---|
| The answer names the company | Explicit mention | The entity appeared in the answer |
| A visible link points to its page | Citation | The page was referenced |
| The answer suggests choosing it | Recommendation | It was proposed for a particular need |
| The answer describes a capability | Factual claim | A statement can be checked against evidence |
| The name appears only in the question | Aided question | The user supplied the entity |
Handle descriptions without a name carefully
A passage may resemble a brand’s positioning without identifying it. That is not enough to call it a mention. Several companies can offer similar capabilities or use similar phrases. Counting inferred references can inflate a report in ways that a second reviewer cannot reproduce.
If an answer gives a unique product URL but no visible name, the record can carry a source-domain match. If a reviewer believes an unnamed description identifies the company, store that judgment separately with its rationale. Leave the explicit-mention field unchanged. This preserves the difference between observable text and interpretation.
Aliases need similar care. Build a list of known product names, former names and common abbreviations. For ambiguous short names, inspect nearby context. A word used in its everyday sense should not become a brand reference because a text matcher found it. Give uncertain matches a review state rather than forcing them into yes or no.
Read the citation before assigning credit
A link near a sentence does not automatically prove that the page supports every part of the sentence. Open the source when the relationship matters. Find the relevant passage, check its date and note whether the claim is actually present. If the page has changed since collection, record that limitation.
The W3C provenance overview is a useful conceptual reference for tracing origins and relationships. For a monitoring sheet, that can be as simple as connecting the answer record to the cited page and the reviewer’s support check. The important feature is an inspectable trail, not a complex vocabulary.
Consider an answer that says a fictional tool supports offline scheduling and cites its homepage. The homepage describes scheduling, but says nothing about offline use. Record the citation as observed and the offline claim as unsupported by that source. Do not erase the citation, and do not mark the claim verified merely because a link exists.
Separate tone from truth
A warm description can contain a factual error. A critical description can be accurate. Sentiment and accuracy therefore need different fields. This becomes especially important when an answer discusses a limitation, such as a product being available only in certain markets.
A useful accuracy review uses labels such as supported, contradicted, insufficient evidence and not checked. Require a source for supported or contradicted decisions. “Insufficient evidence” is a legitimate result; it tells the team what research is missing. It should not be rewritten as false merely because the claim is inconvenient.
NIST’s Generative AI Profile includes confabulation among the risks of generative systems. That is a reason to check material assertions, not a reason to assume every unfavorable statement is invented. The reviewer still needs to establish what the current evidence says.
Define the denominator before reporting a percentage
A mention rate needs a clear denominator. One workable definition is the number of completed eligible answers containing an explicit mention divided by all completed eligible answers in the defined sample. Label this as a sample result. It is not the percentage of all conversations in which the brand appeared.
Report failed or unavailable collections separately. If a scheduled run contains missing answers, readers should see the gap. Decide in advance how refusals and nonresponsive answers are handled. Otherwise two analysts can produce different rates from the same archive while both believe they followed the rules.
Keep aided and unaided questions separate in the report. Also separate markets or languages when the sample was designed to examine them independently. A combined number can be useful as a summary, but it should not conceal that one set directly names the brand and another asks an open category question.
Test the taxonomy with a second reader
Before automating, ask two people to classify a small batch independently. Compare disagreements. Was a recommendation merely a description? Did a source link get mistaken for a brand mention? Was an alias ambiguous? Revise the instructions where those disagreements reveal a missing rule.
Save a few agreed examples as a reference set. Include an obvious mention, a citation without a mention, a conditional recommendation and an unresolved entity match. When the rules change, version the guide and note whether historical records have been reclassified. A new rule can change a total even when no answer changed.
The first useful task is modest: select one brand, gather a small set of complete answer records and label each signal separately. Review the ambiguous cases before creating a chart. That work gives later monitoring a stable vocabulary and makes a discussion about “more mentions” much more precise.

