Request Proposal Toggle Menu

Go Fish Research: What 4,599 ChatGPT Responses Reveal

Go Fish Research: What 4,599 ChatGPT Responses Reveal featured cover image

We Asked ChatGPT the Same Set of Questions Every Day for Two Months. Did the Answers Stay the Same?

Go Fish Digital analyzed 4,599 ChatGPT responses over two months. We measured how often recorded mentions changed and which names stayed most visible within a fixed set of questions.

Key findings

Most questions focused on choosing providers, understanding services and comparing alternatives. Here is what happened to the recorded mentions:

  • Mention lists often changed substantially from one day to the next. In 60% of comparisons with at least one previous-day mention, at least half of those mentions were no longer recorded the next day.
  • The leading group stayed consistent over two months. The same five names ranked in the top five during both the first 31 days (days 1 to 31) and the final 32 days (days 32 to 63) of the study.
  • One day’s result did not tell the whole story. A name could disappear from an individual mention list and still remain among the study’s most-mentioned names overall.

How much can one day’s ChatGPT answer tell you?

A company appears in a recorded mention list today and is missing from it tomorrow. Before treating that omission as a decline, you need to know how often the same question produces a different list and how much of that list usually changes.

We examined ChatGPT’s responses to the same 73 questions asked every day for two months. Our research asked: When ChatGPT receives the same questions each day, how much do mentions change, and do the same names keep leading?

The daily lists often lost a large share of their previous mentions. Yet the top five names were the same across both halves of the study. An individual answer and a two-month leaderboard told different parts of the story.

How we ran the study

We analyzed 4,599 ChatGPT responses over two months: 73 fixed questions × 63 days, with one recorded response per question per day. Comparing each question’s responses on consecutive days gave us 4,526 comparisons.

We used one anonymized client’s question set in one region to keep the business context consistent. This helped us examine daily variation without mixing different clients’ markets and questions. It did not isolate model updates or other causes of change.

We counted the names recorded for each ChatGPT response in our dataset, with each name counted once per response. These recorded mentions included companies, tools, platforms and some generic labels. We checked name matches across all saved answers and reviewed a small diagnostic sample of 31 responses. This sample could not establish an overall error or recommendation rate. A mention does not necessarily mean ChatGPT recommended it.

What kinds of questions did we ask, and how consistent were their mentions?

Most questions fell into three categories:

  • Commercial: finding providers, comparing costs and making hiring decisions.
  • Informational: understanding services, implementation and measurement.
  • Competitor: comparing providers or looking for alternatives.

Each question had one ChatGPT response per day throughout the study, and its category stayed the same.

CategoryQuestionsComparisons losing at least half of prior mentionsTop-five names shared across study halves
Commercial481,808 / 2,796 (64.7%)5 of 5
Informational19512 / 941 (54.4%)4 of 5
Competitor4105 / 248 (42.3%)5 of 5

The loss calculation includes only comparisons where the previous day’s response had at least one recorded mention. Leader overlap compares each category’s top five across the two halves of the study.

Commercial questions had more substantial daily mention changes than informational questions, but their leading group remained consistent. Commercial questions also made up roughly two-thirds of the panel, giving them the greatest weight in the overall results.

These differences describe this question set. The competitor group contained only four questions, three of which already named a provider, potentially influencing repeated mentions.

The full analysis also included one question labeled “branded” and one without a category.

Finding 1: Daily changes often affected a substantial share of mentions

Counting any change treats a one-name substitution and a completely different list equally.

Consider an illustrative response with ten mentions. If nine return tomorrow and one is replaced, the list retains 90% of its previous mentions. It still counts as changed. To see whether changes were small or extensive, we measured the share of yesterday’s mentions missing from today’s list.

Of the 4,108 comparisons with at least one previous-day mention, 2,464 lost at least half of those mentions: 60.0%. The other 418 comparisons started with an empty mention list. We excluded them from this calculation because there were no previous mentions from which to calculate a loss percentage.

Figure 1. All categories use 4,108 comparisons of ChatGPT responses with previous-day mentions. “Lost” means absent from the next mention list recorded in the dataset. Retaining every previous mention still allows new mentions to be added.

Only 10.6% retained every previous mention. Another 29.4% lost some but fewer than half, 44.4% lost at least half but not all, and 15.6% retained none. A response in the first category could still add new mentions. Keeping every previous mention and keeping an identical list are different measures.

We then counted how many previous-day mentions returned across the whole dataset. Each name recorded for each earlier response counted once. Of 27,106 prior-day mention instances, 12,062 reappeared for the same question the following day: a pooled retention rate of 44.5%. Every prior mention has equal weight in this calculation. It is different from averaging each response’s retention percentage.

Across all 4,526 daily comparisons, 91.3% involved at least one addition or removal. That figure includes small edits. The 60.0% result answers the more specific question of how often at least half of yesterday’s mentions were no longer recorded.

We chose the half-loss threshold to describe the size of changes. It is not an industry benchmark for acceptable volatility. “Missing” also means absent from the next mention list. It does not confirm absence from the answer text, as the export checks in Methodology and limitations show.

Finding 2: The same leading names persisted across both halves of the study

We next ranked names by the number of ChatGPT responses with a mention across all 73 questions. Did the leaders change as much as the individual lists?

We divided the study into two separate periods: the first 31 days, containing 2,263 responses, and the final 32 days, containing 2,336 responses. Each period received its own ranking based on mention counts.

The same five names occupied the top five in both periods. Nine of the ten top-ten names were also shared. Some changed positions. The leading group stayed similar even though its order did not stay fixed.

Figure 2. Each week contains 511 ChatGPT responses. Names A to E are the five most-mentioned exact names across the full study. The chart shows their weekly mention rates; the separate rankings establish the top-five overlap. Mentions are not recommendations.

The two most-mentioned names appeared in 32.7% and 32.2% of all responses, respectively. Each stayed in the daily top five throughout the two-month study. They did not need to appear in every answer, or every answer to a particular question, to lead the overall results.

These percentages use all responses as the denominator. Because an answer can mention several names, their rates do not add up to 100%. They measure frequency within our questions, not market share or purchase preference.

The first week compared with the remaining eight weeks

The first week’s five leading names were also the top five across the remaining eight weeks combined. The group stayed the same, but two names exchanged positions. This does not mean those names held the top five spots in every individual week.

The first seven days contained 511 responses. The remaining 56 days contained 4,088. Every question had equal coverage within each period, and none of the first week’s answers appeared in the later comparison.

ComparisonTop-five members sharedTop-ten members shared
First 7 days versus remaining 565 of 59 of 10
First 31 days versus final 325 of 59 of 10

Figure 3. Rankings across 511 first-week responses and 4,088 responses from the remaining eight weeks combined. Names A to E match Figure 2. The same five names led both periods, although their order changed.

Name D moved from fifth to third, while Name E moved from third to fifth. Names A, B and C kept their positions. Only one member of the top ten differed between the two periods.

Keeping these periods separate matters. Comparing the first week with the entire study would include the same 511 answers on both sides. Here, the later results came entirely from subsequent days.

This was a retrospective comparison. We did not run a forecasting experiment or test how many days every business needs to monitor. It does not establish that seven days is sufficient.

Look for the pattern over time

Seeing a different list after repeating a ChatGPT question is only part of the picture. In our study, daily mention lists often changed while the same names led over longer periods. One snapshot can flag something to investigate. The pattern across days and weeks helps show whether a missing mention is an isolated result or a sustained drop.

Finding 3: Visibility tracking needs more than a single answer

A name missing from the recorded mentions for one ChatGPT response does not, by itself, show that visibility is declining. In this study, individual mention lists changed frequently while the same names continued to lead over longer periods.

To understand performance, track how often a name appears across the same questions over time. Check whether a drop continues across repeated observations or is limited to a particular day.

Rankings alone are also incomplete. A name can stay near the top while appearing less often, so track its mention rate alongside its position. A consistent leading group does not mean its visibility percentages stayed unchanged.

Finally, check the questions that matter most to the business. An overall score combines results from many questions and can hide differences between them. Track both the wider pattern and those individual question histories to understand where visibility is holding, improving or declining.

Keep the question set consistent, and record the region and any available ChatGPT model or session settings. If questions are added or removed, check how that affects the result before interpreting a change in visibility.

Separate questions about choosing or comparing providers from questions about understanding a topic. When reporting daily changes, distinguish mentions retained, lost and added so that a one-name substitution is distinguishable from losing most previous mentions.

Methodology and limitations

We analyzed all 4,599 saved ChatGPT responses across 73 fixed questions and 63 days, with no missing question-days or duplicate question-day records. We excluded one filter footer per export and filled in no missing values. The 10,517 summary rows described these same responses. Every row matched our recalculations at its displayed precision. Separate Python calculations agreed on 25 checks, and change-size counts were independently verified.

Our mention rates included all 73 daily responses. Summary visibility rates included only responses with recorded mentions, using 58 to 72 per day. Exact names were counted separately, so spelling variations may remain separate.

Recorded mentions can differ even when saved answers match. In 30 groups containing 87 observations, the same question had identical saved answers on multiple days. These occurred within the first six days and had distinct run identifiers. Comparing every pair within these groups produced 98 comparisons; 22 had different mention lists. The cause is unknown. These comparisons reused responses, so this is not an overall error rate. Some answers with empty mention fields also contained names. Our findings therefore measure recorded mentions, not confirmed changes in answer text.

Adjacent-day comparisons reuse responses and are not independent trials. With one saved response per question per day, we cannot separate variation between runs from changes over time. ChatGPT model versions, session settings and collection retries were unavailable. We chose thresholds and comparison periods after collecting the data.

These results cover one client’s selected question set during this period. They do not establish patterns for other clients, industries or platforms, what caused a mention, or its effect on purchases.

Dataset identities remain anonymized. The analysis code, change-magnitude script, prompt-category script, analysis log and text-validation method document the calculations and limitations.

About Ashish Jacob