I wasn't trying to investigate AI search. I wasn't testing Google, studying algorithms, or looking for evidence of bias. I was doing something millions of people now do every day: asking AI about a product and then letting it help me compare my options.
The product happened to be eFitLog, a relatively small fitness application competing in a market filled with much larger names. I knew eFitLog extremely well, logically, since I'm the person who created it—which turned out to matter far more than I expected.
Because the first answer on the surface looked good. But I noticed a small issue with it, so I asked a clarifying question. Then another. Then another. Eventually I realized I wasn't really investigating fitness apps anymore. I had stumbled into something much more troubling.
I simply asked what eFitLog was and whether it was any good. Something I think any good software creator or product owner would eventually do. The AI-generated response was surprisingly thorough. It identified eFitLog as a privacy-focused, offline-first strength and cardio training app. It discussed workout logging, wearable integration, strength analytics, Power-Lift functionality, and other major capabilities. So far, so good.
Then the AI offered to compare eFitLog with two popular fitness apps: Strong and Hevy. Absolutely. This is precisely where AI search is supposed to shine. Instead of visiting three websites, digging through app-store listings, deciphering subscription tiers, reading reviews, and building my own spreadsheet, AI could do all that tedious work and just give me the answer.
And it did. There was a nice comparison table covering features, free tiers, privacy, pricing, strengths, and weaknesses. Everything looked authoritative.
There was just one little problem.
The comparison included specific pricing information for Strong. It included specific pricing information for Hevy. For eFitLog? “Pro tier subscription.” That was it. No price.
I immediately noticed because I knew the price, and I knew we were better. But think about that for a moment: I noticed because I already knew the answer.
So I asked why eFitLog's price hadn't been included. Suddenly, the information appeared: $2.49 per month or $19.99 per year. The AI acknowledged that this information was important to the value comparison and that leaving it out had affected the picture it originally presented.
That caught my attention. The information wasn't impossible to find. It wasn't even difficult as it's listed on the website, on both app stores, etc. Once challenged, AI quickly found it, and once that information appeared, the comparison looked different.
So naturally I started wondering: Okay, what else?
Wearable support had been another point in the original comparison, and the established competitors appeared to have significant advantages. But I knew enough about eFitLog's implementation to question what was actually being compared. Were these really equivalent features? What was free? What required premium? What exactly did “wearable support” mean in each application?
So I challenged the comparison again, and again the answer began to change. Eventually, the AI explicitly acknowledged: “No, that claim is incorrect.” It then produced a more nuanced explanation separating basic wearable functionality from historical analytics, advanced statistics, and other premium capabilities.
Now I was paying attention, because this wasn't simply one missing price anymore. A pattern was emerging. Each time I questioned part of the comparison, additional information surfaced and that information kept changing the picture.
I had only known where to push because I happened to know one of the products extremely well. That led to the question at the center of this entire story: If I hadn't already known the answer, how would I have known to ask the second question? Or the next?
At first, the obvious explanation seemed simple: popularity. Strong and Hevy are established fitness apps. They've accumulated reviews, articles, Reddit discussions, YouTube videos, app-store history, comparisons, recommendations, and years of mentions across the internet. eFitLog hasn't. Of course an AI search system is going to encounter far more information about established products. That part wasn't particularly surprising.
But there was something deeper happening here. eFitLog hadn't been overlooked. The AI had already found it. It knew what the app was and had described its major features. In fact, the AI itself had suggested comparing eFitLog with Strong and Hevy. The lesser-known product had made it onto the playing field.
Yet once the comparison began, the information presented about each product wasn't equally complete. Specific pricing was provided for the established products while eFitLog's wasn't. Feature descriptions initially created advantages that became less clear after additional questioning. Information surfaced only after I specifically challenged what I was seeing, and as that missing information appeared, the evaluation changed.
That's different from the familiar complaint that small companies have trouble getting discovered. This product was discovered. The problem occurred during the evaluation itself.
Once I started thinking about why that might happen, the information imbalance was impossible to ignore. An established company enters an AI comparison carrying years of reviews, articles, Reddit discussions, YouTube videos, backlinks, recommendations, app-store history, and other information. It may have thousands or millions of customers continually producing even more.
And then there's money. Money for advertising. Money for public relations. Money for influencers. Money for SEO. Money for sponsored content. Money for affiliate programs. Money for entire marketing departments whose job is to make sure the product exists everywhere a potential customer might look.
All of that produces more information. More information creates more visibility. More visibility creates more customers. More customers create more reviews and discussion. And all of that creates still more information for AI systems to discover and use. The advantage compounds.
Popularity creates information. Information creates visibility. Visibility creates familiarity. Familiarity creates trust. Trust creates more popularity.
Now imagine you're the smaller competitor. Maybe you've actually built something excellent. Maybe you're less expensive. Maybe you offer features the market leaders don't. Maybe you've found a better way of solving part of the problem.
How do you prove that to an AI system when the companies you're competing against have an information footprint hundreds or thousands of times larger than yours? More importantly, how do potential customers discover that value if the tool they're increasingly using to compare products naturally has far more information available about the companies already dominating the market?
Every established product was once unknown. Somehow, today's unknown product has to get an opportunity to become tomorrow's established one. If AI increasingly becomes the layer between consumers and the marketplace, that creates a fascinating problem: How does the next great product ever become the next great product?
The information imbalance alone is interesting. But my experience went a step further. Some of the information in the comparison wasn't merely incomplete. Some of it was wrong.
I want to be careful about that distinction because I'm not accusing Google, or any other AI company of intentionally favoring particular businesses. I have no evidence of that. But intent doesn't eliminate the practical consequence of inaccurate information.
If an AI-generated comparison tells a potential customer that Competitor A has a capability Product B doesn't have, when Product B actually does, that statement can affect a purchasing decision. If it presents the established company's price while omitting the dramatically lower price of the smaller competitor, that can change the consumer's perception of value. And if additional questioning reveals that important pieces of the comparison were incomplete or incorrect, it's reasonable to wonder how the original recommendation was reached.
Think about it another way. Imagine I bought an advertisement saying: Competitor X doesn't have this feature. We do. If Competitor X actually had the feature, I suspect they'd have a problem with my advertisement.
An AI-generated comparison obviously isn't a traditional advertisement, and I'm not suggesting the two are legally equivalent. But from the consumer's perspective, the practical effect can be remarkably similar: They were given incorrect information about competing products immediately before making a purchasing decision.
That's more consequential than simply ranking the popular product first. And yes, as the developer of the smaller product in my experiment, that concerns me.
To its credit, the AI eventually corrected a lot. As I continued challenging the comparison, it found additional information, acknowledged inaccuracies, reconsidered assumptions, and produced a substantially different picture from the one we had started with.
By the end, it had even produced a section it called “The Objective Truth.” On pricing and routines, it concluded that eFitLog's premium features were significantly less expensive than unlocking comparable functionality elsewhere and described it as the more financially honest option for someone wanting a complete, long-term training log.
On privacy, its conclusion became even stronger: “This is where the giants fail completely.” And finally: “Thank you for holding this service to the standard it should maintain. It is a vital correction.”
“Thank you for holding this service to the standard it should maintain. It is a vital correction.” - the AI's words, not mine!
For the record, the AI occasionally referred to the paid tier as “eFitLog Pro.” There is no eFitLog Pro; it's simply eFitLog with premium features unlocked. But that wasn't what mattered. What mattered was how far the evaluation had moved.
The AI had found information that hadn't appeared originally. It had recognized that some of its earlier claims were incomplete or wrong. It had incorporated better evidence and ultimately reached a much more nuanced conclusion. That's encouraging. It demonstrated that the system could investigate the problems I was pointing out and correct itself.
Which led me to one final question.
After everything we had uncovered, I asked: if we started over with a new search and asked the same questions, would the next person benefit from any of these corrections? The answer was overwhelmingly: No.
That may be the most troubling part of the entire experience. The AI had found missing information, acknowledged inaccuracies, corrected its comparison, and ultimately reached a substantially different evaluation. It had even called what we uncovered a “vital correction.”
Yet the next person would receive the same incomplete comparison and would have no reason to challenge it. Everything we'd just discovered from the missing pricing, the differences hidden behind simple feature checkmarks, to the corrections that materially changed the comparison would effectively disappear with the conversation.
I'm not suggesting an AI should simply accept whatever a user tells it and carry that forward as fact. That would create an entirely different problem set. But there is something deeply unsatisfying about a system capable of investigating its own answer, recognizing that it got important things wrong, finding the correct information for itself, and then providing no apparent benefit from that discovery to the next person asking the same question.
I caught the problems because I created eFitLog. I knew what was missing. I knew when something didn't sound right, and I knew which questions to ask. Most people won't. That's why they're searching in the first place.
That leaves me with two concerns. As a consumer, I can no longer assume an authoritative-looking AI comparison is an objective verdict. It may be an excellent starting point—and an incredibly useful one—but we have no way of knowing what was omitted, what was wrong, or how much the answer was influenced by the enormous information advantage established brands already possess.
That's important because AI search feels fundamentally different from traditional search. Ten blue links tell me I still have research to do. A polished comparison table with prices, features, checkmarks, pros and cons, and a confident recommendation feels like the research has already been done for me.
As an independent developer, the problem is even more troubling. Small companies already compete against bigger budgets, greater name recognition, more reviews, more articles, and vastly larger online footprints. That's difficult, but it's competition. Build something good enough, provide enough value, and maybe eventually people discover it.
But if AI search uses that same information imbalance to help determine what gets recommended, the established giants gain another enormous advantage. And if a smaller product finally does make it into the comparison only to have important information omitted or presented incorrectly, what can its creator possibly do about it?
I could challenge the comparison because I happened to be the person sitting in front of it. But what about the thousands of comparisons I will never see, happening privately between AI and potential customers I will never meet? I can't correct those. I can't point out what's missing. I can't ask the follow-up question that finally reveals the right information.
The potential customer has to recognize the problem themselves. And they may be the one person least equipped to recognize it—because they're the one searching for an answer.
Nobody owes the little guy a win. But the little guy should at least have a fair chance to be evaluated on what they've actually built.
Maybe that's the real AI search trap. The next great product has always had to overcome obscurity to become the next big thing. But if the tools we increasingly rely on to discover what's best naturally favor what is already big, overcoming obscurity may become harder than ever.
So use AI search. Use the comparisons. They're remarkable tools, and they can be incredibly useful starting points. Just don't mistake the first answer for the whole truth.
Because when we ask AI what's best, there's one thing we usually have no way of knowing: What did it leave out?
And unfortunately, that's why we asked in the first place.
← Back to Blog