This is speculation about where the AI industry is heading. Given how fast things change, there’s a high probability that this prediction will age like milk, but here we go.
In the last few weeks, I’ve started to get the impression that maintaining a moat in the model-building business is very difficult, if not impossible. Open-weight models lag 6 to 12 months behind closed-source models, but they eventually catch up.
For example, Llama 3 reached GPT-4-level performance on MMLU roughly one year later. This chart is a bit outdated and doesn’t include the most recent models, like DeepSeek V3, which is not only open-weight but also significantly cheaper (although it’s fair to assume that the real cost is understated).
Le Chat, based on Mistral flagship LLM (Mistral Large), is the latest to drop in the App Store, surpassing ChatGPT and DeepSeek, which also briefly overtook ChatGPT a few weeks ago.
One reason why it’s difficult to maintain a moat is that everyone has access to the same public data to train models. Access to multimillion-dollar data centers and the most advanced GPUs seemed to be a decent moat, but DeepSeek has shown that these limitations can be overcome with a mix of human ingenuity and free-riding on OpenAI’s achievements through distillation of their LLMs.
LLMs so far seem different from search. Since the early 2000s, Google has built an unshakable moat in search, and no competitor has been able to displace it. We haven’t seen a proliferation of search challengers at Google's scale. Only Bing has seriously attempted to compete, but it has never surpassed a 5% market share.
It’s not entirely clear why this is the case. The transformer architecture behind all modern AI models is public knowledge, just as PageRank, the indexing algorithm invented by Google, was published in early academic papers by Larry Page and Sergey Brin. Search is a more mature industry, and it was built around advertising from the start. It’s possible that Google’s early lead allowed it to accumulate user behavior data, leading to better algorithms and results, which attracted more users and advertisers, reinforcing a positive feedback loop that created network effects, user lock-in, and made Google a verb.
It’s true that OpenAI seems to always be at the forefront of innovation, with Operator and, more recently, Deep Research:
Deep research is OpenAI's next agent that can do work for you independently—you give it a prompt, and ChatGPT will find, analyze, and synthesize hundreds of online sources to create a comprehensive report at the level of a research analyst.
Those with ChatGPT Pro ($200/month) have tried it, and the results are mixed. In general, it works well when there is enough public information that isn’t contaminated by nonsense slop. On niche topics, Deep Research seems to perform better because it primarily functions as a retrieval tool. Historically, LLMs have struggled with generative tasks on niche topics because those topics make up a small portion of training data, increasing the likelihood of hallucinatory nonsense.
That said, I don’t see why other players wouldn’t be able to replicate Operator or Deep Research. DeepSeek R1, a reasoning model, made waves by achieving performance similar to OpenAI’s o1 at a much lower cost, but models from Google and Anthropic are also very close in performance.
So the question remains: where is the moat? How can companies maintain a sustainable competitive advantage without the constant, breakneck race we’re seeing now? The answer could lie in scarcity and secrecy. As Ben Thompson from Stratechery notes:
Infinite content generation increases the value of digital scarcity and verification, just as infinite transparency increases the value of secrecy. […] Secrecy is its own form of friction, the purposeful imposition of scarcity on valuable knowledge.
If all models are trained on the same data and have access to the same public sources at inference time, they will inevitably converge in performance. At that point, integrating well with proprietary data will be the only place to look for alpha: secrecy creates scarcity, and scarcity is where value lies.
It seems the VC industry is already thinking along these lines. David Jayatillake shared on The Joe Reis Show:
VCs have been saying “you need to be building your software in Rust. Your company will be worth more money.”
VCs are worried that if companies are building software that's easily reproducible, then those companies actually aren't very valuable. And so one of the ways to make the software harder to reproduce […] is to maybe not use like some of the more well used languages like Python or JavaScript, and use something that’s a bit more difficult to develop in, like Rust.
This is ironic, to say the least. The AI boom has been fueled by LLMs that scraped vast amounts of public data, including source code. Many companies have been created in the AI space in the last few years. Now, those very same companies are trying to defend themselves from AI scraping their codebase.
I still don’t have a clear idea of what form this will take or who the winners and losers will be. Enterprises have a lot of proprietary information, so companies that perform well in the enterprise space, have already an established relationship, and benefit from high trust, like Microsoft, Amazon, or Oracle, could be among the winners. Maybe we will see more and more in house development of AI capabilities.
Individual users also have proprietary information. Not a large amount on an individual basis, but highly relevant to each of them. Here, it’s harder to predict who will achieve the best product-market fit (PMF), as users can switch between service providers more freely.
Do you like this post? Of course you do. Share it on Twitter/X, LinkedIn and HackerNews






