Tracing the ghost in the whitepaper’s code.
Last week, a quiet signal pulsed through the crypto and AI crosshairs: Vals AI, a startup promising to be the third-party referee for large language models, closed a $40 million Series A led by a16z at a $400 million valuation. The headline reads like a standard infrastructure bet—money flows to the pick-and-shovel sellers. But the narrative buried beneath the press release is far more disruptive. Vals claims that OpenAI, Anthropic, Google, Meta, and xAI are now citing its evaluation results in their model cards. If true, this is not just a funding round; it is the moment AI evaluation ceased to be an afterthought and became a pillar of the industry's trust architecture.

Yet, as someone who has spent two decades watching the crypto industry cycle through narratives—from ICO whitepapers to DeFi yield alchemy—I’ve learned that the sale of a story often precedes the proof of its substance. The ghost in the code is not the technology; it is the belief that the technology can be trusted.
Context: The Broken Benchmark.
AI evaluation has long been a game of mirrors. Benchmarks like GSM8K, HumanEval, and MMLU have become so widely used that model trainers inadvertently—or deliberately—optimize for them. The result is a landscape where a model can score 90% on a static test yet fail miserably on a real-world task that requires nuanced reasoning across a codebase. The industry has known this for years. The solution, however, has remained elusive because the entities with the most to gain from objective evaluation—the model builders themselves—are also the ones controlling the tests.
Enter Vals AI. Its founder, Vals Smith, built a system that pulls real-world development tasks from historical GitHub pull requests. The idea is elegant: instead of a canned dataset, the model is evaluated on its ability to solve problems that actual developers faced. The test is hidden from the model, and the pass/fail decision is automated. The company then extends this methodology across domains—finance, legal, medical—to assess a model's production readiness. It is a classic engineering innovation: not a new algorithm, but a new way of using existing pieces to solve a systemic problem.
Weaving trust into the immutable ledger.
From a narrative perspective, Vals is not selling a tool; it is selling the scarcity of trust. In a market where every AI lab claims superiority, a third-party seal of approval becomes the differentiator. The company’s pitch recalls the early days of smart contract auditing. Before Trail of Bits and OpenZeppelin, projects could self-certify their security. The market demanded independence, and a cottage industry of auditors was born. Today, no serious DeFi protocol launches without an audit. Vals aspires to be that for AI.
However, the analogy is imperfect. Auditors are paid by the projects they audit, creating an inherent conflict of interest. Vals, too, faces this tension. The company’s revenue model is B2B: it sells evaluation services to enterprises and, presumably, to the very model providers it evaluates. The article from the monitoring channel, which reached me through a Web3 signal, notes that a16z’s portfolio companies are likely to become Vals’ first clients. This does not invalidate the product, but it does raise the question: can a subsidiary of a venture capital firm remain independent?
Core: The Narrative Mechanism and Sentiment Analysis.
Let me dissect the numbers. The $400 million valuation is not supported by disclosed revenue. The company claims its “2025 revenue has already reached 8 times the full-year 2024 revenue”—a statement that is either a typo, a misquote, or a deliberate obfuscation. The original Chinese article, parsed through a monitoring channel, admits the ambiguity. Even if we interpret it as “8x growth,” the base is unknown. It could be $100,000 to $800,000, or $1 million to $8 million. Neither justifies a $400 million price tag unless a16z is buying a thesis: that AI evaluation will become a mandatory infrastructure layer, akin to DNS or CDN, with network effects and high switching costs.
That thesis is plausible but fragile. The market for AI evaluation is currently crowded with open-source alternatives (Aider, SWE-bench, HumanEval), academic labs, and internal tools from the hyperscalers. Vals’ competitive advantage is its network of integrations with GitHub and the claim that its private test sets are not contaminated. But as I know from auditing ICOs in 2017, the first-mover narrative is often a mirage. The real value lies in the data moat—the accumulation of evaluation results across many models and tasks. If Vals can aggregate a corpus of verified performance data that no single lab can replicate, it becomes a reference standard. That is the dream.
The pixel that holds a soul.
Yet, the soul of evaluation is not just data; it is the human judgment that defines what “correct” means. The article reveals a hidden risk: Vals likely relies on manual annotation for financial, legal, and medical tasks. The cost structure of that human labor is undisclosed. Scaling human review while maintaining speed and consistency is a challenge that has broken many “AI for AI” startups. My own experience with the Plain English DeFi series taught me that human interpretation is not a bug—it is the feature. But it is also a bottleneck.
Contrarian: The Blind Spots.
The most counter-intuitive angle is that Vals’ success might actually undermine the case for third-party evaluation. If every major lab begins citing Vals, the evaluation becomes a de facto standard. But standards can be gamed. The article warns that “hidden tests could be reverse-engineered by model providers.” That is not paranoia; it is the history of adversarial AI. Once a dataset becomes high-stakes, the incentive to overfit or leak it becomes enormous. Vals would need to constantly generate new, private tests, which is an arms race that may not be economically sustainable.
Moreover, the blockchain community that is now buzzing about this news should be skeptical. The monitoring channel that published the original analysis is a Web3 insider source. The article itself is a rigorous seven-dimensional analysis—a format I recognize from both crypto due diligence and AI research—but it is based on a single company press release. The confidence rating for every dimension is C: “insufficient independent verification.” That is a red flag for anyone who has seen a whitepaper promise met with a rug pull.
Chasing the myth through the ledger’s fog.
Let me anchor this with a personal story. In 2021, I launched a small NFT collection called “Melbourne Memories,” embedding essays about gentrification into the metadata. It sold out, but the cultural impact was limited. The real lesson was that metadata—the layer of description and verification—is where meaning is made. Vals is trying to build the metadata layer for AI. But metadata is only as trustworthy as the authority that maintains it. And authority can be a single point of failure.
Takeaway: The Next Narrative.
The future of AI evaluation may not be a centralized service like Vals, but a decentralized protocol that uses cryptographic commitments to ensure test integrity. Imagine a system where test sets are hashed and published on-chain, and model providers submit their outputs with zero-knowledge proofs. Such a system would be tamper-proof and verifiable by anyone. It would be the “audit” equivalent of an immutable ledger. Vals could be the catalyst that pushes the industry toward that vision, or it could become the victim of its own success, too embedded in the venture capital ecosystem to remain neutral.
The echo of a promise unkept.
For now, the $40 million is a bet on a narrative. As a crypto editor who has seen the myth of “peer-to-peer electronic cash” become a Wall Street toy, I know that narratives can be powerful but ephemeral. The ghost in the whitepaper’s code is still there. We are still chasing a myth through the fog of a ledger that has not yet been built. The difference is that this time, the stakes are not just financial—they are existential. If we cannot trust AI evaluation, we cannot trust AI. And if we cannot trust AI, we cannot trust the future we are building.