The most valuable infrastructure in the AI gold rush may not be compute or models—it's the ability to independently verify them. Vals AI's $40 million Series A from a16z at a $400 million valuation is a bet on that thesis. But as a crypto analyst who spent 2017 auditing ICO whitepapers and 2020 modeling DeFi solvency, I see parallels that the market is ignoring. Liquidity is the only truth in a volatile market, and the same applies to AI evaluation: without verifiable data, the hype is just noise.
Context: The Benchmarking Crisis
Large language models are evaluated on datasets like GSM8K and HumanEval, but these benchmarks are increasingly contaminated—training data leaks into the test sets, and model vendors optimize specifically for them. The result is a gap between leaderboard scores and real-world performance. Vals AI aims to solve this by using real-world tasks from GitHub pull requests as dynamic evaluation sets. Instead of static questions, they extract development tasks from arbitrary repositories, run hidden tests, and score models on actual code completion. They also extend to finance, law, and medical domains. The company claims that OpenAI, Anthropic, Google, Meta, and xAI reference their model cards, and that revenue has grown 8x this year.
Core: The Infrastructure Play
At first glance, Vals AI is a benchmarking tool. But the structural play is deeper. The evaluation becomes a gateway for enterprise AI adoption. Companies can connect their own codebase, run tests, and see how different models perform on their specific tasks. This transforms model selection from a leaderboard gamble to a metrics-driven procurement decision. The product sits between developers and model vendors, capturing a data moat: every evaluation reveals which models fail on which tasks, creating a proprietary dataset of failure modes. This is the same logic that drove AWS's early dominance—not just selling compute, but capturing the metadata of how it's used.
But here's the risk that my pre-mortem framework flags: verification centralization. If a single entity controls the evaluation layer, it becomes a bottleneck for model deployment. a16z's investment gives Vals access to its portfolio companies, but it also creates a conflict of interest. Vals is both the auditor and the auditor's customer. In crypto, we learned that trustless systems require multiple independent validators. The same applies here. The article from the blockchain monitoring channel that reported this news highlights the absence of third-party verification of Vals's own claims. The revenue multiple is ambiguous, the customer count is undisclosed, and the technical details of how they avoid test set contamination remain opaque. Risk is not avoided; it is priced and hedged. The market is pricing this as a $400 million infrastructure bet, but the hedge is missing.
Contrarian: The Decoupling Thesis
The contrarian angle is that AI evaluation will decentralize, not consolidate. The same forces that pushed crypto from centralized exchanges to self-custody will apply here. Enterprise clients, especially in regulated industries like finance and healthcare, will demand verified evaluation results that are immutable and auditable. This is where blockchain-based verification could enter. Imagine an on-chain registry of evaluation results, where each test is hashed, and the model's performance is timestamped. Vals AI's current model is a centralized oracle—it reports truth, but cannot prove it in a trustless way. The market will eventually demand a verifiable compute layer for AI evaluation, similar to how Proof of Compute protocols are emerging for AI training. This is a natural extension of my 2026 work on AI-crypto computational markets.
Moreover, the regulatory angle cannot be ignored. The Tornado Cash sanctions set a precedent that writing code can be a crime. If AI models are used for harmful outputs, the evaluation platform that certified them could face liability. Vals AI's position as a third-party certifier exposes it to legal risk that the market is not pricing. The model card references from OpenAI and Anthropic may be one-time experiments, not recurring revenue. The 8x revenue growth could be a single large contract from a VC portfolio company. These are the same signals I saw in 2017 ICOs with inflated revenue projections based on a single whale investor.
Takeaway: Positioning for the Cycle
Vals AI is a symptom of a maturing market that needs verification infrastructure. But the centralized approach is a temporary solution. The long-term value lies in decentralized, verifiable evaluation—a layer that any model can be tested against, and any user can audit. The current bull market euphoria around AI is masking the technical flaws in the evaluation stack. Smart contracts execute, they do not negotiate. The same rigor that crypto applied to financial verification must apply to AI verification. The question is not whether Vals AI will succeed, but whether the market will learn its lesson before the next crash.