
CodeBuddy vs Claude Code: The Benchmark That Exposes Crypto's Agent Blind Spot
CryptoFox
17-11. That's the scoreline from Tencent's own benchmark. CodeBuddy loses to Claude Code in a head-to-head comparison across 7 models and 28 tasks. The headline is easy: Tencent's agent is weaker. But the real story is what the data reveals about the fragility of agent benchmarks in crypto—and why the industry's growing reliance on AI agents is built on a foundation of sand.
For crypto developers, this isn't just a tech rivalry. It's a risk signal. The agents that audit your smart contracts, manage your yield strategies, and execute your trades are only as good as their execution layer. Tencent's WorkBuddy Bench test drives that point home. But the methodology is as shaky as a unaudited DeFi protocol.
Context: Why Crypto Should Care About Tencent's Agent Benchmark
Tencent's WorkBuddy Bench is a new agent benchmark that tests four task categories: coding, web, office, and security. The test uses the same underlying model but switches between two 'harnesses'—CodeBuddy (Tencent's own) and Claude Code (Anthropic's). The result: a 17-11 win for Claude Code across all comparisons. On coding tasks, Claude Code swept 7-0. On web and office, CodeBuddy won 4-3 each. On security, Claude Code won 4-3.
But here's the crypto connection. AI agents are flooding the blockchain space. From automated trading bots on Solana to smart contract auditing agents on Ethereum, the industry is racing to deploy LLM-powered tools. Yet there is no standardized benchmark for these agents. Most projects rely on self-reported metrics or generic benchmarks like SWE-bench. Tencent's test is one of the first to isolate the harness effect—the execution layer that orchestrates tool calls, context management, and task decomposition. For crypto, this is critical. If a harness can boost performance by 10 points on the same model, then the choice of agent framework could mean the difference between a secure audit and a drained treasury.
Core: The Data That Matters
Let's break down the numbers. 7 models, 4 categories, 28 comparisons. The math is self-consistent: 17+11=28. Claude Code dominates coding—7 out of 7 models performed better with Claude Code's harness. This is not a fluke. It's a systematic advantage in the execution layer for code tasks. For crypto, coding tasks include writing Solidity, debugging Vyper, or auditing cross-chain bridges. If Claude Code's harness is superior for these, then every crypto developer using a CodeBuddy-based agent is leaving performance on the table.
But the story gets more nuanced. In web and office tasks, CodeBuddy wins 4-3 each. This suggests a scenario-specific advantage—likely due to deeper integration with Tencent's ecosystem (WeChat, Tencent Docs, etc.). For crypto, this matters less. The highest-value tasks for crypto agents are coding (smart contracts, exploits) and security (vulnerability detection). In both, Claude Code wins. Security: 4-3 for Claude Code. That's a narrow margin, but still a loss.
What's the key insight? The harness effect is independent of the underlying model. The same model switches harness and scores shift by over 10 points—this is a massive effect size. In my experience auditing over 500 token contracts during the 2017 ICO boom, I saw a similar pattern: the execution layer (the smart contract logic, the calling sequence) often mattered more than the base asset's tokenomics. The same principle applies here. The agent's harness is the bottleneck.
s static.
But caution is required. The benchmark is self-reported by Tencent. The task set is small—260 tasks across 4 categories. The coding tasks may be biased toward Claude Code's environment (e.g., Unix shell, Git operations). The office tasks may favor CodeBuddy's integration with Tencent's suite. For crypto, the real-world tasks are different: on-chain data queries, MEV detection, gas optimization, cross-chain interaction. None of these are tested. The benchmark's external validity is low.
s static.
Contrarian Angle: The Blind Spot in the Blind Spot
Everyone will focus on the winner. But the real unreported angle is this: the benchmark itself is a symptom of the crypto industry's growing reliance on unverified AI agents. We are rushing to deploy agents for trading, auditing, and governance without rigorous, independent testing. Tencent's test, for all its flaws, is at least a step toward transparency. Most crypto projects do not even have that.
Consider the implications for DeFi. If a protocol uses an AI agent to monitor lending pools or adjust liquidation parameters, a flawed harness could lead to catastrophic losses. The 2022 Terra collapse showed what happens when systemic risk is ignored. Agent harnesses are the new systemic risk. They are the hidden layer that can amplify or mitigate errors. And the industry is not watching.
Second contrarian point: the benchmark's focus on harness may be a distraction. The real value for crypto is not in the agent's ability to execute code but in its ability to interact with decentralized infrastructure—reading on-chain state, signing transactions, managing private keys. None of these are tested. The benchmark is a sandbox. Crypto needs real-world stress tests.
Third: the market may be overvaluing coding agents. The most lucrative crypto applications for AI are not coding assistants but automated risk management and compliance. In 2025, with institutional adoption accelerating, the demand is for agents that can interpret regulatory frameworks (like MiCA) and generate compliance checklists. CodeBuddy's office and web strengths may be more relevant for that use case than Claude Code's coding prowess. Yet the benchmark's categories are generic.
s static.
Takeaway: What to Watch Next
The battle for agent supremacy is just beginning. But don't trust a single vendor's benchmark. The crypto industry needs independent, on-chain verified agent performance tests. I want to see third-party replication of WorkBuddy Bench results. I want to see task sets that include real-world DeFi scenarios: flash loan detection, arbitrage optimization, multi-sig management. Until then, the 17-11 scoreline is a data point, not a verdict.
Watch for these signals: (1) Third-party reproductions of the benchmark. (2) Release of the task set and model list. (3) Crypto-native agent benchmarks from projects like Gaianet or Autonolas. (4) Real-world failure rates of agents in production. The market will sort winners when real funds are at stake.
Data over destiny. The numbers don't lie, but they also don't tell the whole story. For crypto, the agent revolution is still in its infancy. Don't let a single benchmark shape your strategy. Build your own tests. Audit your own workflows. The only benchmark that matters is the one that protects your capital.