When Tencent quietly released the WorkBuddy Bench last week, most of the crypto world was busy chasing the next AI agent token pumping on a KOL narrative. But the numbers buried in that benchmark tell a story that cuts straight to the core of the AI-agent stack—and it's a story that will reshape how we value crypto AI projects in this bull market.
Claude Code won 17 out of 28 head-to-head comparisons against CodeBuddy. In coding tasks, it swept 7-0 across all seven underlying models. The same models, the same tasks, but a different agent execution harness—and the results diverged by over 10 points in some cases. Code talks, but stories sell. The story here is that the agent execution layer, not the model, is the new bottleneck.
Context: The Pivot from Model Hype to Execution Reality
The crypto AI narrative has been dominated by model performance—ETH stakers comparing GPT-4o to Claude 3.5, token prices rising on fine-tuning announcements. But the real value creation in AI agents, especially in crypto, happens at the orchestration layer: how the agent manages context, calls tools, handles multi-step workflows, and integrates with on-chain protocols. Tencent's WorkBuddy Bench, despite being a self-published, small-sample benchmark (260 tasks, 4 categories), provides the first systematic evidence that the harness matters more than the underlying model in specific scenarios.

Based on my experience auditing DeFi protocols for oracle latency issues, I've seen how a seemingly minor change in execution order can cascade into a liquidation event. The same principle applies here: an agent's harness is its execution environment, and a poorly designed harness will fail even with the best model. The crypto community, drunk on bull market euphoria, tends to overlook these technical details when evaluating AI agent projects. The token price is the narrative; the actual agent performance is the code.
Core: The Harness Effect and Its Crypto Implications
Let's dive into the data. Tencent's benchmark compared seven unnamed models across four task categories—coding, web, office, and security—using two harnesses: their own CodeBuddy and Claude Code. The key finding: in coding, Claude Code won all 7 comparisons (7-0). In web and office, CodeBuddy won 4-3 each. In security, Claude Code won 4-3. The aggregate: Claude Code 17, CodeBuddy 11.
The math is self-consistent: 7 models × 4 categories = 28 comparisons. The fact that the same model, when swapped between harnesses, showed such a stark difference—especially the 7-0 sweep in coding—strongly suggests that the harness is a separate performance variable. This is not trivial. In crypto, AI agents are being deployed for automated trading, DeFi yield optimization, and governance participation. If the harness can swing performance by 10+ points, then the choice of execution framework is as critical as the choice of model.
Narrative is the new liquidity. In the bull market, liquidity flows to the strongest narrative. The narrative of "model-first" is deeply embedded in crypto AI. Projects like Bittensor, Fetch.ai, and others tout their model capabilities. But this benchmark suggests that the narrative is shifting: the execution layer is becoming the differentiator. The crypto AI projects that will survive the next cycle are those that optimize for agent orchestration, not just model fine-tuning.
Consider the analogy to DeFi. In 2020, everyone was obsessed with TVL (total value locked) as the metric of success. But the real alpha was in the execution layer—the smart contract architecture, the oracle design, the liquidation mechanism. The projects that survived the 2022 crash were those with robust execution layers, not just high TVL. Similarly, AI agent projects that focus on harness quality—tool integration, context management, error recovery—will outperform those that simply wrap a powerful model.
Contrarian: The Benchmark's Hidden Biases
Before you FOMO into the next AI agent token based on this benchmark, consider the contrarian angle. The WorkBuddy Bench is a self-published benchmark by Tencent. The tasks are not open-sourced, the models are not disclosed, and the scoring methodology is opaque. The 4-3 wins in web and office tasks could be a result of task selection that favors Tencent's ecosystem—enterprise WeChat, Tencent Docs, Tencent Meeting. In other words, the benchmark may have an inherent bias toward CodeBuddy's home turf, while the coding tasks (which are more standardized) show the true gap.
Furthermore, the sample size is small. 260 tasks across 4 categories means roughly 65 tasks per category. A 4-3 win is statistically indistinguishable from noise. The 7-0 sweep in coding, however, is robust—it's a 1 in 128 probability if the true win rate were 50%. So the coding gap is real, but the other categories may be artifacts of task design.
Hype decays; utility endures. The utility of an AI agent in crypto is not measured by benchmark scores on arbitrary tasks, but by its ability to execute real-world workflows—like monitoring a Uniswap pool, executing a flash loan, or participating in a DAO vote. The benchmark tells us that Claude Code has a superior coding harness, but for crypto agents, the critical tasks are often more about web interactions (DEX interfaces, wallet connections) and security (private key management, transaction signing). The fact that CodeBuddy holds its own in web and security categories suggests it may be more suitable for crypto-specific agent use cases.
The Machine Economy Blueprint
This brings me to a broader thesis that I've been developing since 2025, when I first started researching autonomous agent economies. The next bull run will not be driven by human speculation alone, but by machine economies—agents that pay each other for services, execute microtransactions, and optimize their own resource allocation. In such a world, the execution layer becomes the economic substrate. The harness is not just a technical tool; it's the operating system for machine-to-machine commerce.
Tencent's benchmark, despite its flaws, provides a glimpse into this future. The fact that a single harness change can swing performance by 10+ points means that the agent's "digital reflexes" are programmable. In crypto, we already have programmable money (smart contracts). The next step is programmable agents that can interact with that money in a secure, efficient, and context-aware manner.
Based on my audit experience of DeFi protocols, I've seen how fragile the execution layer can be. A single missing check in a smart contract can lead to a multi-million dollar exploit. Similarly, an agent harness that fails to handle a timeout or a reorg can cause catastrophic losses. The crypto AI projects that will win are those that invest in harness engineering—not just model training.

Takeaway: The Next Narrative Shift
So where does this leave us? The benchmark is a signal, not a verdict. It tells us that the competition is shifting from the model layer to the execution layer. In crypto, this means the value of AI agent tokens will increasingly depend on the quality of their agent orchestration, not just the base model.
Watch for projects that are building agent-specific execution environments—like autonomous agent frameworks that integrate with on-chain data, or tool-calling systems that support multi-chain operations. The projects that have already optimized their harness will be the ones that capture the next wave of liquidity.

Narrative is the new liquidity. Right now, the narrative is still about model performance. But the data is clear: the agent execution layer is the hidden variable. The question is not whether the market will realize this, but when. And when it does, the tokens that reflect execution-layer value will be the ones that outperform.
As for CodeBuddy, it lost the coding battle, but it may win the crypto war—if Tencent can leverage its ecosystem for web and security tasks. But that's a story for another cycle. For now, the takeaway is simple: stop buying the model, start buying the harness.