Build1 publisher2 min readPublished
CLQT measures a 0.30 gap between what LLM trading agents say and what they allocate
Bo Qu and Mingguang Chen seal every gather-analyze-decide-execute-reflect cycle into a recompute-verifiable audit chain and score five capability axes from it. In four weeks of live paper trading the gap came to 0.23.
The Engineer · Build desk

What happened
- CLQT enforces point-in-time data access through a hard TimeGate and models multi-component institutional transaction and financing costs; both are conditions of the evaluation.
- It computes a five-axis capability scorecard from the decision trail, covering Coherence, Acuity, Composure, Discipline and Reliability, with coherence scored partly by a held-out LLM judge.
- Validation ran a contamination-controlled, year-long multi-model backtest with a 13-configuration ablation grid, alongside a four-week live broker paper-trading track on post-cutoff data.
- Across both tracks the agents' allocations failed to follow their own stated analysis, a gap the authors put at +0.30 in the backtest and +0.23 in live paper trading.
- Once the modelled costs were charged, the agents beat defensive baselines but did not clearly outperform the index.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint A recompute-verifiable chain changes what counts as a submission: a claimed score can be rejected by rerunning the trail, so a published equity curve no longer discharges the burden.
- decision With the same model run under a committee scaffold and a single orchestrator, any team comparing two models has to hold the scaffold fixed or publish both numbers.
- exposure Oversight that reads an agent's written rationale is inspecting an artifact the paper says does not track the allocations, which puts review processes built on explanation quality at risk.
- precedent The six-condition coverage claim sets the bar the next trading benchmark will be held to, and gives reviewers a checklist to reject one that meets four of them.
Look-ahead bias is the default condition of a backtest. Pull a price or fundamentals series today and you get the revised version, restatements included, published after the date the agent is pretending to trade on. An agent that beats the index by reading a later revision has proved something about the harness. The abstract puts the consequence plainly: apparent alpha can dissolve once look-ahead bias and realistic trading costs are controlled [4].
Sealing each cycle into a recompute-verifiable audit chain is the part of this design worth copying [8]. A cumulative return curve has to be taken on trust. A score built from the chain can be rejected by a third party who reruns the trail.
In the campaign the capability leader was not the Sharpe leader [14]. The value of individual modules showed up on the capability and behavioral axes in cases where returns alone could not separate the configurations [16]. An ablation grid scored only on return will report that such a component does nothing.
The stating-versus-doing gap is the figure I would take into a design review. It ran +0.30 in the year-long backtest and +0.23 in the live track [15], which is 0.07 lower, about 23 percent smaller [22], over a window roughly a thirteenth as long [23]. The abstract does not say what scale the gap is measured on.
Qu and Chen are explicit about what one study cannot settle. They write that a credible model ranking would require many models, regimes and repeated runs, beyond the scope of this work [20]. Their standard for calling a model superior is dominance on the relevant capability axes, consistently across sub-periods. Until that holds, they write, the benchmark's product is "a reusable map of limitations" [21].
The coverage claim is a claim about everyone else's benchmarks. To the authors' knowledge, no existing benchmark satisfies all six conditions simultaneously: enforcing point-in-time data semantics, modelling realistic transaction and financing costs across asset classes, evaluating decisions at the portfolio level, measuring strategy consistency across rounds, accumulating structured multi-tier memory, and treating tool orchestration as a benchmarkable capability [18]. Six conditions at once is the bar CLQT will be measured against as well. For its cost result to transfer to your own agent, the modelled transaction and financing components [6] would have to match your asset mix, your holding periods and your broker's terms. Your data layer would also have to refuse post-decision revisions the way the TimeGate does [5].
What to watch
- Whether the full paper publishes the cost-model parameters and the model list, so a third party can rerun the audit chain and reproduce the scorecard.