Build1 distinct publisher3 min readUpdated
The $2,000 pot is noise. The scoring null, 20,000 random traders per asset re-drawn daily on the path that actually happened, is the part other leaderboards should copy.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Run FINCHAL's scoring formula against its own headline test and the two numbers do not sit where you would expect. Beating the ceiling puts an entrant at p = 0.95, and the score at that point is -log10(0.05), or about 1.30 [18]. The published landmarks on that scale are 2.0 for a one-in-100 outcome and 3.0 for one-in-1,000 [11]. Clearing luck therefore lands near the bottom of the range VIDRAFT chose to display, and by construction 1,000 of the 20,000 random players clear it in every market [19]. Reaching 3.0 means finishing ahead of 19,980 of them [20].
The asset-specific null is what makes those thresholds mean anything. The pre-season ceilings run from 86.6% on Bitcoin to 9.2% on gold [8], a factor of 9.4 between the two [21]. An 80% year in Bitcoin sits inside the range random position-flipping could have produced; 12% in gold sits outside it [9]. A single cross-asset winner was tried and dropped after six normalisation methods left the tails still favouring particular markets, so the money splits into four separate $500 awards [3].
Nor is the distribution fixed. FINAL-Bench says it rebuilds the random-player returns daily on the market path that has actually occurred, specifically to stop a broad rally lifting every long-biased entrant [10]. Beta gets re-priced as the season runs: an agent that was long through a rally is measured against 20,000 random traders who lived through the same rally.
The 1x exposure cap does the equivalent job on the other axis. VIDRAFT said an uncapped simulation returned 48,763% cumulative, which is what a leaderboard looks like when it has quietly become a contest over position size [6].
The eight self-tests on the scorer are where the lineage shows [13]. The interface withholds future prices from agents [12], but the harder leak is on the scoring side, and the test VIDRAFT calls critical checks that an entrant who opens a position on the same bar as a price jump earns nothing from that jump [13]. One shifted index would let the scorer read the future while still printing plausible numbers [14]. That is the failure FINAL-Bench documented on August 22, when the same drug-prediction task came out 0.211 AUROC apart depending on whether the split was time-based or random [15]. Leakage does not announce itself. It just improves the results.
Which is why the pot is the least interesting line in the announcement. Split four ways across a 122-day season, $500 per asset works out at roughly $4.10 per asset-day [22]; nobody staffs an agent for that, and runtimewire's own read is that the design work is the draw [17]. What is missing is ownership. The announcement runs under a SeaWolf-AI account displayed as VIDRAFT_LAB alongside the FINAL-Bench organization, whose page links to VIDRAFT, and the operational relationship among those four names is not defined in the public materials [16]. For a contest whose entire value rests on the credibility of one scorer, the party answerable for that scorer should be nameable.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Position exposure is capped at 1x and values outside the permitted range are clipped.
VIDRAFT said an early simulation without the exposure cap produced a cumulative return of 48,763%, showing how quickly a leaderboard can become a contest over position size.
Min-sik Kim, CEO of VIDRAFT, opened a financial forecasting contest on Monday that lets AI agents take positions directly and sets a deliberately high bar for calling their returns skill.
FINCHAL will award $2,000 across four markets after a 122-day season ending December 24.
VIDRAFT abandoned an overall cross-asset winner after testing six normalization methods and finding that differences in the tails continued to favour particular markets; FINCHAL instead awards $500 separately for each asset.
FINCHAL entrants submit a position between -1.0 and +1.0 for Nvidia, Bitcoin, gold through the GLD exchange-traded fund, or crude oil through USO; -1.0 is fully short, zero is flat, +1.0 is fully long, fractional positions are allowed, and each position remains active until it is replaced.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Mechanically detailed but single-sourced
The cluster contains one publisher working from the project's own Hugging Face announcement, so mechanism detail is unusually specific (position encoding, 1x clipping, 20,000-player null model, -log10(1 - p) scoring, eight self-tests, four MCP tools) while every quantitative figure is self-reported and unreplicated. The derived arithmetic on ceilings, resolution and prize-per-asset-day is checkable from published numbers, which raises internal consistency without adding independent confirmation.
Launch-stage, no participation data
Observed adoption is limited to the launch itself and to benchmark figures the project published about its own simulations. There are no entrant counts, agent submissions, third-party integrations, or partner deployments in the supplied material, and the season had only just opened at publication.
Design promise ahead of results
The framing is comparatively disciplined — the source itself calls the prize modest, flags that founder credentials are company-supplied, and notes the undefined organizational structure. Mild overstatement remains because the null-model design is presented as a template other leaderboards should adopt while no live season results, entrant data or independent audit exist yet, and because the scale's resolution at 20,000 draws is tighter than the one-in-1,000 language suggests.
Vendor-published launch with self-supplied credentials
Every substantive figure originates with the party promoting the contest: VIDRAFT announced FINCHAL on its own Hugging Face accounts, supplied the 48,763% uncapped-simulation figure, published its own luck ceilings, and provided unverified founder credentials while pursuing an institutional research-funding ambition. The publishing accounts and their relationships are themselves not fully defined, and the coverage derives from that primary announcement.
Clear mechanics, unverified performance
Confidence is moderate: the design facts are specific, internally consistent and arithmetically checkable, so what FINCHAL intends to measure is well established. It is capped by the single-publisher cluster, wholly self-reported figures, absent participation evidence, and an undefined organizational structure behind the benchmark.
build
The split, not the model: 0.606 to 0.818 AUROC without touching a hyperparameter1 distinct publisher
build
The judge is the product: building the scorer for an AI malaria drug leaderboard1 distinct publisher
build
Unsloth's 10% quant claim is really about which machines can run a 27B model1 distinct publisher
build
Hugging Face's $13B process puts most teams' model pipeline under a single owner2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 24, 2026