Build1 distinct publisher3 min readPublished
The release ships the 15-year news crawl, 372,907 translated STEM problems, the filtering code and the training recipe, so an outsider can check the contamination scans instead of taking the benchmark table on trust.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The ArmSTEM verification loop is the interesting engineering here, worth reading before the score table. COPA machine-translated English math and science problems into Armenian, masking numbers and mathematical notation first so the translation could not corrupt them, then asked an independent model to solve the Armenian version and reproduce the original answer [13]. When the solver failed, it got the English original as a control, and items that defeated it in both languages were retained with a `solver_limited` label rather than counted as verified translations [14]. That label is the good part. It separates "the translation broke the problem" from "the problem is hard", a distinction most synthetic-data pipelines collapse into a single discard bucket.
The loop's blind spot is a translation that stays solvable, still lands the original answer, and is wrong in some way the answer never exposes. Human review covered 300 items: two native Armenian speakers assessed logical coherence and correctness separately, agreed on every verdict, accepted 299, and recorded a Cohen's kappa of 1.0 [15]. Against 372,907 problems, that sample is 0.08% [22]. 324,323 of the problems carry step-by-step solutions [12], leaving 48,584 without them [21]. COPA itself notes the process remains partly dependent on model-based judgment [19].
ArmWeb is 4.37 million documents and about 3.3 billion tokens measured under the Gemma-4 tokenizer [6], roughly 755 tokens per document [20], which is news-article length and what a 15-year crawl of Armenian news sites should produce [7]. Set against a 9-billion-parameter model, 3.3 billion tokens is 0.37 tokens per parameter [24]. This is an adaptation corpus, and anyone planning to train Armenian from scratch on it should price the shortfall before starting.
The contamination figures are what should change how you read the ranking. COPA reports benchmark matches in 7.9% of CulturaX-hy documents, 10.9% of HPLT-v2-hy and 17.4% of FineWeb-2-hy [9]; its own post-split pool sat at 3.3%, or 147,101 documents, which 13-gram decontamination then removed [10]. FineWeb-2-hy's rate is 5.3 times COPA's pre-cleaning rate [23]. COPA says arm-gemma-e4b recorded the highest mean score among the open Armenian models it evaluated [4]. For that to transfer, the evaluated set has to contain the model you would otherwise have picked, and you have to accept a decontaminated model being scored beside models whose Armenian slices overlapped the test items. Both conditions weigh against the ranking's transferability.
On provenance the release goes further than most: each released record carries source and URL metadata [8], and COPA describes this as the first open Armenian LLM published with its complete training data and recipe [5]. Provenance metadata establishes where a document came from; whether it can be redistributed is a separate question, and the material as published does not set out licensing terms for fifteen years of crawled news [25]. The validation approach has history behind it, too: Arakelyan's SynDARin paper proposed generating and validating question-answering datasets in low-resource languages and tested it on a 1,200-sample Armenian set [17]. COPA sells decision-intelligence products to institutional and commercial clients [18], so the release doubles as a credential. A team defending a model to a regulator will want provenance and licence both, and this release has shipped just the one.
Ranked by verification strength, evidence, and original report placement.
COPA says the resulting model, arm-gemma-e4b, recorded the highest mean score among the open Armenian models it evaluated.
ArmWeb contains 4.37 million documents and about 3.3 billion tokens under the Gemma-4 tokenizer.
The paper's contamination analysis reports benchmark matches in 7.9% of CulturaX-hy documents, 10.9% of HPLT-v2-hy and 17.4% of FineWeb-2-hy.
ArmWeb's post-split pool had a 3.3% contamination rate, representing 147,101 documents, before COPA removed the matches through 13-gram decontamination.
In the paper's human-evaluation results, two native Armenian speakers separately assessed a 300-item sample for logical coherence and correctness, agreed on every verdict, accepted 299 examples and recorded a Cohen's kappa of 1.0.
Erik Arakelyan and an eight-person COPA team released a 9-billion-parameter Armenian base model alongside the datasets, code and training recipe used to build it, according to the team's Hugging Face release article. The team includes Khatun Avetisyan, Meri Davtyan, Heghine Grigoryan, Nane Khachatryan, Hayk Shahsuvaryan, Henrik Sergoyan and Vahan Martirosyan.
Distinct publishers with included, body-backed reporting in this cluster.
runtimewire.com
1 article · September 4, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Unsloth's 10% quant claim is really about which machines can run a 27B model1 distinct publisher
build
Hugging Face's $13B process puts most teams' model pipeline under a single owner2 distinct publishers
build
Intel puts its Arc GPU operating knowledge inside the coding agent already installed1 distinct publisher
invest
Washington pitches Carolina Principles to G20, urging no new AI rules or bodies1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Precise, traceable, narrated once
The numbers are specific to the point of being falsifiable — 147,101 contaminated documents, 13-gram decontamination, kappa 1.0 on 300 items, a named tokenizer for the token count — and the artefacts they describe are public, which is unusual and counts for a lot. What holds the score down is that all of it arrives through a single outlet reading COPA's own paper, including the contamination rates attributed to CulturaX, HPLT and FineWeb. The release is verifiable; it has not yet been verified.
Shipped, not yet taken up
What exists is the release event and a self-run scoreboard, both dated to the same day. Nobody reports a download count, a fork, a downstream fine-tune or an Armenian product built on ArmWeb, and the two rival Armenian models appear only as benchmark rows rather than as users of anything here. That is the expected state for a day-old research drop, and it is also the whole of the adoption record.
A step ahead of its own proof
The framing is more careful than most, and the openness argument largely delivers. Two edges overreach. 'First open Armenian LLM with complete training data and recipe' is a priority claim with no survey behind it, and 'highest mean score among open Armenian models' means 0.500 against 0.477 for its own starting checkpoint, on a suite COPA assembled, while losing to that checkpoint on two of six tasks. Set against that, the corpus work is arguably undersold: a decontaminated 3.3-billion-token Armenian crawl with URL-level provenance outlives whichever model happens to sit on it.
Vendor holding its own scorecard
COPA sells decision-intelligence products, and Runtimewire says outright that the Armenian release gives that commercial work a public technical foundation. The team therefore chose the benchmark suite, ran the contamination scans on both its own corpus and its competitors', and supplied the first-ever framing. The countervailing incentive is real and unusual: publishing the crawl, the filters and the recipe hands critics the tools to embarrass you, which is not what a team optimising purely for headlines does.
One narrator, auditable subject
Confidence sits mid-range for an unusual reason. Normally a single-publisher story about a company's own paper would sit lower, but the substance is a data release whose claims can be tested by anyone who downloads it, and the reporting is granular rather than promotional. The pull downward: no second outlet, no outside replication, no usage signal, and a silence on licensing for fifteen years of scraped publisher text that nobody has yet asked about.