Science1 distinct publisher2 min readUpdated
Artificial Analysis' AA-Briefcase grades deliverables built from four multi-week projects, but each task starts with no memory of the model's own earlier submissions. Continuity stays unmeasured.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
Independence is a grading convenience with a cost that lands on whoever reads the results. If a task starts cold, the only material a later week's deliverable can rest on is the shared source pool, never the spreadsheet the model itself produced two weeks earlier [4]. What survives is a measure of two things: whether an agent can find requirements nobody flagged inside thousands of files, and whether it hands back the right artifact in the right format [1][5][8]. Both are worth a number. Neither is a week of work.
The listing states plainly that each scenario is a multi-week workflow the agent works through in sequence, with several tasks per week [3]. It does not say how a task in week four receives what week two produced, and describes no mechanism for passing a prior deliverable forward [17]. Either the brief restates the needed inputs, or a reference version is supplied, or the tasks were written not to need one. Those are three different benchmarks wearing the same score.
The arithmetic is worth stating. Four scenarios and 91 tasks average about 23 tasks per scenario [10], and each is a fresh start, so a model is scored 91 times on beginning and zero times on continuing [11].
Pairwise judging is house style here. GDPval-AA v2 derives Elo ratings from blind pairwise comparisons of output across 44 occupations [15], and two of AA-Briefcase's three check types use the same construction, one for rigour and one for presentation [6]. Ranking is what that yields. An absolute standard, the kind a client applies when accepting or rejecting a memo, comes only from the binary rubric [12].
Artificial Analysis already grades state elsewhere: its EnterpriseOps-Gym implementation scores agents on the final state of the underlying databases after multi-step workflows across eight business domains [14]. A database is checkable. A forecast built on last week's flawed model is not, at least not without deciding in advance whether the agent should inherit its own mistake or notice and repair it. That is the harder evaluation to build, and it is the one an operator running an agent unattended from Monday to Friday needs.
What is public so far is the shape rather than the substance. A fifth scenario sits on Hugging Face to demonstrate structure, submission and grading, explicitly excluded from official results [7], while the same catalogue lists a private long-horizon evaluation built on business workflows requiring spreadsheets, presentations and memos [13]. Carry-over, whenever someone builds it, is the line between a model that writes deliverables and an agent that holds a project.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
AA-Briefcase evaluates models across four multi-week knowledge work projects comprising thousands of input files and 91 tasks in total.
Across the scenarios, models must complete realistic professional workflows in fields such as data science, product management, and corporate strategy.
Each scenario is a multi-week workflow that the agent works through in sequence, each week holding several tasks, and every task is a deliverable graded against a rubric of checks.
Although tasks within a scenario share files and context across weeks, models currently complete each task in an independent run, without carrying over their own prior submissions.
The binary pass-or-fail check asks whether the model followed the task instructions, identified requirements hidden across source files, used the correct evidence, and reached the right conclusions.
Two of the three check types are pairwise comparisons against another model's submission: one asking which deliverable is more thorough, analytically rigorous and well-supported, the other which is more professionally presented.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed primary methodology, no external corroboration or results
Every factual claim traces to the benchmark owner's own documentation, which is the authoritative source for its design and is unusually specific (four scenarios, 91 tasks, three named check types, exact rubric question, file-type and tool-call reporting axes). Evidence strength is capped because there is one publisher, no model scores in the supplied text, no grader or comparison-pool disclosure, and no third party has verified the implementation.
No uptake evidence in supplied sources
The supplied material shows only the publisher's own actions — the methodology page and a demonstrative Hugging Face scenario. There are no model results, download or usage figures, third-party citations, or vendor references to AA-Briefcase, so nothing in the sources supports an adoption level and none is inferred.
Mildly overstated framing, limitation self-disclosed
Positioning around 'multi-week' projects and, in the adjacent catalogue, 'frontier agentic capability in long-horizon knowledge work' implies continuity that the scoring does not test, since each of the 91 tasks is an independent run. The gap is small and positive rather than large because the publisher states the constraint plainly in the same paragraph as the scope claim, and because two of three grading dimensions are openly described as relative rather than absolute.
Benchmark owner documenting and promoting its own evaluation suite
Artificial Analysis authors, operates and publishes the evaluation it describes, and the same page markets a broad catalogue of its other agentic benchmarks, including a private frontier evaluation that cannot be independently inspected. That gives it a commercial interest in the perceived rigour and reach of its suite. Countervailing signals are the candid disclosure of the independent-run limitation and the explicit labelling of the public scenario as non-scoring.
Design facts firm, capability conclusions unsupported
Confidence is high that the described design and its stated limitation are accurate, because they come from the benchmark's operator in specific terms and the derived findings are simple consequences of stated counts. It is low for anything beyond design: no scores, grader details, comparison-pool composition, or independent replication are present, and adoption is unmeasurable from the supplied material.
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
science
Frontier scores at $2/$6: Grok 4.6 ties GPT-5.6 on one evaluator's index for a fifth the output price1 distinct publisher
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
build
Artificial Analysis moves eval onto your data, and turns model choice into procurement1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 22, 2026