Skip to content

Science1 publisher3 min readPublished

When the method is a purchase: BioScience team says tool opacity needs an audit trail

An international group writing in BioScience says AI models, satellite feeds and proprietary sensors have put unverifiable steps inside the method itself. Every fix they propose is lab-side.

The Scientist · Science desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A new study published in BioScience by an international team of scientists finds that scientists increasingly rely on powerful data sources and tools that they often cannot fully understand, inspect or verify, addressing reproducibility, trust and the future of scientific research when critical technologies shape science without being fully open to scrutiny.
  • Ivan Jaric, a researcher at the University of Paris-Saclay and lead author of the study, said many of these tools are true black boxes that keep the processes behind their results largely hidden, and are often owned by private companies that intentionally limit access to information about how their systems operate or process data, guided by proprietary constraints and commercial aims.
  • The paper identifies several types of black boxes becoming widely used in ecology and conservation, with large language models and other AI technologies among the most prominent, increasingly used to analyse massive data sets, interpret satellite imagery and model ecosystems.
  • Researchers often have little or no access to the data used to train these AI systems, the underlying algorithms, direct system testing, or an understanding of how and why they generate particular outputs.
  • Many remote sensing products rely on proprietary processing that researchers cannot fully access and verify, while some wildlife tracking devices provide only processed animal locations while withholding the underlying raw data.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

An international team writing in BioScience argues that ecology and conservation research now routinely depends on tools whose users cannot fully understand, inspect or verify them [1]. That relocates the reproducibility problem: it is no longer mainly a question of how a group analysed its data, but of what the group bought and what the vendor declined to disclose.

The paper's list of black boxes maps onto a procurement catalogue rather than a statistics syllabus. Large language models and other AI systems are used to analyse large data sets, interpret satellite imagery and model ecosystems, while researchers often have little or no access to the training data, the algorithms, direct system testing, or any account of why a given output was produced [3][4]. Many remote sensing products depend on proprietary processing that cannot be fully accessed or verified, and some wildlife tracking devices return processed animal locations while withholding the raw data [5]. Search engines and social media platforms, now used as biodiversity and human-behaviour data sources, run on hidden algorithms and shifting policies that can introduce unknown biases [6]. Social surveys have the same exposure: private contractors recruit participants and run the instrument, often without telling the researcher how respondents were selected, how data quality was maintained, or whether answers were affected by AI agent interference [7]. That is five distinct supplier categories sitting inside the method [14].

Lead author Ivan Jarić of the University of Paris-Saclay said the systems are often owned by private companies that intentionally limit access to information about how they operate, guided by proprietary constraints and commercial aims [2]. Co-author Karen Anderson of the University of Exeter added that the cause is not only commercial: modern tools are becoming so technically complex that users, and in some cases their developers, may struggle to scrutinise how they work [8]. The authors say publish-or-perish incentives, growing data volumes and urgent environmental crises all deepen the dependency [9]. Their stated worry is cumulative rather than dramatic. Alongside monopoly risk, obstacles to open science and susceptibility to manipulation, they warn that if key analytical steps cannot be inspected or repeated, confidence in findings erodes gradually [10]. As AI systems become more capable and autonomous, they expect results to become harder to interpret and verify [13].

The remedies are worth reading as an operational checklist, because none of them requires a vendor to open anything. The authors recommend prioritising open-source software and hardware where possible, benchmarking proprietary tools against transparent data sets, comparing results across multiple methods, documenting training data, pipelines, versions, settings and especially tool limitations, and broadening awareness of the problem [11]. All five are actions the purchasing lab performs on its own [15]. Human oversight, they say, should remain central throughout the research process [12]. In practice that means the limitations section stops being a courtesy and becomes a register of which steps in the pipeline the group cannot reproduce without a supplier.

What to watch: whether journals and funders start requiring tool version, setting and limitation disclosure as a condition of submission, and whether vendors respond with auditable benchmarks. Also worth noting what this paper does not supply. The account of the study reports no estimate of how many studies now rely on black-box components, or how much measurable damage that has already done to replication [16], which leaves the argument resting on the categories it identifies rather than on a count.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories