Skip to content

Science1 publisher3 min readPublished

Vision models fine-tuned on Cornell's labeled reef video still struggle to read wild fish behavior

Cornell's WildFin team found standard vision models still misread fish behavior after tuning on nine hours of reef video that took about 2,000 hours to make. The team blames scarce wildlife footage in model training and is asking ecologists to release their field video, a remedy this study has not yet tested.

The Scientist · Science desk

Illustration accompanying Vision models fine-tuned on Cornell's labeled reef video still struggle to read wild fish behavior

What happened

  • When Abigail Grassick, a Cornell doctoral candidate, first ran existing vision models on her reef footage, they could not keep track of an individual fish, let alone spot feeding or predator evasion.
  • Grassick and a team of undergraduates labeled each frame by picking from a list of behaviors such as foraging, feeding and fighting.
  • The team hopes computer scientists will use WildFin to build and test models that interpret fish behavior better.
  • Grassick presented the work at the European Conference on Computer Vision in Malmo, Sweden, on Sept. 8, and the paper is posted on arXiv.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • cost Until models can code behavior, ecologists pay for each new hour of usable reef video in human labor at roughly WildFin's rate, before any analysis begins.
  • contradiction Hein's data-scarcity explanation and the fine-tuning result pull apart: nine added labeled hours did not rescue performance, so the evidence does not yet show that more footage alone fixes it.
  • capability Competing models can now be scored on the same frame-labeled wild footage, so claims of progress on fish behavior become comparable.
  • decision Labs with archived field video now have a concrete use for it, and releasing it would let model builders test whether more wild footage helps, the question WildFin's own fine-tuning left open.

Producing the dataset took about 1,400 hours in the field and 600 hours of annotation, and Grassick said the work was incredibly time-consuming [4]. For each finished hour of footage, that comes to roughly 156 hours underwater [3] and 67 hours at a screen [4]. In all, it is about 222 hours of human work per hour of labeled video [2].

Andrew Hein, an associate professor of computational biology at Cornell and a co-author, says the models perform poorly because they are trained on very little wildlife footage [6]. "People are training these computer vision models on internet-scale data. Well, what fraction of internet-scale data are scientific imagery? The answer is, it's a tiny fraction," he said [7]. The explanation is plausible. This study has not tested it. When the team fine-tuned several standard models on WildFin to classify behavior, performance stayed poor [5]. Nine hours is a small addition to an internet-scale training diet [3], and one fine-tuning run cannot say whether ninety hours or nine hundred would be enough.

Jennifer Sun, an assistant professor of computer science at Cornell, points to a second problem. She had used computer vision to interpret animal behavior in the lab, but wild footage, with roving bands of fish, variable lighting and shifting camera angles, was harder [8]. "When I first talked to Andrew, I never anticipated how hard his videos were to analyze," Sun said. "He has hundreds of fish, and it's hard to know if it's even the same fish in subsequent frames." [9] That is a tracking failure, and it comes before any behavior label. A model cannot say a fish is feeding if it has lost track of which fish it was watching.

The dataset contains something close to a control for this. Besides the crowded bommie clips, it includes video of a diver following a single fish across the seafloor, much like footage divers shoot on vacation, collected and annotated by colleagues at the University of Colorado Boulder [11]. Comparing scores on the two kinds of footage would show how much of the failure comes from crowds and how much from the behaviors themselves. The phys.org account does not report which models were tried, what they scored, or how they did on each footage type.

Every clip comes from sites off Curacao [12]. Hein's goal goes beyond that region. "My hope is that we really can unlock the value in citizen science footage," he said [14]. A model that tops a WildFin leaderboard will have shown it can handle Curacao bommies. Tourist and naturalist video from other reefs is the production case, and it would need testing of its own.

Hence the request for more footage. The team invites ecologists to release their own field video, much of it sitting on hard drives on lab shelves, as training data [16]. Cleaning that footage is a job in itself, and Sun envisions artificial intelligence agents being used to convert it [17]. "It's a multidisciplinary problem that needs to be addressed," Grassick said. "You need statistics people, machine-learning people and ecologists all working together to figure out how to make this process less expensive, and how to still answer our questions." [15]

What to watch

  • Per-model scores in the arXiv paper, especially the gap between crowded bommie clips and single-fish diver footage.
  • Whether other labs release field video, and whether models trained on the larger pool improve on WildFin.
  • Tests on reef footage from outside Curacao, which would show whether WildFin-tuned models carry over to citizen-science video.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories