Build1 publisher3 min readPublished
UCLA's optical deepfake screen classifies at least 15 videos in one pass of light
UCLA researchers' hybrid optical-neural processor screens at least 15 video streams per pass of light at 97.79% accuracy on Celeb-DF. The experimental eLight study pitches it as a cheap first filter whose savings hinge on how many genuine videos reach the heavier detectors.
The Engineer · Build desk

What happened
- In the 15-video Celeb-DF experiment the system flagged 99.86% of manipulated videos and correctly recognised 95.72% of genuine ones.
- After its digital encoder was fine-tuned on 50 videos generated with Google's Veo 3, the system scored 94.80% on a held-out set of 105 real and 105 generated videos.
- Under noise, blur, JPEG compression and misaligned optics, detection held comparatively steady in several tested conditions and lost accuracy under severe degradation.
- Swapping in a lighter digital encoder cut estimated end-to-end energy use by 37.8% to 41.7% against the study's digital baseline, but also lowered accuracy and specificity.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost The second-stage compute bill scales with genuine-video volume, because the filter forwards a fixed share of real uploads however few fakes arrive.
- constraint A platform cannot bank the lighter encoder's energy cut and the reported specificity together; every cheaper first stage pushes more genuine video into the expensive one.
- exposure Each new generator implies a fine-tuning cycle on the encoder, and footage from a generator it has not been tuned on falls outside what the study measured.
The light handles the back half of the pipeline. According to The Neuron's account of the paper, a digital encoder extracts features from sampled frames and writes them as a pattern onto a spatial light modulator [5]. Light then passes through an optical decoder, where propagation and diffraction perform part of the classification, and sensors turn the output intensity into one score per video [6]. The stream count comes from layout [1]. Separate regions of the optics take separate videos, so one pass classifies all of them [7]. Everything before the modulator runs on ordinary electronics [8].
The 97.79% Celeb-DF accuracy [2] is the exact midpoint of the 15-video run's two rates: (99.86 + 95.72) / 2 = 97.79 [3]. Accuracy lands on that midpoint when fake and genuine videos carry equal weight. For the figure to transfer to a platform, half of its uploads would have to be fakes. On a stream where genuine video dominates, specificity decides the result. A 95.72% specificity forwards 4.28%, about four in every hundred genuine videos, to the second stage [1].
The researchers pitch the processor as a first stage that flags suspicious footage before heavier digital detectors run [3]. I think tilting it toward sensitivity is right for that job. The stated reasoning is that a fake missed at the first stage may never be checked again, while a genuine video flagged by mistake can still go to another detector [11]. At 99.86% sensitivity, the filter lets through about 14 of every 10,000 manipulated videos [2]. Any saving then depends on how much content gets flagged and how much compute the second-stage detector needs [12].
The energy accounting deserves credit. The optical decoder draws roughly 1.38 to 4.11 millijoules per video in one configuration [15]. The September 22 paper also counts the electronics: total consumption is higher because the digital encoder accounts for much of the energy [4][16]. Reporting the decoder's millijoules alone would have made a better slide. Shrinking that encoder produced the energy cuts, and the smaller builds lost specificity [17][18]. Each point of specificity lost sends one more genuine video in a hundred to the expensive detector, so part of the first-stage saving is spent again downstream. The Neuron's write-up does not give the specificity of the lighter builds.
The Veo 3 test shows the encoder can be retrained for a new generator with a small batch, and it was scored on a held-out set of 210 videos [13][4]. The write-up states that the result does not show the same accuracy carrying over to an unfamiliar generator without fine-tuning [14]. The team also ran black-box adversarial attacks, in which an attacker can query the detector but cannot inspect its parameters [21].
Deployment comes down to two conditions. Total screening cost has to fall while false negatives stay within the platform's required threshold [19]. Hardware cost and the volume passed to second-stage detectors decide whether overall resource use drops at all [19].
What to watch
- Published specificity for the lighter-encoder builds; that number decides how much of the 37.8% to 41.7% energy cut survives at the second stage.
- A test on a generator the encoder was never fine-tuned on, the case the Veo 3 result leaves open.
- Any hardware cost estimate for the spatial light modulator and optical decoder at platform scale.