Skip to content

Invest1 publisher3 min readPublished

S32 leads $140 million round for Basecamp Research's DNA from a million unsequenced species

The Series C values Basecamp Research at about $800 million, roughly 3.6 times everything it has raised, on the strength of a proprietary sequence library and scaling comparisons the company ran internally.

The Investor · Invest desk

Photograph accompanying S32 leads $140 million round for Basecamp Research's DNA from a million unsequenced species
Photo: european-biotechnology.com

What happened

  • Basecamp Research, an AI biotech with offices in London and Boston, closed an oversubscribed $140 million Series C after six years working on the gap in public genomic databases.
  • The round takes the company's total capital raised to $225 million and values it at approximately $800 million.
  • Basecamp is using its EDEN foundation models to design large serine recombinases intended to reprogram a patient's immune cells inside the body, removing the cell-therapy manufacturing chain.
  • The company has set a target of one quadrillion DNA tokens within 18 months, drawing on field partnerships across more than 30 countries on all seven continents.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • constraint A doubling of the corpus every three months for six consecutive quarters ties Basecamp to sequencing and field logistics, and that work absorbs capital well before any designed enzyme reaches a patient.
  • exposure Nvidia's and Anthropic's venture arms both hold equity in a company whose principal asset is training data, putting two AI suppliers on the same side of the table as a prospective customer.
  • decision Cell-therapy incumbents have to weigh further spending on a laboratory process that Basecamp is now funded to try to make unnecessary, and the waitlist gives patients a reason to prefer the attempt.
  • precedent A mark set at 3.6 times all capital raised, on corpus size and internally run scaling curves, gives the next sequencing-led biotech a comparable to quote before it has a molecule in a patient.

The valuation is about 3.6 times every dollar Basecamp has raised in its life [5][1], and the new money accounts for some 17.5 percent of the post-money figure [2]. The first generation of EDEN models used roughly two thirds of the atlas as it currently stands [3]. The 18-month goal is about 67 times what the company holds today, which works out to a doubling every three months or so for six straight quarters [4].

The cost side is the laboratory bill. Approved cancer cell therapies cost $400,000 or more per patient in lab work before a hospital bill enters the picture [7][8]. The entire Series C equals the laboratory cost of 350 patients at that figure [5]. The same manufacturing step creates waitlists. TechTimes reports that a substantial share of blood cancer patients who qualify for existing CAR-T treatments die before their cells are ready on manufacturing waitlists [8].

The case for the data rests on concentration. Basecamp's own analysis puts about 68 percent of the sequence volume in the Sequence Read Archive in five species, with humans alone at around 54 percent [10]. That leaves the other four species some 14 points between them [6]. The company has spent six years on that gap [18]. Philip Lorenz, Basecamp's CTO, has said that training only on public genome data is like "teaching a language model to understand all human communication using only newspaper articles from 1975" [11].

There is less here for an outside investor to check, because the diligence runs on Basecamp's own evaluations. The steeper scaling the company claims for its proprietary data over public databases came from internal tests [15]. Lorenz also reported that an architecture with lower perplexity on standard AI benchmarks did worse on biological tasks than a Llama-based architecture that scored higher on those same metrics [16]. Basecamp now uses biological task performance at training checkpoints as its primary evaluation signal, setting aside the proxy metrics the broader field uses to compare models [17].

In my view the round prices the data supply, and the enzymes are option value on top of it. TechTimes describes the in-body reprogramming as something the recombinases could do in theory [7], and the company's pitch is explicitly conditional on the platform working in humans [9]. The counter-thesis has serious sponsors. Nvidia's NVentures came back after a January 2026 SAFE note, Anthropic's Anthology Fund joined, and so did the NATO Innovation Fund and the UK's Sovereign AI Fund [4]. None of them need a recombinase in a patient to get value from a corpus other people want to train on. Andy Conrad, the former Verily CEO, takes a board seat as an S32 general partner [3].

If a model trained on public data designs comparable large serine recombinases, the 68-and-54 concentration argument [10] stops supporting a premium, because concentration only matters if it starves enzyme discovery. That would break the first reading. If the corpus target slips, the price is left resting on 9.7 trillion nucleotide tokens and 28 billion parameters [14]. It also rests on field partnerships across more than 30 countries on all seven continents [13], and those have to keep delivering.

What to watch

  • Whether the Trillion Gene Atlas reaches one quadrillion tokens on schedule, and how much of the growth comes from new field partnerships versus deeper sequencing of existing ones.
  • Any peer-reviewed or third-party comparison of EDEN against public-data models on large serine recombinase design.
  • Whether an EDEN-designed recombinase enters a formal preclinical or clinical programme, and who funds it.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories