Basecamp Research announced a $140 million Series C funding round on September 23, 2026. The company will use the money to develop the next generation of EDEN, a foundation model trained on gene sequences collected from nature, and to advance therapies that rewrite cells inside a patient's body.

A great deal of experimentation and validation separates an AI's ability to design a new molecule from that molecule's use as a safe drug. Comparing the company's papers with its public pipeline shows how far it aims to take development with this round of funding.

AD

How can sequences collected in the field correct the bias in public data?

EDEN learns not from text but from base sequences, the strings that represent the DNA of microbes and other organisms. The large model has 28 billion parameters and was trained on 9.7 trillion base tokens collected by Basecamp Research.

According to the EDEN preprint, published in February 2026, the training data included more than 10 billion novel genes and taxonomic units corresponding to more than 1 million species not registered in existing genome databases. These figures reflect the scale of the data at the time the model was trained and do not include sequences the company plans to collect in the future.

Why go out into the field to collect samples itself?

In its BaseData preprint, the company analyzed the public sequence database Sequence Read Archive (SRA) and found that more than 68% of the reads registered there come from just five organisms.

In medical research, sequences from humans and laboratory animals accumulate in large quantities. As training data for finding the diverse enzymes that evolution has produced, however, that leaves a large bias.

The 68% figure is the share of registered reads in the SRA. It does not refer to 68% of the species on Earth or 68% of known genes.

Basecamp Research works with research institutions in various regions to collect samples from soil, the ocean and other environments, and links the resulting gene sequences to information about where they were collected.

The company places particular weight on the genetic relationships that different organisms, such as phages and their hosts, have built up over evolutionary time. For example, when designing an enzyme that recognizes a specific DNA sequence and inserts a gene, the pairing of the enzyme with the DNA sequence it has targeted is an important clue, in addition to the enzyme's own sequence.

Still, a large amount of training data does not by itself show that something works as a therapy.

In an exclusive interview with THE DECODER, Chief Technology Officer Philipp Lorenz said the company held about 15 trillion tokens of data at the time of the interview. The 9.7 trillion tokens in the paper is the amount used earlier to train the model; the 15 trillion is the total holdings, which have grown since.

He also said that roughly one-third of the company's GPU compute is allocated to reinforcement learning experiments. The company has not published figures showing how much this has improved the performance of drug candidates.

In the same interview, Lorenz said that a model trained on the company's own collected data outperformed an identically configured model trained only on public data on a metric for predicting base sequences.

However, details of the datasets used in the comparison and the specific measured values have not been released. Even if a model predicts the next base with high accuracy, that ability does not directly indicate safety or therapeutic effect when administered to patients. To gauge potential as a drug, experimental results from molecules actually generated must be examined separately from the model's evaluation scores.

The mechanism for collecting proprietary data also raises a separate question.

According to the BaseData paper, the company signs agreements with stakeholders in the regions where it collects samples, setting out data usage rights and benefit sharing. It also says it has introduced a system that can pay royalties from the point at which collected data is used to train AI.

The company says it had distributed commercial royalties to 52 beneficiaries in 19 countries by the end of 2024, but the relevant section of the paper does not state the total amount paid or individual amounts. This makes it difficult for outsiders to compare in detail how much benefit is being returned locally.

For a company whose competitive edge lies in proprietary genetic data gathered through field surveys, a system for continuing to collect samples and for returning benefits to the regions where they are collected matters for sustaining the business as well as for the technology.

How far have the gene-insertion experiments confirmed the approach?

One application of EDEN is the design of large serine recombinases.

These enzymes recognize two DNA sequences and cause recombination between them, allowing relatively long genes to be inserted. Basecamp Research trained the model on pairings of naturally occurring enzymes and the DNA sequences they recognize, then generated new enzymes that act on specified targets.

In current CAR-T therapy, the standard approach is to remove cells from a patient, introduce genes outside the body and return them. Basecamp Research, by contrast, aims to introduce therapeutic genes directly into T cells inside the patient's body.

EDEN's foundational pretraining used no human patient or clinical trial data. The model is trained on a large volume of gene sequences from nature and then adjusted for individual uses such as gene insertion.

As a result, the model pretrained on natural sequences also yielded enzyme candidates that work inside human T cells.

Here, however, it is necessary to separate confirming a reaction in human cells from designing an enzyme to hit a new target in the human genome. The two were not confirmed under the same conditions.

Organizing Figure 3 and the main text of the EDEN preprint by the target specified when designing the enzymes and the method of later evaluation gives the following.

Target at design time Where evaluated Result reported in the paper
Known bacterial target Cell-free reaction assay 53.6% of 176 generated candidates showed significant DNA recombination activity
Known bacterial target Primary human T cells Insertion of a CD19-targeting CAR gene confirmed for half of 20 selected candidates
New target in the human genome Cell-free reaction assay Active candidates obtained for all 14 targets, including disease-related sites

For enzymes designed to match chosen positions in the human genome, activity was confirmed in vitro. Evaluating those candidates inside actual human cells, however, is described in the paper as future work.

There is an easily overlooked difference here.

The enzymes for which CAR gene insertion was confirmed in human T cells are candidates designed using known bacterial targets as a guide. For the new targets in the human genome, by contrast, the researchers specified 30-base target sequences, created new candidates and confirmed activity in a cell-free reaction assay.

The result of "half of 20 candidates" for the former therefore cannot be lined up against the activity rate for the latter to compare performance. The targets and the test methods differ.

The paper itself lists validating enzymes designed for new human genome targets in relevant human cells, and examining whether genes are inserted at unintended locations, as future tasks.

The same model also designed antimicrobial peptide candidates.

The paper reports that of 33 synthesized candidates evaluated against bacteria in vitro, 97% showed antimicrobial activity. This is not a result showing that patients with infections were treated.

That EDEN can design several types of molecules is separate from whether each candidate holds up as a safe and effective drug, and each must be validated on its own.

AD

The next question is delivery and safety, not model performance

The public pipeline as of September 2026 lists six therapeutic programs that Basecamp Research is pursuing in-house.

Four are at the lead optimization stage: in vivo CAR-T for blood cancers and autoimmune diseases, gene therapy for phenylketonuria, and antimicrobial peptides for drug-resistant bacteria. Two are at the discovery stage: CAR-T for solid tumors and a peptide for diabetes.

Based on this list, no program has yet advanced to preclinical or clinical trials.

The September funding announcement says promising preclinical results have been obtained in several areas. However, in vitro and cell-based experimental results shown in the papers cannot confirm that candidate substances can be safely delivered inside a patient's body.

For in vivo CAR-T, what matters is whether the molecules needed for editing can be delivered to the targeted T cells. It is also necessary to examine whether they act on other cells and whether genes are inserted at unintended positions in the DNA.

The paper's authors likewise list evaluating off-target gene insertion and validating enzymes designed for new targets in appropriate human cells as future tasks.

Basecamp Research has shown that it can use diverse gene sequences from nature as AI training data and design enzymes and antimicrobial peptides matched to chosen targets.

In judging whether the $140 million just raised will lead to actual therapies, the model's parameter count and the volume of training data are not the only things that matter.

Can an enzyme designed for a new target in the human genome actually insert a gene at the intended site inside human cells? How far can unintended activity be suppressed? And can the candidates in the public pipeline advance to preclinical trials?

Experimental results on these questions will show how far EDEN's technology can carry into drug discovery.