Teaching Gemma to read DNA, without teaching it DNA
To isolate high-value traits, plant geneticists traditionally rely on Bulk Segregant Analysis—crossing opposing parent lines, cultivating the offspring across a growing season, sorting plants into trait-bearing pools, and sequencing their genomes. While this process narrows the trait to a broad chromosomal neighborhood, it leaves scientists with an unannotated spreadsheet containing tens of thousands of candidate mutations, including letter swaps (Single Nucleotide Polymorphisms, or SNPs), insertions, and deletions.
Faced with this bottleneck, a natural question arises: why not simply train a generalist model like Gemma directly on raw DNA?
Developing reliable genomic capabilities within a generalist LLM introduces trade-offs: text-oriented tokenization may be inefficient for nucleotide sequences, supporting long inputs does not inherently capture long-range regulatory interactions, and genomic fine-tuning risks eroding the model’s core reasoning, coding, and tool-use faculties. Conversely, dedicated genomic language models (GLMs) like BOTANIC-1 excel at sequence representations but are difficult for non-technical researchers to deploy, query, and integrate into daily lab workflows.
Living Models resolves this by pairing the two: Gemma 4 serves as a conversational and agentic abstraction layer, translating natural-language biological inquiries into structured queries against BOTANIC-1’s inference endpoints.